VLDB 2026 Research / reviewers in the wild / expert
Alexander Kotov 0001
dblp:47/8024 · also Alexander S. Kotov
· DBLP profile ↗
32ranked-venue papers
9as first author
2since 2021 · last 2026
0000-0002-9872-6605ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 28 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
16 papers |
Information retrieval · 62% Knowledge graphs · 16% Data mining · 11% | |
| Artificial intelligence
3 papers |
Question answering and dialogue systems · 63% Vision and language · 37% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › search engines › semantic search
entity retrieval |
2.5 | 5 | 2026 | Conversational Entity Retrieval from a Knowledge Graph using Aggregation of Fine-grained Relevance Signals with Graph Convolutions and Self-Attention · WSDM 2026 Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph · WWW 2024 DBpedia-Entity v2: A Test Collection for Entity Search · SIGIR 2017 |
Knowledge graphs
knowledge graph querying |
1.8 | 2 | 2026 | Conversational Entity Retrieval from a Knowledge Graph using Aggregation of Fine-grained Relevance Signals with Graph Convolutions and Self-Attention · WSDM 2026 Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph · WWW 2024 |
Information retrieval
retrieval models |
1.2 | 5 | 2018 | Deep Neural Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and Images · WSDM 2018 Embedding-based Query Expansion for Weighted Sequential Dependence Retrieval Model · SIGIR 2017 Parameterized Fielded Term Dependence Models for Ad-hoc Entity Retrieval from Knowledge Graph · SIGIR 2016 |
Information retrieval › search engines › semantic search › entity retrieval
entity ranking |
0.8 | 1 | 2024 | Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph · WWW 2024 |
Information retrieval › retrieval models › language model
term dependency models |
0.5 | 2 | 2017 | Embedding-based Query Expansion for Weighted Sequential Dependence Retrieval Model · SIGIR 2017 Fielded Sequential Dependence Model for Ad-Hoc Entity Retrieval in the Web of Data · SIGIR 2015 |
Information retrieval › query reformulation
query expansion |
0.4 | 2 | 2017 | Embedding-based Query Expansion for Weighted Sequential Dependence Retrieval Model · SIGIR 2017 Tapping into knowledge base for concept feedback: leveraging conceptnet to improve search results for difficult queries · WSDM 2012 |
Information retrieval
cross-modal retrieval |
0.3 | 1 | 2018 | Deep Neural Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and Images · WSDM 2018 |
Information retrieval
multimodal retrieval |
0.3 | 1 | 2018 | Deep Neural Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and Images · WSDM 2018 |
Information retrieval
evaluation |
0.3 | 1 | 2017 | DBpedia-Entity v2: A Test Collection for Entity Search · SIGIR 2017 |
Recommender systems
point-of-interest recommendation |
0.3 | 1 | 2017 | Probabilistic Social Sequential Model for Tour Recommendation · WSDM 2017 |
Information retrieval › relevance feedback
pseudo-relevance feedback |
0.3 | 1 | 2017 | Embedding-based Query Expansion for Weighted Sequential Dependence Retrieval Model · SIGIR 2017 |
Information retrieval › evaluation
relevance judgment |
0.3 | 1 | 2017 | DBpedia-Entity v2: A Test Collection for Entity Search · SIGIR 2017 |
Recommender systems › domain-specific recommendation
route recommendation |
0.3 | 1 | 2017 | Probabilistic Social Sequential Model for Tour Recommendation · WSDM 2017 |
Recommender systems
sequential recommendation |
0.3 | 1 | 2017 | Probabilistic Social Sequential Model for Tour Recommendation · WSDM 2017 |
Information retrieval › evaluation
test collection |
0.3 | 1 | 2017 | DBpedia-Entity v2: A Test Collection for Entity Search · SIGIR 2017 |
Recommender systems › domain-specific recommendation
travel recommendation |
0.3 | 1 | 2017 | Probabilistic Social Sequential Model for Tour Recommendation · WSDM 2017 |
Information retrieval › search engines › semantic search › entity retrieval
knowledge graph entity retrieval |
0.2 | 1 | 2016 | Parameterized Fielded Term Dependence Models for Ad-hoc Entity Retrieval from Knowledge Graph · SIGIR 2016 |
Natural language and speech › Question answering and dialogue systems
conversational search |
0.2 | 1 | 2024 | Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge Graph · WWW 2024 |
Data mining › clustering › spectral clustering
approximate spectral clustering |
0.2 | 1 | 2015 | Multi-level Approximate Spectral Clustering · ICDM 2015 |
Data mining
clustering |
0.2 | 1 | 2015 | Multi-level Approximate Spectral Clustering · ICDM 2015 |
Data mining › text mining
sentiment analysis |
0.2 | 1 | 2015 | Parametric and Non-parametric User-aware Sentiment Topic Models · SIGIR 2015 |
Data mining › clustering
spectral clustering |
0.2 | 1 | 2015 | Multi-level Approximate Spectral Clustering · ICDM 2015 |
Data mining › text mining
topic modeling |
0.2 | 1 | 2015 | Parametric and Non-parametric User-aware Sentiment Topic Models · SIGIR 2015 |
Algorithms and data structures › matrix approximation
low-rank approximation |
0.2 | 1 | 2015 | Multi-level Approximate Spectral Clustering · ICDM 2015 |
Web and social media mining › event detection
burst detection |
0.1 | 1 | 2011 | Mining named entities with temporally correlated bursts from multilingual web news streams · WSDM 2011 |
Information retrieval › interactive information retrieval › session search
cross-session search |
0.1 | 1 | 2011 | Modeling and analysis of cross-session search tasks · SIGIR 2011 |
Data mining › text mining
information extraction |
0.1 | 1 | 2011 | Mining named entities with temporally correlated bursts from multilingual web news streams · WSDM 2011 |
Information retrieval
query suggestion |
0.1 | 1 | 2011 | Modeling and analysis of cross-session search tasks · SIGIR 2011 |
Information retrieval › user behavior
search behavior |
0.1 | 1 | 2011 | Modeling and analysis of cross-session search tasks · SIGIR 2011 |
Data mining › text mining
text stream mining |
0.1 | 1 | 2011 | Mining named entities with temporally correlated bursts from multilingual web news streams · WSDM 2011 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.5neural architecture · 1.5LSTM · 1.5learning to rank · 1.2self-attention · 1.0graph convolution · 1.0structured hinge loss · 0.7gated neural architecture · 0.7entity retrieval · 0.3entity linking · 0.3sampling strategies · 0.2low-rank approximation · 0.2markov-modulated poisson process · 0.1dynamic programming · 0.1ranking · 0.1content indexing · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conversational Entity Retrieval from a Knowledge Graph using Aggregation of Fine-grained Relevance Signals with Graph Convolutions and Self-AttentionabstractThe recently introduced task of Conversational Entity Retrieval from a Knowledge Graph (CER-KG) presents unique research challenges due to the complexity of the queries along with the necessity to consider KG structure and the context of an information-seeking dialog. This paper proposes a novel approach to CER-KG that first constructs a sub-graph around each candidate response entity, which includes its neighboring KG components, such as other entities, literals, categories and predicates, and then scores and ranks each candidate answer entity with Diverse Relevance signal Aggregation via Graph cONvolution (DRAGON), a novel learning-to-rank neural architecture for CER-KG. Unlike previous approaches to CER-KG, DRAGON directly takes a large number of fine-grained relevance signals as input and learns to effectively aggregate and transform those signals into the ranking scores of candidate response entities. In particular, a set of sparse and structured vectors of relevance features used as input to DRAGON measure lexical and semantic similarity between a query in the current turn or responses from the past turns of an information-seeking dialog and each node in the candidate response entity's sub-graph. DRAGON then propagates the relevance signals in feature vectors around the sub-graph using graph convolution layers and aggregates those signals into the candidate response entity ranking score with multi-head attention and fully-connected layers. This design enables DRAGON to attenuate noisy relevance signals from the local KG neighborhood during propagation and attend to the signals from the most important nodes in the candidate entity sub-graph. Our results demonstrate that DRAGON yields significant gains in retrieval accuracy over the previously proposed approach for CER-KG and performs comparably to a much larger fine-tuned cross-encoder architecture. Mona Zamiri, Alexander Kotov 0001 |
WSDM | 2 |
| 2024 | Benchmark and Neural Architecture for Conversational Entity Retrieval from a Knowledge GraphabstractThis paper introduces a novel information retrieval (IR) task of Conversational Entity Retrieval from a Knowledge Graph (CER-KG), which extends non-conversational entity retrieval from a knowledge graph (KG) to the conversational scenario. The user queries in CER-KG dialog turns may rely on the results of the preceding turns, which are KG entities. Similar to the conversational document IR, CER-KG can be viewed as a sequence of interrelated ranking tasks. To enable future research on CER-KG, we created QBLink-KG, a publicly available benchmark that was adapted from QBLink, a benchmark for text-based conversational reading comprehension of Wikipedia. As an initial approach to CER-KG, we experimented with Transformer- and LSTM-based query encoders in combination with the Neural Architecture for Conversational Entity Retrieval (NACER), our proposed feature-based neural architecture for entity ranking in CER-KG. NACER computes the ranking score of a candidate KG entity by taking into account diverse lexical and semantic matching signals between various KG components in its neighborhood, such as entities, categories, and literals, as well as entities in the results of the preceding turns in dialog history. The reported experimental results reveal the key challenges of CER-KG along with the possible directions for new approaches to this task. Mona Zamiri, Yao Qiang, Fedor Nikolaev, Dongxiao Zhu, Alexander Kotov 0001 |
WWW | 5 |
| 2020 | Joint Word and Entity Embeddings for Entity Retrieval from a Knowledge Graph
Fedor Nikolaev, Alexander Kotov 0001 |
ECIR (1) | 2 |
| 2019 | Tensor Decomposition for Sub-typing of Complex Diseases based on Clinical and Genomic DataabstractIt has long been understood that stratification of patients into fine-grained cohorts is a foundation of accurate diagnosis and effective treatment of complex diseases, such as cancer. Nevertheless, cancer therapies still fail or cause unnecessary suffering to many patients, which suggests that our current understanding of cancer sub-types needs to be refined. In this paper, we propose CLIGEN, a novel computational pipeline for high-throughput data-driven stratification of patients with a complex disease into cohorts corresponding to multi-modal disease sub-types based on clinical and genomic data. We applied CLIGEN to discover breast cancer sub-types based on the clinical and genomic data of 503 patients with breast ductal carcinoma in the Cancer Genome Atlas (TCGA). Quantitative and qualitative evaluation of the breast cancer sub-types discovered by CLIGEN indicate that they are biologically meaningful and correlate with clinical outcomes, such as patient survival time. Diana Diaz, Aliccia Bollig-Fischer, Alexander Kotov 0001 |
BIBM | 3 |
| 2019 | Bayesian approach to incorporating different types of biomedical knowledge bases into information retrieval systems for clinical decision support in precision medicineabstractBy providing clinicians with information regarding treatment options for molecular sub-types of complex diseases with genetic origin, such as cancer, information retrieval (IR) systems play an important role in precision medicine. In this paper, we propose Bayesian Precision Medicine (BPM), a novel probabilistic framework for query expansion in information retrieval systems for Clinical Decision Support (CDS) in Precision Medicine (PM). Such systems can assist clinicians with selecting personalized treatment of complex diseases based on the patients' genomic data, such as gene mutations. In particular, we focus on a clinical decision support scenario in which clinicians provide two types of information in their queries: (1) short description of a patient's case, which may contain information regarding the type of cancer that a patient has as well as symptoms and demographics, and (2) gene mutations, which may contain gene names, mutation code and type of mutation. The goal of an IR system in this scenario is to rank biomedical articles from a large collection, such as the MEDLINE, based on their relevance to the provided query. One of the main challenges faced by IR systems in this scenario is semantic matching of heterogeneous information (gene names, medical terminology and other query keywords) in queries and relevant biomedical articles. To address this challenge, we propose a probabilistic framework that enables mapping gene mutations provided in a given query onto the biomedical concepts that are related to the entire query and can be effectively utilized for query expansion. The BPM obtains candidate query expansion concepts from biomedical knowledge bases, the Unified Medical Language System (UMLS) and the Drug-Gene Interaction Database (DGIdb), as well as the top-ranked MEDLINE articles retrieved for the original query. The BPM then utilizes information from the Catalog of Somatic Mutations in Cancer (COSMIC) and co-occurrence statistics in MEDLINE to assess the relatedness of candidate query expansion concepts to gene mutations and other information provided in a query. Experimental evaluation of the BPM was conducted on a large subset of MEDLINE articles as well as abstracts from the American Association for Cancer Research (AACR) and American Society of Clinical Oncology (ASCO) proceedings. Experimental results on a publicly available benchmark provided by the 2017 TREC precision medicine track indicate that the proposed probabilistic framework is effective at utilizing both genomic and textual information in queries to improve the accuracy of IR systems for CDS in PM through query expansion. Saeid Balaneshinkordan, Alexander Kotov 0001 |
J. Biomed. Informatics | 2 |
| 2018 | Attentive Neural Architecture for Ad-hoc Structured Document RetrievalabstractThe problem of ad-hoc structured document retrieval arises in many information access scenarios, from Web to product search. Yet neither deep neural networks, which have been successfully applied to ad-hoc information retrieval and Web search, nor the attention mechanism, which has been shown to significantly improve the performance of deep neural networks on natural language processing tasks, have been explored in the context of this problem. In this paper, we propose a deep neural architecture for ad-hoc structured document retrieval, which utilizes attention mechanism to determine important phrases in keyword queries as well as the relative importance of matching those phrases in different fields of structured documents. Experimental evaluation on publicly available collections for Web document, product and entity retrieval from knowledge graphs indicates superior retrieval accuracy of the proposed neural architecture relative to both state-of-the-art neural architectures for ad-hoc document retrieval and probabilistic models for ad-hoc structured document retrieval. Saeid Balaneshinkordan, Alexander Kotov 0001, Fedor Nikolaev |
CIKM | 2 |
| 2018 | Utilizing Knowledge Graphs for Text-Centric Information RetrievalabstractThe past decade has witnessed the emergence of several publicly available and proprietary knowledge graphs (KGs). The depth and breadth of content in these KGs made them not only rich sources of structured knowledge by themselves, but also valuable resources for search systems. A surge of recent developments in entity linking and entity retrieval methods gave rise to a new line of research that aims at utilizing KGs for text-centric retrieval applications. This tutorial is the first to summarize and disseminate the progress in this emerging area to industry practitioners and researchers. Laura Dietz, Alexander Kotov 0001, Edgar Meij |
SIGIR | 2 |
| 2018 | Deep Neural Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and ImagesabstractRecent advances in deep learning and distributed representations of images and text have resulted in the emergence of several neural architectures for cross-modal retrieval tasks, such as searching collections of images in response to textual queries and assigning textual descriptions to images. However, the multi-modal retrieval scenario, when a query can be either a text or an image and the goal is to retrieve both a textual fragment and an image, which should be considered as an atomic unit, has been significantly less studied. In this paper, we propose a gated neural architecture to project image and keyword queries as well as multi-modal retrieval units into the same low-dimensional embedding space and perform semantic matching in this space. The proposed architecture is trained to minimize structured hinge loss and can be applied to both cross- and multi-modal retrieval. Experimental results for six different cross- and multi-modal retrieval tasks obtained on publicly available datasets indicate superior retrieval accuracy of the proposed architecture in comparison to the state-of-art baselines. Saeid Balaneshinkordan, Alexander Kotov 0001 |
WSDM | 2 |
| 2017 | Sentence Retrieval with Sentiment-specific Topical Anchoring for Review SummarizationabstractWe propose Topic Anchoring-based Review Summarization (TARS), a two-step extractive summarization method, which creates review summaries from the sentences that represent the most important aspects of a review. In the first step, the proposed method utilizes Topic Aspect Sentiment Model (TASM), a novel sentiment-topic model, to identify aspects of sentiment-specific topics in a collection of reviews. The output of TASM is utilized in the second step of TARS to rank review sentences based on how representative of the most important review aspects their words are. Qualitative and quantitative evaluation of review summaries using two collections indicate the effectiveness of structuring review summaries around aspects of sentiment-specific topics. Jiaxing Tan, Alexander Kotov 0001, Rojiar Pir Mohammadiani, Yumei Huo |
CIKM | 2 |
| 2017 | Embedding-based Query Expansion for Weighted Sequential Dependence Retrieval ModelabstractAlthough information retrieval models based on Markov Random Fields (MRF), such as Sequential Dependence Model and Weighted Sequential Dependence Model (WSDM), have been shown to outperform bag-of-words probabilistic and language modeling retrieval models by taking into account term dependencies, it is not known how to effectively account for term dependencies in query expansion methods based on pseudo-relevance feedback (PRF) for retrieval models of this type. In this paper, we propose Semantic Weighted Dependence Model (SWDM), a PRF based query expansion method for WSDM, which utilizes distributed low-dimensional word representations (i.e., word embeddings). Our method finds the closest unigrams to each query term in the embedding space and top retrieved documents and directly incorporates them into the retrieval function of WSDM. Experiments on TREC datasets indicate statistically significant improvement of SWDM over state-of-the-art MRF retrieval models, PRF methods for MRF retrieval models and embedding based query expansion methods for bag-of-words retrieval models. Saeid Balaneshinkordan, Alexander Kotov 0001 |
SIGIR | 2 |
| 2017 | DBpedia-Entity v2: A Test Collection for Entity SearchabstractThe DBpedia-entity collection has been used as a standard test collection for entity search in recent years. We develop and release a new version of this test collection, DBpedia-Entity v2, which uses a more recent DBpedia dump and a unified candidate result pool from the same set of retrieval models. Relevance judgments are also collected in a uniform way, using the same group of crowdsourcing workers, following the same assessment guidelines. The result is an up-to-date and consistent test collection.To facilitate further research, we also provide details about the pre-processing and indexing steps, and include baseline results from both classical and recently developed entity search methods. Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov 0001, Jamie Callan |
SIGIR | 6 |
| 2017 | Utilizing Knowledge Graphs in Text-centricInformation RetrievalabstractThe past decade has witnessed the emergence of several publicly available and proprietary knowledge graphs (KGs). The increasing depth and breadth of content in KGs makes them not only rich sources of structured knowledge by themselves but also valuable resources for search systems. A surge of recent developments in entity linking and retrieval methods gave rise to a new line of research that aims at utilizing KGs for text-centric retrieval applications, making this an ideal time to pause and report current findings to the community, summarizing successful approaches, and soliciting new ideas. This tutorial is the first to disseminate the progress in this emerging field to researchers and practitioners. Laura Dietz, Alexander Kotov 0001, Edgar Meij |
WSDM | 2 |
| 2017 | Probabilistic Social Sequential Model for Tour RecommendationabstractThe pervasive growth of location-based services such as Foursquare and Yelp has enabled researchers to incorpo- rate better personalization into recommendation models by leveraging the geo-temporal breadcrumbs left by a plethora of travelers. In this paper, we explore Travel path recommendation, which is one of the applications of intelligent urban navigation that aims in recommending sequence of point of interest (POIs) to tourists. Currently, travelers rely on a tedious and time-consuming process of searching the web, browsing through websites such as Trip Advisor, and reading travel blogs to compile an itinerary. On the other hand, people who do not plan ahead of their trip find it extremely difficult to do this in real-time since there are no automated systems that can provide personalized itinerary for travelers. To tackle this problem, we propose a tour recommendation model that uses a probabilistic generative framework to incorporate user's categorical preference, influence from their social circle, the dynamic travel transitions (or patterns) and the popularity of venues to recommend sequence of POIs for tourists. Through comprehensive experiments over a rich dataset of travel patterns from Foursquare, we show that our model is capable of outperforming the state-of-the-art probabilistic tour recommendation model by providing contextual and meaningful recommendation for travelers. Vineeth Rakesh, Niranjan Jadhav, Alexander Kotov 0001, Chandan K. Reddy |
WSDM | 3 |
| 2016 | Scheduling big data workflows in the cloud under budget constraintsabstractBig data is fast becoming a ubiquitous term in both academia and industry and there is a strong need for new data-centric workflow tools and techniques to process and analyze large-scale complex datasets that are growing exponentially. On the other hand, the unbound resource leasing capability foreseen in the cloud facilitates data scientists to wring actionable insights from the data in a time and cost efficient manner. In the data-centric workflow environment, scheduling data processing tasks onto appropriate resources are often driven by the constraints provided by the users. Enforcing a constraint while executing the workflow in the cloud adds a new optimization challenge on how to meet the objective while satisfying the given constraint. In this paper, we propose a new Big dAta woRkflow schEduler uNder budgeT constraint known as BARENTS that supports high-performance workflow scheduling in a heterogeneous cloud computing environment with a single objective to minimize the workflow makespan under a provided budget constraint. Our case study and experiments show the competitive advantages of our proposed scheduler. The proposed BARENTS scheduler is implemented in a new release of DATA VIEW, one of the most usable big data workflow systems in the community. Aravind Mohan, Mahdi Ebrahimi, Shiyong Lu, Alexander Kotov 0001 |
IEEE BigData | 4 |
| 2016 | Sequential Query Expansion using Concept GraphabstractManually and automatically constructed concept graphs (or semantic networks), in which the nodes correspond to words or phrases and the typed edges designate semantic relationships between words and phrases, have been previously shown to be rich sources of effective latent concepts for query expansion. However, finding good expansion concepts for a given query in large and dense concept graphs is a challenging problem, since the number of candidate concepts that are related to query terms and phrases and need to be examined increases exponentially with the distance from the original query concepts. In this paper, we propose a two-stage feature-based method for sequential selection of the most effective concepts for query expansion from a concept graph. In the first stage, the proposed method weighs the concepts according to different types of computationally inexpensive features, including collection and concept graph statistics. In the second stage, a sequential concept selection algorithm utilizing more expensive features is applied to find the most effective expansion concepts at different distances from the original query concepts. Experiments on TREC datasets of different type indicate that the proposed method achieves significant improvement in retrieval accuracy over state-of-the-art methods for query expansion using concept graphs. Saeid Balaneshinkordan, Alexander Kotov 0001 |
CIKM | 2 |
| 2016 | A Comparative Study of Query-biased and Non-redundant Snippets for Structured Search on Mobile DevicesabstractTo investigate what kind of snippets are better suited for structured search on mobile devices, we built an experimental mobile search application and conducted a task-oriented interactive user study with 36 participants. Four different versions of a search engine result page (SERP) were compared by varying the snippet type (query-biased vs. non-redundant) and the snippet length (two vs. four lines per result). We adopted a within-subjects experiment design and made each participant do four realistic search tasks using different versions of the application. During the study sessions, we collected search logs, "think-aloud" comments, and post-task surveys. Each session was finalized with an interview. We found that with non-redundant snippets the participants were able to complete the tasks faster and find more relevant results. Most participants preferred non-redundant snippets and wanted to see more information about each result on the SERP for any snippet type. Yet, the participants felt that the version with query-biased snippets was easier to use. We conclude with a set of practical design recommendations. Nikita Spirin, Alexander Kotov 0001, Karrie Karahalios, Vassil Mladenov, Pavel A. Izhutov |
CIKM | 2 |
| 2016 | An Empirical Comparison of Term Association and Knowledge Graphs for Query Expansion
Saeid Balaneshinkordan, Alexander Kotov 0001 |
ECIR | 2 |
| 2016 | Feedback or Research: Separating Pre-purchase from Post-purchase Consumer Reviews
Alexander Kotov 0001, Aravind Mohan, Shiyong Lu, Paul M. Stieg |
ECIR | 2 |
| 2016 | Parameterized Fielded Term Dependence Models for Ad-hoc Entity Retrieval from Knowledge GraphabstractAccurate projection of terms in free-text queries onto structured entity representations is one of the fundamental problems in entity retrieval from knowledge graphs. In this paper, we demonstrate that existing retrieval models for ad-hoc structured and unstructured document retrieval fall short of addressing this problem, due to their rigid assumptions. According to these assumptions, either all query concepts of the same type (unigrams and bigrams) are projected onto the fields of entity representations with identical weights or such projection is determined based only on one simple statistic, which makes it sensitive to data sparsity. To address this issue, we propose the Parametrized Fielded Sequential Dependence Model (PFSDM) and the Parametrized Fielded Full Dependence Model (PFFDM), two novel models for entity retrieval from knowledge graphs, which infer the user's intent behind each individual query concept by dynamically estimating its projection onto the fields of structured entity representations based on a small number of statistical and linguistic features. Experimental results obtained on several publicly available benchmarks indicate that PFSDM and PFFDM consistently outperform state-of-the-art retrieval models for the task of entity retrieval from knowledge graph. Fedor Nikolaev, Alexander Kotov 0001, Nikita Zhiltsov |
SIGIR | 2 |
| 2016 | A study of the effectiveness of machine learning methods for classification of clinical interview fragments into a large number of categories
Alexander Kotov 0001, April Idalski Carcone, Ming Dong 0001, Sylvie Naar, Kathryn Brogan Hartlieb |
J. Biomed. Informatics | 2 |
| 2015 | Interpretable Probabilistic Latent Variable Models for Automatic Annotation of Clinical Text
Alexander Kotov 0001, April Idalski Carcone, Ming Dong 0001, Sylvie Naar, Kathryn Brogan Hartlieb |
AMIA | 1 |
| 2015 | Geographical Latent Variable Models for Microblog Retrieval
Alexander Kotov 0001, Vineeth Rakesh, Eugene Agichtein, Chandan K. Reddy |
ECIR | 1 |
| 2015 | Multi-level Approximate Spectral ClusteringabstractClustering is a task of finding natural groups in datasets based on measured or perceived similarity between data points. Spectral clustering is a well-known graph-theoretic approach, which is capable of capturing non-convex geometries of datasets. However, it generally becomes infeasible for analyzing large datasets due to relatively high time and space complexity. In this paper, we propose Multi-level Approximate Spectral (MAS) clustering to enable efficient analysis of large datasets. By integrating a series of low-rank matrix approximations (i.e., approximations to the affinity matrix and its subspace, as well as those for the Laplacian matrix and the Laplacian subspace), MAS achieves great computational and spacial efficiency. MAS provides a general framework for fast and accurate spectral clustering, which works with any kernels, various fast sampling strategies and different low-rank approximation algorithms. In addition, it can be easily extended for distributed computing. From a theoretical perspective, we provide rigorous analysis of its approximation error in addition to its correctness and computational complexity. Through extensive experiments we demonstrate superior performance of the proposed method relative to several well-known approximate spectral clustering algorithms. Ming Dong 0001, Alexander Kotov 0001 |
ICDM | 3 |
| 2015 | Parametric and Non-parametric User-aware Sentiment Topic ModelsabstractThe popularity of Web 2.0 has resulted in a large number of publicly available online consumer reviews created by a demographically diverse user base. Information about the authors of these reviews, such as age, gender and location, provided by many on-line consumer review platforms may allow companies to better understand the preferences of different market segments and improve their product design, manufacturing processes and marketing campaigns accordingly. However, previous work in sentiment analysis has largely ignored these additional user meta-data. To address this deficiency, in this paper, we propose parametric and non-parametric User-aware Sentiment Topic Models (USTM) that incorporate demographic information of review authors into topic modeling process in order to discover associations between market segments, topical aspects and sentiments. Qualitative examination of the topics discovered using USTM framework in the two datasets collected from popular online consumer review platforms as well as quantitative evaluation of the methods utilizing those topics for the tasks of review sentiment classification and user attribute prediction both indicate the utility of accounting for demographic information of review authors in opinion mining. Zaihan Yang, Alexander Kotov 0001, Aravind Mohan, Shiyong Lu |
SIGIR | 2 |
| 2015 | Fielded Sequential Dependence Model for Ad-Hoc Entity Retrieval in the Web of DataabstractPreviously proposed approaches to ad-hoc entity retrieval in the Web of Data (ERWD) used multi-fielded representation of entities and relied on standard unigram bag-of-words retrieval models. Although retrieval models incorporating term dependencies have been shown to be significantly more effective than the unigram bag-of-words ones for ad hoc document retrieval, it is not known whether accounting for term dependencies can improve retrieval from the Web of Data. In this work, we propose a novel retrieval model that incorporates term dependencies into structured document retrieval and apply it to the task of ERWD. In the proposed model, the document field weights and the relative importance of unigrams and bigrams are optimized with respect to the target retrieval metric using a learning-to-rank method. Experiments on a publicly available benchmark indicate significant improvement of the accuracy of retrieval results by the proposed model over state-of-the-art retrieval models for ERWD. Nikita Zhiltsov, Alexander Kotov 0001, Fedor Nikolaev |
SIGIR | 2 |
| 2013 | The importance of being socially-savvy: quantifying the influence of social networks on microblog retrievalabstractSocial media users create virtual connections for various reasons: personal and professional. While significant research efforts have been spent on exploring the dynamics of creation of social network connections, little is known about how those connections influence the content generated by social media users. In this work, we quantitatively evaluate the influence of social networks on social media content providers. Additionally, we propose several document expansion methods, which leverage the content generated by the social networks of the authors of social media documents and compare their effectiveness. Experimental results on a large sample of Twitter data indicate that retrieval models discriminatively leveraging social network content for document expansion outperform both traditional, socially-unaware retrieval models and retrieval models that indiscriminatively utilize all social connections. Alexander Kotov 0001, Eugene Agichtein |
CIKM | 1 |
| 2012 | Tapping into knowledge base for concept feedback: leveraging conceptnet to improve search results for difficult queriesabstractQuery expansion is an important and commonly used technique for improving Web search results. Existing methods for query expansion have mostly relied on global or local analysis of document collection, click-through data, or simple ontologies such as WordNet. In this paper, we present the results of a systematic study of the methods leveraging the ConceptNet knowledge base, an emerging new Web resource, for query expansion. Specifically, we focus on the methods leveraging ConceptNet to improve the search results for poorly performing (or difficult) queries. Unlike other lexico-semantic resources, such as WordNet and Wikipedia, which have been extensively studied in the past, ConceptNet features a graph-based representation model of commonsense knowledge, in which the terms are conceptually related through rich relational ontology. Such representation structure enables complex, multi-step inferences between the concepts, which can be applied to query expansion. We first demonstrate through simulation experiments that expanding queries with the related concepts from ConceptNet has great potential for improving the search results for difficult queries. We then propose and study several supervised and unsupervised methods for selecting the concepts from ConceptNet for automatic query expansion. The experimental results on multiple data sets indicate that the proposed methods can effectively leverage ConceptNet to improve the retrieval performance of difficult queries both when used in isolation as well as in combination with pseudo-relevance feedback. Alexander Kotov 0001, ChengXiang Zhai |
WSDM | 1 |
| 2011 | Interactive sense feedback for difficult queriesabstractAmbiguity of query terms is a common cause of inaccurate retrieval results. Existing work has mostly focused on studying how to improve retrieval accuracy by automatically resolving word sense ambiguity. However, fully automatic sense identification and disambiguation is a very challenging task. In this work, we propose to involve a user in the process of disambiguation through interactive sense feedback and study the potential effectiveness of this novel feedback strategy. We propose several general methods to automatically identify the major senses of query terms based on global analysis of document collection and generate concise representations of the discovered senses to the users. This feedback strategy does not rely on initial retrieval results, and thus can be especially useful for improving the results of difficult queries. We evaluated the effectiveness of the proposed methods for sense identification and presentation through simulation experiments and user studies, which both indicate that sense feedback strategy is a promising alternative to the existing interactive feedback techniques such as relevance feedback and term feedback. Alexander Kotov 0001, ChengXiang Zhai |
CIKM | 1 |
| 2011 | Modeling and analysis of cross-session search tasksabstractThe information needs of search engine users vary in complexity, depending on the task they are trying to accomplish. Some simple needs can be satisfied with a single query, whereas others require a series of queries issued over a longer period of time. While search engines effectively satisfy many simple needs, searchers receive little support when their information needs span session boundaries. In this work, we propose methods for modeling and analyzing user search behavior that extends over multiple search sessions. We focus on two problems: (i) given a user query, identify all of the related queries from previous sessions that the same user has issued, and (ii) given a multi-query task for a user, predict whether the user will return to this task in the future. We model both problems within a classification framework that uses features of individual queries and long-term user search behavior at different granularity. Experimental evaluation of the proposed models for both tasks indicates that it is possible to effectively model and analyze cross-session search behavior. Our findings have implications for improving search for complex information needs and designing search engine features to support cross-session search tasks. Alexander Kotov 0001, Paul N. Bennett, Ryen W. White, Susan T. Dumais, Jaime Teevan |
SIGIR | 1 |
| 2011 | Mining named entities with temporally correlated bursts from multilingual web news streamsabstractIn this work, we study a new text mining problem of discovering named entities with temporally correlated bursts of mention counts in multiple multilingual Web news streams. Mining named entities with temporally correlated bursts of mention counts in multilingual text streams has many interesting and important applications, such as identification of the latent events, attracting the attention of on-line media in different countries, and valuable linguistic knowledge in the form of transliterations. While mining "bursty" terms in a single text stream has been studied before, the problem of detecting terms with temporally correlated bursts in multilingual Web streams raises two new challenges: (i) correlated terms in multiple streams may have bursts that are of different orders of magnitude in their intensity and (ii) bursts of correlated terms may be separated by time gaps. We propose a two-stage method for mining items with temporally correlated bursts from multiple data streams, which addresses both challenges. In the first stage of the method, the temporal behavior of different entities is normalized by modeling them with the Markov-Modulated Poisson Process. In the second stage, a dynamic programming algorithm is used to discover correlated bursts of different items, that can be potentially separated by time gaps. We evaluated our method with the task of discovering transliterations of named entities from multilingual Web news streams. Experimental results indicate that our method can not only effectively discover named entities with correlated bursts in multilingual Web news streams, but also outperforms two state-of-the-art baseline methods for unsupervised discovery of transliterations in static text collections. Alexander Kotov 0001, ChengXiang Zhai, Richard Sproat |
WSDM | 1 |
| 2010 | Temporal query log profiling to improve web search rankingabstractTemporal information can be leveraged and incorporated to improve web search ranking. In this work, we propose a method to improve the ranking of search results by identifying the fundamental properties of temporal behavior of low-quality hosts and spam-prone queries in search logs and modeling those properties as quantifiable features. In particular, we introduce the concepts of host churn, a measure of changes in host visibility for user queries, and query volatility, a measure of semantic instability of query results, and propose the methods for construction of temporal profiles from search query logs that can be used for estimation of a set of features based on the introduced concepts. The utility of the proposed concepts has been experimentally demonstrated for two language-independent search tasks: the regression-based ranking of search results and a novel classification problem of detecting spam-prone queries introduced in this work. Alexander Kotov 0001, Pranam Kolari, Lei Duan, Yi Chang 0001 |
CIKM | 1 |
| 2010 | Towards natural question guided searchabstractWeb search is generally motivated by an information need. Since asking well-formulated questions is the fastest and the most natural way to obtain information for human beings, almost all the queries posed to search engines correspond to some underlying questions, which represent the information need. Accurate determination of these questions may substantially improve the quality of search results and usability of search interfaces. In this paper, we propose a new framework for question-guided search, in which a retrieval system would automatically generate potentially interesting questions to the users. Since the answers to such questions are known to exist in search results, these questions can potentially guide the users directly to the answers they are looking for, eliminating the need to scan the documents in the results list. Moreover, in case of imprecise or ambiguous queries, automaticallygenerated questions can naturally engage the users into feedback cycles to refine their information need and guide them towards their search goals. Implementation of the proposed strategy raises new challenges in content indexing, question generation, ranking and feedback. We proposed new methods to address these challenges and evaluated them with a prototype system on a subset of Wikipedia. The evaluation results show the promise of this new question-guided search strategy. Categories andSubjectDescriptors Alexander Kotov 0001, ChengXiang Zhai |
WWW | 1 |