VLDB 2026 Research / reviewers in the wild / expert
Maksim Tkatchenko
dblp:03/9583 · also Maksim Tkachenko
· DBLP profile ↗
10ranked-venue papers
6as first author
3since 2021 · last 2023
0000-0001-6687-0525ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Semantically Constitutive Entities in Knowledge Graphs
Chong Cher Chia, Maksim Tkatchenko, Hady Wirawan Lauw |
DEXA (1) | 2 |
| 2023 | Robust Bidirectional Poly-MatchingabstractA fundamental problem in many scenarios is to match entities across two data sources. It is frequently presumed in prior work that entities to be matched are of comparable granularity. In this work, we address one-to-many or poly-matching in the scenario where entities have varying granularity. A distinctive feature of our problem is its bidirectional nature, where the ‘one’ or the ‘many’ could come from either source arbitrarily. Moreover, to deal with diverse entity representations that give rise to noisy similarity values, we incorporate novel notions of receptivity and reclusivity into a robust matching objective. As the optimal solution to the resulting formulation is proven computationally intractable, we propose more scalable yet still performant heuristics. Experiments on multiple real-life datasets showcase the effectiveness and outperformance of our proposed algorithms over baselines. Ween Jiann Lee, Maksim Tkatchenko, Hady Wirawan Lauw |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Robust BiPoly-Matching for Multi-Granular EntitiesabstractEntity matching across two data sources is a prevalent need in many domains, including e-commerce. Of interest is the scenario where entities have varying granularity, e.g., a coarse product category may match multiple finer categories. Previous work in one-to-many matching generally presumes the ‘one’ necessarily comes from a designated source and the ‘many’ from the other source. In contrast, we propose a novel formulation that allows concurrent one-to-many bidirectional matching in any direction. Beyond flexibility, we also seek matching that is more robust to noisy similarity values arising from diverse entity descriptions, by introducing receptivity and reclusivity notions. In addition to an optimal formulation, we also propose an efficient and performant heuristic. Experiments on multiple real-life datasets from e-commerce sources showcase the effectiveness and outperformance of our proposed algorithms over baselines. Ween Jiann Lee, Maksim Tkatchenko, Hady Wirawan Lauw |
ICDM | 2 |
| 2019 | CompareLDA: A Topic Model for Document ComparisonabstractA number of real-world applications require comparison of entities based on their textual representations. In this work, we develop a topic model supervised by pairwise comparisons of documents. Such a model seeks to yield topics that help to differentiate entities along some dimension of interest, which may vary from one application to another. While previous supervised topic models consider document labels in an independent and pointwise manner, our proposed Comparative Latent Dirichlet Allocation (CompareLDA) learns predictive topic distributions that comply with the pairwise comparison observations. To fit the model, we derive a maximum likelihood estimation method via augmented variational approximation algorithm. Evaluation on several public datasets underscores the strengths of CompareLDA in modelling document comparisons. Maksim Tkatchenko, Hady Wirawan Lauw |
AAAI | 1 |
| 2018 | Searching for the X-Factor: Exploring Corpus Subjectivity for Word EmbeddingsabstractWe explore the notion of subjectivity, and hypothesize that word embeddings learnt from input corpora of varying levels of subjectivity behave differently on natural language processing tasks such as classifying a sentence by sentiment, subjectivity, or topic.Through systematic comparative analyses, we establish this to be the case indeed.Moreover, based on the discovery of the outsized role that sentiment words play on subjectivity-sensitive tasks such as sentiment classification, we develop a novel word embedding SentiVec which is infused with sentiment information from a lexical resource, and is shown to outperform baselines on such tasks. Maksim Tkatchenko, Chong Cher Chia, Hady Wirawan Lauw |
ACL (1) | 1 |
| 2017 | Comparative Relation Generative ModelabstractOnline reviews are important decision aids to consumers. Other than helping users to evaluate individual products, reviews also support comparison shopping by comparing two (or more) products based on a specific aspect. However, making a comparison across two different reviews, written by different authors, is not always equitable due to the different standards and preferences of authors. Therefore, we focus on comparative sentences, whereby two products are compared directly by a review author within a sentence. We study the problem of comparative relation mining. Given a set of comparative sentences, each relating a pair of entities, our objective is three-fold: to interpret the comparative direction in each sentence, to identify the aspect of each sentence, and to determine the relative merits of each entity with respect to that aspect. This requires mining comparative relations at two levels of resolution: at the sentence level, and at the entity level. Our insight is that there is a significant synergy between the two levels. We propose a generative model for comparative text, which jointly models comparative directions at the sentence level, and ranking at the entity level. This model is tested comprehensively on Amazon reviews dataset with good empirical outperformance over pipelined baselines. Maksim Tkatchenko, Hady Wirawan Lauw |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Plackett-Luce Regression Mixture Model for Heterogeneous RankingsabstractLearning to rank is an important problem in many scenarios, such as information retrieval, natural language processing, recommender systems, etc. The objective is to learn a function that ranks a number of instances based on their features. In the vast majority of the learning to rank literature, there is an implicit assumption that the population of ranking instances are homogeneous, and thus can be modeled by a single central ranking function. In this work, we are concerned with learning to rank for a heterogeneous population, which may consist of a number of sub-populations, each of which may rank objects differently. Because these sub-populations are not known in advance, and are effectively latent, the problem turns into simultaneously learning both a set of ranking functions, as well as the latent assignment of instances to functions. To address this problem in a joint manner, we develop a probabilistic graphical model called Plackett-Luce Regression Mixture or PLRM model, and describe its inference via Expectation-Maximization algorithm. Comprehensive experiments on publicly-available real-life datasets showcase the effectiveness of PLRM, as opposed to a pipelined approach of clustering followed by learning to rank, as well as approaches that assume a single ranking function for a heterogeneous population. Maksim Tkatchenko, Hady Wirawan Lauw |
CIKM | 1 |
| 2015 | A Convolution Kernel Approach to Identifying Comparisons in TextabstractMaksim Tkachenko, Hady Lauw. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Maksim Tkatchenko, Hady Wirawan Lauw |
ACL (1) | 1 |
| 2014 | Generative Modeling of Entity Comparisons in TextabstractUsers frequently rely on online reviews for decision making. In addition to allowing users to evaluate the quality of individual products, reviews also support comparison shopping. One key user activity is to compare two (or more) products based on a specific aspect. However, making a comparison across two different reviews, written by different authors, is not always equitable due to the different standards and preferences of individual authors. Therefore, we focus instead on comparative sentences, whereby two products are compared directly by a review author within a single sentence. Maksim Tkatchenko, Hady Wirawan Lauw |
CIKM | 1 |
| 2013 | Introducing Baselines for Russian Named Entity Recognition
Rinat Gareev, Maksim Tkatchenko, Valery D. Solovyev, Andrey Simanovsky, Vladimir Ivanov 0001 |
CICLing (1) | 2 |