Alexey Tikhonov

dblp:82/8978 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
13since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 AutoVyaz: Automating the Formation of Slavic Calligraphy Ligatures
abstract
This paper introduces a procedural technique for the automated generation of Slavic Vyaz, a traditional Cyrillic calligraphy style known for its ornamental complexity and cultural significance. By developing a specialized description of font geometry that includes anchor points and edges to define character shapes, our approach facilitates the creation of ligatures and character deformations that maintain the aesthetic and rhythmic qualities of Vyaz. We propose several heuristics for arranging text layouts that enable the generation of calligraphic patterns, which can be customized and adapted to various design needs. Our technique’s capability to simulate traditional calligraphy is demonstrated through comparative analyses and examples of preliminary results.
Alexey Tikhonov
FDG1
2024 Humor Mechanics: Advancing Humor Generation with Multistep Reasoning
Alexey Tikhonov, Pavel Shtykovskiy
ICCC1
2023 Connecting degree and polarity: An artificial language learning study
abstract
We investigate a new linguistic generalisation in pre-trained language models (taking BERT Devlin et al. 2019 as a case study).We focus on degree modifiers (expressions like slightly, very, rather, extremely) and test the hypothesis that the degree expressed by a modifier (low, medium or high degree) is related to the modifier's sensitivity to sentence polarity (whether it shows preference for affirmative or negative sentences or neither).To probe this connection, we apply the Artificial Language Learning experimental paradigm from psycholinguistics to a neural language model.Our experimental results suggest that BERT generalizes in line with existing linguistic observations that relate degree semantics to polarity sensitivity, including the main one: low degree semantics is associated with preference towards positive polarity.
Lisa Bylinina, Alexey Tikhonov, Ekaterina Garmash
EMNLP2
2023 Fine-tuning transformers: Vocabulary transfer
Vladislav D. Mosin, Igor Samenko, Borislav Kozlovskii, Alexey Tikhonov, Ivan P. Yamshchikov
Artif. Intell.4
2022 Transformers in the loop: Polarity in neural models of language
abstract
Representation of linguistic phenomena in computational language models is typically assessed against the predictions of existing linguistic theories of these phenomena.Using the notion of polarity as a case study, we show that this is not always the most adequate set-up.We probe polarity via so-called 'negative polarity items' (in particular, English any) in two pretrained Transformer-based models (BERT and GPT-2).We show that -at least for polaritymetrics derived from language models are more consistent with data from psycholinguistic experiments than linguistic theory predictions.Establishing this allows us to more adequately evaluate the performance of language models and also to use language models to discover new insights into natural language grammar beyond existing linguistic theories.This work contributes to establishing closer ties between psycholinguistic experiments and experiments with language models.
Lisa Bylinina, Alexey Tikhonov
ACL (1)2
2022 The driving forces of polarity-sensitivity: Experiments with multilingual pre-trained neural language models
Lisa Bylinina, Alexey Tikhonov
CogSci2
2022 BERT in Plutarch's Shadows
abstract
The extensive surviving corpus of the ancient scholar Plutarch of Chaeronea (ca.45-120 CE) also contains several texts which, according to current scholarly opinion, did not originate with him and are therefore attributed to an anonymous author Pseudo-Plutarch.These include, in particular, the work Placita Philosophorum (Quotations and Opinions of the Ancient Philosophers), which is extremely important for the history of ancient philosophy.Little is known about the identity of that anonymous author and its relation to other authors from the same period.This paper presents a BERT language model for Ancient Greek.The model discovers previously unknown statistical properties relevant to these literary, philosophical, and historical problems and can shed new light on this authorship question.In particular, the Placita Philosophorum, together with one of the other Pseudo-Plutarch texts, shows similarities with the texts written by authors from an Alexandrian context (2nd/3rd century CE)."I do not need a friend who changes when I change and who nods when I nod; my shadow does that much better."(Plutarch, Quomodo adulator ab amico internoscatur 53b 10)
Ivan P. Yamshchikov, Alexey Tikhonov, Yorgos Pantis, Charlotte Schubert, Jürgen Jost
EMNLP2
2022 HeadlineCause: A Dataset of News Headlines for Detecting Causalities
abstract
Detecting implicit causal relations in texts is a task that requires both common sense and world knowledge. Existing datasets are focused either on commonsense causal reasoning or explicit causal relations. In this work, we present HeadlineCause, a dataset for detecting implicit causal relations between pairs of news headlines. The dataset includes over 5000 headline pairs from English news and over 9000 headline pairs from Russian news labeled through crowdsourcing. The pairs vary from totally unrelated or belonging to the same general topic to the ones including causation and refutation relations. We also present a set of models and experiments that demonstrates the dataset validity, including a multilingual XLM-RoBERTa based model for causality detection and a GPT-2 based model for possible effects prediction.
Ilya Gusev, Alexey Tikhonov
LREC2
2022 EENLP: Cross-lingual Eastern European NLP Index
abstract
Motivated by the sparsity of NLP resources for Eastern European languages, we present a broad index of existing Eastern European language resources (90+ datasets and 45+ models) published as a github repository open for updates from the community. Furthermore, to support the evaluation of commonsense reasoning tasks, we provide hand-crafted cross-lingual datasets for five different semantic tasks (namely news categorization, paraphrase detection, Natural Language Inference (NLI) task, tweet sentiment detection, and news sentiment detection) for some of the Eastern European languages. We perform several experiments with the existing multilingual models on these datasets to define the performance baselines and compare them to the existing results for other languages.
Alexey Tikhonov, Alex Malkhasov, Andrey Manoshin, George-Andrei Dima, Réka Cserháti, Md. Sadek Hossain Asif, Matt Sárdi
LREC1
2022 When Less Is More: Systematic Analysis of Cascade-Based Community Detection
abstract
Information diffusion, spreading of infectious diseases, and spreading of rumors are fundamental processes occurring in real-life networks. In many practical cases, one can observe when nodes become infected, but the underlying network, over which a contagion or information propagates, is hidden. Inferring properties of the underlying network is important since these properties can be used for constraining infections, forecasting, viral marketing, and so on. Moreover, for many applications, it is sufficient to recover only coarse high-level properties of this network rather than all its edges. This article conducts a systematic and extensive analysis of the following problem: Given only the infection times, find communities of highly interconnected nodes. This task significantly differs from the well-studied community detection problem since we do not observe a graph to be clustered. We carry out a thorough comparison between existing and new approaches on several large datasets and cover methodological challenges specific to this problem. One of the main conclusions is that the most stable performance and the most significant improvement on the current state-of-the-art are achieved by our proposed simple heuristic approaches agnostic to a particular graph structure and epidemic model. We also show that some well-known community detection algorithms can be enhanced by including edge weights based on the cascade data.
Liudmila Ostroumova, Alexey Tikhonov, Nelly Litvak
ACM Trans. Knowl. Discov. Data2
2021 Style-transfer and Paraphrase: Looking for a Sensible Semantic Similarity Metric
abstract
The rapid development of such natural language processing tasks as style transfer, paraphrase, and machine translation often calls for the use of semantic similarity metrics. In recent years a lot of methods to measure the semantic similarity of two short texts were developed. This paper provides a comprehensive analysis for more than a dozen of such methods. Using a new dataset of fourteen thousand sentence pairs human-labeled according to their semantic similarity, we demonstrate that none of the metrics widely used in the literature is close enough to human judgment in these tasks. A number of recently proposed metrics provide comparable results, yet Word Mover Distance is shown to be the most reasonable solution to measure semantic similarity in reformulated texts at the moment.
Ivan P. Yamshchikov, Viacheslav Shibaev, Nikolay Khlebnikov, Alexey Tikhonov
AAAI4
2021 Artificial Neural Networks Jamming on the Beat
abstract
This paper addresses the issue of long-scale correlations that is characteristic for symbolic music and is a challenge for modern generative algorithms. It suggests a very simple workaround for this challenge, namely, generation of a drum pattern that could be further used as a foundation for melody generation. The paper presents a large dataset of drum patterns alongside with corresponding melodies. It explores two possible methods for drum pattern generation. Exploring a latent space of drum patterns one could generate new drum patterns with a given music style. Finally, the paper demonstrates that a simple artificial neural network could be trained to generate melodies corresponding with these drum patters used as inputs. Resulting system could be used for end-to-end generation of symbolic music with song-like structure and higher long-scale correlations between the notes.
Alexey Tikhonov, Ivan P. Yamshchikov
COMPLEXIS1
2021 Systematic Analysis of Cluster Similarity Indices: How to Validate Validation Measures
abstract
Many cluster similarity indices are used to evaluate clustering algorithms, and choosing the best one for a particular task remains an open problem. We demonstrate that this problem is crucial: there are many disagreements among the indices, these disagreements do affect which algorithms are preferred in applications, and this can lead to degraded performance in real-world systems. We propose a theoretical framework to tackle this problem: we develop a list of desirable properties and conduct an extensive theoretical analysis to verify which indices satisfy them. This allows for making an informed choice: given a particular application, one can first select properties that are desirable for the task and then identify indices satisfying these. Our work unifies and considerably extends existing attempts at analyzing cluster similarity indices: we introduce new properties, formalize existing ones, and mathematically prove or disprove each property for an extensive list of validation indices. This broader and more rigorous approach leads to recommendations that considerably differ from how validation indices are currently being chosen by practitioners. Some of the most popular indices are even shown to be dominated by previously overlooked ones.
Martijn Gösgens, Alexey Tikhonov, Liudmila Ostroumova
ICML2
2020 Paranoid Transformer: Reading Narrative of Madness as Computational Approach to Creativity
Yana Agafonova, Alexey Tikhonov, Ivan P. Yamshchikov
ICCC2
2020 Drum Beats and Where To Find Them: Sampling Drum Patterns from a Latent Space
Alexey Tikhonov, Ivan P. Yamshchikov
ICCC1
2019 Style Transfer for Texts: Retrain, Report Errors, Compare with Rewrites
abstract
Alexey Tikhonov, Viacheslav Shibaev, Aleksander Nagaev, Aigul Nugmanova, Ivan P. Yamshchikov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alexey Tikhonov, Viacheslav Shibaev, Aleksander Nagaev, Aigul Nugmanova, Ivan P. Yamshchikov
EMNLP/IJCNLP (1)1
2019 Community Detection through Likelihood Optimization: In Search of a Sound Model
abstract
Community detection is one of the most important problems in network analysis. Among many algorithms proposed for this task, methods based on statistical inference are of particular interest: they are mathematically sound and were shown to provide partitions of good quality. Statistical inference methods are based on fitting some random graph model (a.k.a. null model) to the observed network by maximizing the likelihood. The choice of this model is extremely important and is the main focus of the current study. We provide an extensive theoretical and empirical analysis to compare several models: the widely used planted partition model, recently proposed degree-corrected modification of this model, and a new null model having some desirable statistical properties. We also develop and compare two likelihood optimization algorithms suitable for the models under consideration. An extensive empirical analysis on a variety of datasets shows, in particular, that the new model is the best one for describing most of the considered real-world complex networks according to the likelihood of observed graph structures.
Liudmila Ostroumova, Alexey Tikhonov
WWW2
2019 Learning Clusters through Information Diffusion
abstract
When information or infectious diseases spread over a network, in many practical cases, one can observe when nodes adopt information or become infected, but the underlying network is hidden. In this paper, we analyze the problem of finding communities of highly interconnected nodes, given only the infection times of nodes. We propose, analyze, and empirically compare several algorithms for this task. The most stable performance, that improves the current state-of-the-art, is obtained by our proposed heuristic approaches, that are agnostic to a particular graph structure and epidemic model.
Liudmila Ostroumova, Alexey Tikhonov, Nelly Litvak
WWW2
2018 Guess who? Multilingual Approach For The Automated Generation Of Author-Stylized Poetry
abstract
This paper addresses the problem of stylized text generation in a multilingual setup. A version of a language model based on a long short-term memory (LSTM) artificial neural network with extended phonetic and semantic embeddings is used for stylized poetry generation. The quality of the resulting poems generated by the network is estimated through bilingual evaluation understudy (BLEU), a survey and a new cross-entropy based metric that is suggested for the problems of such type. The experiments show that the proposed model consistently outperforms random sample and vanilla-LSTM baselines, humans also tend to associate machine generated texts with the target author.
Alexey Tikhonov, Ivan P. Yamshchikov
SLT1
2014 Crawling Policies Based on Web Page Popularity Prediction
Liudmila Ostroumova, Ivan Bogatyy, Arseniy Chelnokov, Alexey Tikhonov, Gleb Gusev
ECIR4
2014 Parameter-free discovery and recommendation of areas-of-interest
abstract
The task of discovering places of interest is a key step for many location-based recommendation tasks. In this paper we propose a fully unsupervised and parameter-free approach to deal with this problem based on the collection of geotagged photos. While previous papers are mostly devoted to discovering points (POI), we focus on areas of interest (AOI). Recommendation of better matches the traditional tourist goals and allows to robustly incorporate the interests of many users resulting in less subjective recommendations. The typical question that can be answered with the algorithm is formulated as "Where can one spend T minutes/hours walking around to observe as many attractive places as possible?"
Dmitry Laptev, Alexey Tikhonov, Pavel Serdyukov, Gleb Gusev
SIGSPATIAL/GIS2
2013 Studying page life patterns in dynamical web
abstract
With the ever-increasing speed of content turnover on the web, it is particularly important to understand the patterns that pages' popularity follows. This paper focuses on the dynamical part of the web, i.e. pages that have a limited lifespan and experience a short popularity outburst within it. We classify these pages into five patterns based on how quickly they gain popularity and how quickly they lose it. We study the properties of pages that belong to each pattern and determine content topics that contain disproportionately high fractions of particular patterns. These developments are utilized to create an algorithm that approximates with reasonable accuracy the expected popularity pattern of a web page based on its URL and, if available, prior knowledge about its domain's topics.
Alexey Tikhonov, Ivan Bogatyy, Pavel Burangulov, Liudmila Ostroumova, Vitaliy Koshelev, Gleb Gusev
SIGIR1
2010 Analyzing conversations with dynamic graph visualization
abstract
In this paper, we consider the problem of analysis and visualization of online conversations (chat histories, email archives, etc.). We present a dynamic graph drawing algorithm based on modification of multidimensional scaling. The algorithm builds a layout of sequence of graphs and produces a slice view of the evolution of online communications. The method have been applied for visualization of two real-world datasets. We show how to use these visualizations for analyzing and extracting hidden temporal patterns from online conversation data.
Sergey Pupyrev, Alexey Tikhonov
ISDA2