Slava Jankin Mikhaylov

dblp:169/9715 · also Slava J. Mikhaylov, Slava Jankin, Slava Mikhaylov · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2024
0000-0001-6915-177XORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2024 From Searches to Sneezes: Evaluating Digital Indicators for Allergic Rhinitis Surveillance in the United Kingdom1
abstract
Background: Allergic rhinitis (AR) affects millions globally. Traditional surveillance methods often lag behind real-time disease activity, hindering timely interventions. Digital epidemiology offers potential for more responsive AR monitoring. Objective: To assess the potential of various digital indicators for AR surveillance in the United Kingdom. Methods: We analyzed weekly data from January 2016 to January 2024, including Google Trends (GT) data, Twitter Frequency (TF), self-reported medication use from the MASK-Air app, and clinical AR incidence data. We employed Spearman correlations, Granger causality tests, SARIMAX, linear regression, and Random Forest regression. Results: Google Trends data showed the strongest correlation with AR incidence (Spearman correlation: 0.73) and the highest Granger causality score (10.33). Regression models using GT data alone performed almost as well as combined models. Twitter data provided slight improvements in test set performance. Self-reported medication use data showed weak correlations with AR incidence and did not significantly improve predictive models. Conclusions: Our findings demonstrate the strong potential of Google Trends data for AR surveillance in the UK. While other digital indicators showed limited predictive power at the national level, social media data may offer value in addressing geographical specificity. Integrating digital surveillance methods, particularly GT data, into public health strategies could enhance AR management and improve outcomes for affected individuals.
Krishnamoorthy Manohara, Slava Jankin Mikhaylov, Jean Bousquet, Paulina Garcia Corral, Jorge Roa, Hannah Béchara, Bernardo Sousa-Pinto
IEEE Big Data2
2022 Applying NLP Techniques to Classify Businesses by their International Standard Industrial Classification (ISIC) Code
abstract
The application of machine learning has played an important role in several aspects of text classification across domains, and has brought with it great changes to the current state of the art. In this paper, we propose a novel application of NLP techniques to classify entities by their International Standard Industrial Classification (ISIC) code based on descriptions provided by the business owners themselves and the names of said businesses. Faced with the issues of irregularity and a small amount of noisy training data, we employ several different NLP models and data enhancement strategies. We identify DistilBERT, under normalized loss, as the best model for our task, with a 77.9% average accuracy on a 56 label multiclass classification task.
Hannah Béchara, Shuzhou Yuan, Slava Jankin Mikhaylov
IEEE Big Data4