Anton Vakhrushev

dblp:301/8308 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2024
0009-0007-1797-8688ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Kernel, tree and ensemble methods · 46% Probabilistic and Bayesian machine learning · 36% Learning paradigms · 14%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Kernel, tree and ensemble methods
gradient boosting
1.322024
Uplift Modelling via Gradient Boosting · KDD 2024
SketchBoost: Fast Gradient Boosted Decision Tree for Multioutput Problems · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.812024
Uplift Modelling via Gradient Boosting · KDD 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › heterogeneous treatment effect estimation
uplift modeling
0.812024
Uplift Modelling via Gradient Boosting · KDD 2024
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosted decision trees
0.612022
SketchBoost: Fast Gradient Boosted Decision Tree for Multioutput Problems · NeurIPS 2022
Machine learning › Learning paradigms
multi-output learning
0.612022
SketchBoost: Fast Gradient Boosted Decision Tree for Multioutput Problems · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

gradient boosting · 1.3GPU acceleration · 0.6
YearPublicationVenuePosition
2024 Uplift Modelling via Gradient Boosting
abstract
The Gradient Boosting machine learning ensemble algorithm, well-known for its proficiency and superior performance in intricate machine learning tasks, has encountered limited success in the realm of uplift modeling. Uplift modeling is a challenging task that necessitates a known target for the precise computation of the training gradient. The prevailing two-model strategies, which separately model treatment and control outcomes, are encumbered with limitations as they fail to directly tackle the uplift problem.
Bulat Ibragimov, Anton Vakhrushev
KDD2
2024 Cross-Domain Latent Factors Sharing via Implicit Matrix Factorization
abstract
Data sparsity has been one of the long-standing problems for recommender systems. One of the solutions to mitigate this issue is to exploit knowledge available in other source domains. However, many cross-domain recommender systems introduce a complex architecture that makes them less scalable in practice. On the other hand, matrix factorization methods are still considered to be strong baselines for single-domain recommendations. In this paper, we introduce the CDIMF, a model that extends the standard implicit matrix factorization with ALS to cross-domain scenarios. We apply the Alternating Direction Method of Multipliers to learn shared latent factors for overlapped users while factorizing the interaction matrix. In a dual-domain setting, experiments on industrial datasets demonstrate a competing performance of CDIMF for both cold-start and warm-start. The proposed model can outperform most other recent cross-domain and single-domain models. We also provide the code to reproduce experiments on GitHub.
Abdulaziz Samra, Evgeny Frolov, Alexey Vasilev, Alexander Grigorevskiy, Anton Vakhrushev
RecSys5
2022 SketchBoost: Fast Gradient Boosted Decision Tree for Multioutput Problems
abstract
Gradient Boosted Decision Tree (GBDT) is a widely-used machine learning algorithm that has been shown to achieve state-of-the-art results on many standard data science problems. We are interested in its application to multioutput problems when the output is highly multidimensional. Although there are highly effective GBDT implementations, their scalability to such problems is still unsatisfactory. In this paper, we propose novel methods aiming to accelerate the training process of GBDT in the multioutput scenario. The idea behind these methods lies in the approximate computation of a scoring function used to find the best split of decision trees. These methods are implemented in SketchBoost, which itself is integrated into our easily customizable Python-based GPU implementation of GBDT called Py-Boost. Our numerical study demonstrates that SketchBoost speeds up the training process of GBDT by up to over 40 times while achieving comparable or even better performance.
Leonid Iosipoi, Anton Vakhrushev
NeurIPS2