VLDB 2026 Research / reviewers in the wild / expert
Ronay Ak
dblp:123/3573
· DBLP profile ↗
11ranked-venue papers in the field
2as first author
7since 2021 · last 2025
0000-0003-2768-6535ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5 (1 first)Big Data, Cloud & Distributed Data Systems · 4 (1 first)Data Mining & Knowledge Discovery · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Text Embedding Models with Positive-aware Hard-negative Mining
Gabriel de Souza Pereira Moreira, Radek Osmulski, Ronay Ak, Benedikt Schifferer, Even Oldridge |
CIKM | 4 |
| 2025 | Boost the Performance of Tabular Data Models with GPU Accelerated Feature EngineeringabstractFeature engineering remains a crucial technique for improving the performance of models trained on tabular data. Unlike computer vision and natural language processing, where deep learning models automatically extract hierarchical features from raw data, the most accurate tabular models, such as gradient boosted decision trees, still benefit significantly from manually crafted features. This is demonstrated in Team NVIDIA's many first-place data science competition victories. Chris Deotte, Ronay Ak |
KDD (2) | 2 |
| 2022 | Reducing the Friction for Building Recommender Systems with MerlinabstractRecommender Systems (RecSys) are the engine of the modern internet and the catalyst for human decisions. The goal of a recommender system is to generate relevant recommendations for users from a collection of items or services that might interest them. Building a recommendation system is challenging because it requires multiple stages (item retrieval, filtering, ranking, ordering) to work together seamlessly and efficiently during training and inference. The biggest challenges faced by new practitioners are the lack of understanding around what RecSys look like in the real world and the difficulty in transitioning from the simple Matrix Factorization (MF) to more complex deep learning architectures with multiple input features, neural components and prediction heads. Sara Rabhi, Ronay Ak, Marc Romeijn, Gabriel de Souza Pereira Moreira, Benedikt Schifferer |
KDD | 2 |
| 2022 | Training and Deploying Multi-Stage Recommender SystemsabstractIndustrial recommender systems are made up of complex pipelines requiring multiple steps including feature engineering and preprocessing, a retrieval model for candidate generation, filtering, a feature store query, a ranking model for scoring, and an ordering stage. These pipelines need to be carefully deployed as a set, requiring coordination during their development and deployment. Data scientists, ML engineers, and researchers might focus on different stages of recommender systems, however they share a common desire to reduce the time and effort searching for and combining boilerplate code coming from different sources or writing custom code from scratch to create their own RecSys pipelines. Ronay Ak, Benedikt Schifferer, Sara Rabhi, Gabriel de Souza Pereira Moreira |
RecSys | 1 |
| 2022 | Building and Deploying a Multi-Stage Recommender System with MerlinabstractNewcomers to recommender systems often face challenges related to their lack of understanding of how these systems operate in real life. In most online content related to this topic, the focus is on models and algorithms that score items based on the user’s preferences. However, the recommender model alone does not comprise everything needed for serving optimized recommender systems that meet the company’s business objectives. An industry-standard recommender system involves a number of steps, including data preprocessing, defining and training recommender models, as well as filtering and business logic for serving. In this work, we propose the four-stage recommender system, an industry-wide design pattern we have identified for production recommender systems. The four-stage pipeline includes an item retrieval step that prepares a small subset of relevant items for scoring. The filtering stage then cleans up the subset of items based on business logic such as removing out-of-stock or previously seen items. As for the ranking component, it uses a recommender model to score each item in the presented list based on the preferences of the user. In the final step, the scores are re-ordered to provide a final recommendation list aligned with other business needs or constraints such as diversity. In particular, the presented demo demonstrates how easy it is to build and deploy a four-stage recommender system pipeline using the NVIDIA Merlin open-source framework. Karl Higley, Even Oldridge, Ronay Ak, Sara Rabhi, Gabriel de Souza Pereira Moreira |
RecSys | 3 |
| 2021 | End-to-End Session-Based Recommendation on GPUabstractIn recent years, several deep learning-based algorithms have been proposed for recommendation systems while its adoption in industry deployments have been steeply growing. In particular, NLP-inspired approaches have been successfully adapted for sequential and session-based recommendation problems, which are important for many domains like e-commerce, news and streaming media. In this regard, this hands-on tutorial will offer to the participants: (I) an introduction on the main concepts and algorithms for session-based recommendation, (II) how to build, train and evaluate a session-based recommendation model based on RNN and Transformer architectures, and (III) how to speed up with GPUs the entire RecSys pipeline which encompasses feature engineering, preprocessing, training, evaluation and inference using NVIDIA Merlin - an open source ecosystem for large-scale deep learning recommender systems. Gabriel de Souza Pereira Moreira, Sara Rabhi, Ronay Ak, Benedikt Schifferer |
RecSys | 3 |
| 2021 | Transformers4Rec: Bridging the Gap between NLP and Sequential / Session-Based RecommendationabstractMuch of the recent progress in sequential and session-based recommendation has been driven by improvements in model architecture and pretraining techniques originating in the field of Natural Language Processing. Transformer architectures in particular have facilitated building higher-capacity models and provided data augmentation and training techniques which demonstrably improve the effectiveness of sequential recommendation. But with a thousandfold more research going on in NLP, the application of transformers for recommendation understandably lags behind. To remedy this we introduce Transformers4Rec, an open-source library built upon HuggingFace’s Transformers library with a similar goal of opening up the advances of NLP based Transformers to the recommender system community and making these advancements immediately accessible for the tasks of sequential and session-based recommendation. Like its core dependency, Transformers4Rec is designed to be extensible by researchers, simple for practitioners, and fast and robust in industrial deployments. Gabriel de Souza Pereira Moreira, Sara Rabhi, Jeongmin Lee 0001, Ronay Ak, Even Oldridge |
RecSys | 4 |
| 2017 | Automatic localization of casting defects with convolutional neural networksabstractAutomatic localization of defects in metal castings is a challenging task, owing to the rare occurrence and variation in appearance of defects. Convolutional neural networks (CNN) have recently shown outstanding performance in both image classification and localization tasks. We examine how several different CNN architectures can be used to localize casting defects in X-ray images. We take advantage of transfer learning to allow state-of-the-art CNN localization models to be trained on a relatively small dataset. In an alternative approach, we train a defect classification model on a series of defect images and then use a sliding classifier method to develop a simple localization model. We compare the localization accuracy and computational performance of each technique. We show promising results for defect localization on the GRIMA database of X-ray images (GDXray) dataset and establish a benchmark for future studies on this dataset. Max Ferguson, Ronay Ak, Y. Tina Lee, Kincho H. Law |
IEEE BigData | 2 |
| 2015 | Data analytics and uncertainty quantification for energy prediction in manufacturingabstractMany industries are applying various methods for optimizing energy use across the manufacturing life cycle. These methods are either physics-based or data-driven. Manufacturing systems generate a vast amount of data from operations and in simulations. Advances in data collection systems and data analytics (DA) tools have enabled the development of predictive analytics for energy prediction. Many of these prediction methods do not account for the uncertainty quantification-UQ (both in data and model). This work addresses the issue of uncertainty in predictive analytics. This work focuses on metal cutting processes and presents a Neural Networks (NNs) model to predict the required energy consumption during the manufacturing of a part on a milling machine. The model accounts for the uncertainty associated with both the manufacturing processes parameters, and assumptions in building the prediction model. To achieve this, prediction intervals are estimated instead of point predictions. In order to increase the ability to generalize over new datasets, an ensemble model of neural networks (NNs) is used, and the k-nearest-neighbors (k-nn) approach is applied to identify similar patterns between training and test datasets to increase the accuracy of the results by using local information from the closest patterns of the training sets. Case study results demonstrate consistency and high prediction precision as compared to the individual NNs of the ensembles. Moreover, it is shown that with advanced data collection and processing techniques, one can construct a prediction model to predict the energy consumption of a machine tool for machining a part with multiple operations and process parameters. Ronay Ak, Raunak Bhinge |
IEEE BigData | 1 |
| 2015 | Analysis and optimization in smart manufacturing based on a reusable knowledge base for process performance modelsabstractIn this paper, we propose an architectural design and software framework for fast development of descriptive, diagnostic, predictive, and prescriptive analytics solutions for dynamic production processes. The proposed architecture and framework will support the storage of modular, extensible, and reusable Knowledge Base (KB) of process performance models. The approach requires the development of automatic methods that can translate the high-level models in the reusable KB into low-level specialized models required by a variety of underlying analysis tools, including data manipulation, optimization, statistical learning, estimation, and simulation. We also propose an organization and key structure for the reusable KB, composed of atomic and composite process performance models and domain-specific dashboards. Furthermore, we illustrate the use of the proposed architecture and framework by performing diagnostic tasks on a composite performance model. Alexander Brodsky 0001, Guodong Shao, Mohan Krishnamoorthy, Anantha Narayanan, Daniel A. Menascé, Ronay Ak |
IEEE BigData | 6 |
| 2015 | A neural network meta-model and its application for manufacturingabstractManufacturing generates a vast amount of data both from operations and simulation. Extracting appropriate information from this data can provide insights to increase a manufacturer's competitive advantage through improved sustainability, productivity, and flexibility of their operations. Manufacturers, as well as other industries, have successfully applied a promising statistical learning technique, called neural networks (NNs), to extract meaningful information from large data sets, so called big data. However, the application of NN to manufacturing problems remains limited because it involves the specialized skills of a data scientist. This paper introduces an approach to automate the application of analytical models to manufacturing problems. We present an NN meta-model (MM), which defines a set of concepts, rules, and constraints to represent NNs. An NN model can be automatically generated and manipulated based on the specifications of the NN MM. In addition, we present an algorithm to generate a predictive model from an NN and available data. The predictive model is represented in either Predictive Model Markup Language (PMML) or Portable Format for Analytics (PFA). Then we illustrate the approach in the context of a specific manufacturing system. Finally, we identify future steps planned towards later implementation of the proposed approach. David Lechevalier, Steven Hudak, Ronay Ak, Y. Tina Lee, Sebti Foufou |
IEEE BigData | 3 |