Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Revital Eres

dblp:57/5784 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Runtime systems and virtual machines · 44% Compilers and program optimization · 44% Programming languages and type systems · 13%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
dynamic optimization
0.212013
JIT technology with C/C++: Feedback-directed dynamic recompilation for statically compiled languages · ACM Trans. Archit. Code Optim. 2013
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.212013
JIT technology with C/C++: Feedback-directed dynamic recompilation for statically compiled languages · ACM Trans. Archit. Code Optim. 2013

Methods — techniques the papers use, named apart from their topics

online profiling · 0.2intermediate representation · 0.2
YearPublicationVenuePosition
2024 Data-Prep-Kit: getting your data ready for LLM application development
abstract
Data preparation is the first and a very important step towards any Large Language Model (LLM) development. This paper introduces an easy-to-use, extensible, and scale-flexible open-source data preparation toolkit called Data Prep Kit (DPK). DPK is architected and designed to enable users to scale their data preparation to their needs. With DPK they can prepare data on a local machine or effortlessly scale to run on a cluster with thousands of CPU Cores. DPK comes with a highly scalable, yet extensible set of modules that transform natural language and code data. If the user needs additional transforms, they can be easily developed using extensive DPK support for transform creation. These modules can be used independently or pipelined to perform a series of operations. In this paper, we describe DPK architecture and show its performance from a small scale to a very large number of CPUs. The modules from DPK have been used for the preparation of Granite Models [1] [2]. We believe DPK is a valuable contribution to the AI community to easily prepare data to enhance the performance of their LLM models or to fine-tune models with Retrieval-Augmented Generation (RAG).
Boris Lublinsky, Alexy Roytman, Shivdeep Singh, Constantin Adam, Abdulhamid Adebayo, Sungeun An, Yuan Chi Chang, Xuan-Hong Dang, Nirmit Desai, Michele Dolfi, Hajar Emami-Gohari, Revital Eres, Takuya Goto, Dhiraj Joshi, Yan Koyfman, Mohammad Nassar, Hima Patel, Paramesvaran Selvam, Syed Yousaf Shah, Saptha Surendran, Daiki Tsuzuku, Petros Zerfos, Shahrokh Daijavad
IEEE Big Data13
2018 PM aware storage engine for MongoDB
abstract
With the maturity of Persistant Memories (PM) such as storage class memory technologies, e.g., STT-MRAM, PCM, ReRAM and 3DXpoint, we expect to see practical implementation of data structures, data stores and databases for use-cases such as IoT, mobile, and cloud. We developed a PM-aware storage engine for MongoDB which leverages PM hardware capabilities such as byte addressability and persistency. With our storage engine we see improved latency, less write amplification, less capacity and simpler implementation due to the fact that some code paths become unnecessary compared to past implementations.
Moshe Hershcovitch, Revital Eres, Adam J. McPadden
SYSTOR2
2013 JIT technology with C/C++: Feedback-directed dynamic recompilation for statically compiled languages
abstract
The growing gap between the advanced capabilities of static compilers as reflected in benchmarking results and the actual performance that users experience in real-life scenarios makes client-side dynamic optimization technologies imperative to the domain of static languages. Dynamic optimization of software distributed in the form of a platform-agnostic Intermediate-Representation (IR) has been very successful in the domain of managed languages, greatly improving upon interpreted code, especially when online profiling is used. However, can such feedback-directed IR-based dynamic code generation be viable in the domain of statically compiled, rather than interpreted, languages? We show that fat binaries, which combine the IR together with the statically compiled executable, can provide a practical solution for software vendors, allowing their software to be dynamically optimized without the limitation of binary-level approaches, which lack the high-level IR of the program, and without the warm-up costs associated with the IR-only software distribution approach. We describe and evaluate the fat-binary-based runtime compilation approach using SPECint2006, demonstrating that the overheads it incurs are low enough to be successfully surmounted by dynamic optimization. Building on Java JIT technologies, our results already improve upon common real-world usage scenarios, including very small workloads.
Dorit Nuzman, Revital Eres, Sergei Dyshel, Marcel Zalmanovici, José G. Castaños
ACM Trans. Archit. Code Optim.2
2004 Permuted and Scaled String Matching
Ayelet Butman, Revital Eres, Gad M. Landau
SPIRE2
2004 Scaled and permuted string matching
Ayelet Butman, Revital Eres, Gad M. Landau
Inf. Process. Lett.2
2003 A Combinatorial Approach to Automatic Discovery of Cluster-Patterns
Revital Eres, Gad M. Landau, Laxmi Parida
WABI1