EDBT 2026 Demo / reviewers in the wild / expert
Karim Youssef
dblp:44/9456
· DBLP profile ↗
13ranked-venue papers
11as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Metis: Agentic Knowledge Synthesis for Explainable I/O Performance in HPC SystemsabstractI/O performance explainability in HPC requires contextual characterization across the full software and system stack. Contextual characterization identifies the semantics and runtime role of I/O functions. State-of-the-art contextual characterization is still largely manual, but remains highly valuable for explaining bottlenecks and guiding optimization. However, manual contextual characterization is difficult to scale and hard to reproduce as I/O libraries and cross-layer interactions grow in complexity. We present Metis, a framework for systematic characterization of HPC I/O functions that uses agentic LLMs to integrate heterogeneous evidence sources and quantify agents agreement. Across evaluation, Metis improves held-out-category generalization over an MCP Tool baseline (0.90 vs. 0.35), reduces runtime (27.7 s vs. 84.5 s) while increasing throughput (41.8 vs. 14.28 functions/min), and sustains high verifier throughput under federated scaling (330K–1.18M functions/s). These results demonstrate that Metis is an effective and practical approach for explainable characterization of complex HDF5 behavior, enabling more trustworthy and reproducible HPC I/O analysis. Karim Youssef, Sarah Neuwirth, Neeraj Rajesh, Hariharan Devarajan |
HPDC | 1 |
| 2026 | WADO: A Distributed WORM Storage Service for Asynchronous Data OperationsabstractAI-driven scientific workloads increasingly depend on data-intensive input pipelines, where deep learning frameworks must ingest and transform large datasets from hierarchical HPC storage. Existing system-centric data services improve movement and locality between the parallel file system (PFS), node-local storage, and memory. However, they do not directly optimize how input pipeline operations execute across scopes, stage overlap, and resource-specific parallelism. As scale grows, this gap causes worker stalls, contention, and poor hardware utilization. We present WADO, a distributed write-once-read-many (WORM) object-store runtime for data-centric workloads that closes this gap through three coordinated mechanisms: scope-centric processing, explicit pipeline decomposition, and interference-aware explicit parallelism. WADO dynamically maps operations to execution scopes, overlaps stages such as I/O, communication, and transformations, and applies contention-aware concurrency control to match hardware behavior at runtime. Our evaluation shows three main findings: (1) scope-centric processing preserves throughput under scale, improving mixed-operation throughput by up to 1.65 × ; (2) explicit pipeline decomposition converts serialized wait into overlapped progress, delivering up to 2.16 × higher sustained bandwidth; and (3) interference-aware explicit parallelism improves effective bandwidth by up to 4.4 × by avoiding oversubscription collapse. On Unet3D model training, these mechanisms translate to end-to-end gains, improving data loading performance by 4.1 × compared to baseline PyTorch on Lustre, and 1.51 × compared to DYAD, enabled by deeper pipelining, adaptive parallelism, and near-data transformation offloading. Karim Youssef, Hariharan Devarajan, Nikoli Dryden, Roger A. Pearce |
SSDBM | 1 |
| 2026 | Optimizing Management of Persistent Data Structures in High-Performance AnalyticsabstractLarge-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application's control, and the significantly high storage footprint of such snapshots. To address these limitations, we presentPrivateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. We integratedPrivateerintoMetall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning.Privateeroptimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system.Privateeralso optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression. Karim Youssef, Keita Iwabuchi, Maya B. Gokhale, Wu-chun Feng, Roger A. Pearce |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | Scalable and Maintainable Distributed Sequence Alignment Using SparkabstractThe exponential growth of genomic data presents a challenge to bioinformatics research. NCBI BLAST, a popular pairwise sequence alignment tool, does not scale with the hundreds of gigabytes (GB) of sequenced data. Therefore, mpiBLAST was widely adopted and scaled up to 65,536 processors. However, mpiBLAST is tightly coupled with an obsolete NCBI BLAST version, creating a challenge to upgrading mpiBLAST with the ever-changing NCBI BLAST code. Recent parallel BLAST implementations, like SparkBLAST, use parallelism wrappers separate from NCBI BLAST to overcome this issue. However, query partitioning, a parallel method that duplicates the genome database on each compute node, makes SparkBLAST scale poorly with databases larger than a single node's memory. Thus, no parallel BLAST utility simultaneously addresses performance, scalability, and software maintainability. To fill this gap, we introduce SparkLeBLAST, a parallel BLAST tool that uses the Spark framework and efficient data partitioning to combine mpiBLAST's performance and scalability with SparkBLAST's simplicity and maintainability. SparkLeBLAST democratizes scalable genomic analysis for domain scientists without extensive distributed computing experience. SparkLeBLAST runs up to 6.68× faster than SparkBLAST. SparkLeBLAST also accelerates taxonomic assignment of COVID-19 genomic diversity analysis by 20.9× as it speeds up the BLAST search component by 88.6× using 128 compute nodes. Karim Youssef, Yusuf Elnady, Eli Tilevich, Wu-chun Feng |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | Metall: A persistent memory allocator for data-centric analytics
Keita Iwabuchi, Karim Youssef, Kaushik Velusamy, Maya B. Gokhale, Roger A. Pearce |
Parallel Comput. | 2 |
| 2022 | Enabling Scalable and Extensible Memory-Mapped Datastores in UserspaceabstractExascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we presentUMap, a scalable and extensible userspace service for memory-mapping datastores. Through decoupled queue management, concurrency aware adaptation, and dynamic load balancing,UMapenables application performance to scale even at high concurrency. We evaluateUMapin data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show thatUMapas a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9${\times}$. We performed two case studies to demonstrateUMap's extensibility. First, a new datastore residing in remote memory is incorporated intoUMapas an application-specific plugin. Second, we present a persistent memory allocatorMetallbuilt atopUMapfor unified storage/memory. Ivy Bo Peng, Maya B. Gokhale, Karim Youssef, Keita Iwabuchi, Roger A. Pearce |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | SparkLeBLAST: Scalable Parallelization of BLAST Sequence Alignment Using SparkabstractThe exponential growth of genomic data presents challenges in analyzing and computing on such biological data at scale. While NCBI’s BLAST is a widely used pairwise sequence alignment tool, it does not scale to large datasets that are hundreds of gigabytes (GB) in size. To address this scalability problem, mpiBLAST emerged and became widely used, enabling scaling to 65,536 processes. However, mpiBLAST suffers from being tightly coupled with a specific implementation of BLAST, rendering it difficult to upgrade with the ever-evolving NCBI BLAST code. To address this shortcoming, recent parallel BLAST tools, such as SparkBLAST, consist of wrappers that are decoupled from the BLAST code but suffer from poor scalability with large sequence databases. Thus, there does not exist any parallel BLAST tool that can simultaneously address the issues of performance, scalability, programmability, and upgradability. To address this void, we propose SparkLeBLAST, a parallel BLAST tool that leverages our performance modeling and the Spark framework to deliver the performance and scalability of mpiBLAST and the ease of programming and upgradability of SparkBLAST, respectively. Ultimately, SparkLeBLAST delivers a 10x speedup relative to the state-of-the-art SparkBLAST and nearly a 2x speedup relative to the latest version of mpiBLAST. Karim Youssef, Wu-chun Feng |
CCGRID | 1 |
| 2015 | Identification and Localization of One or Two Concurrent Speakers in a Binaural Robotic ContextabstractThis paper presents a method of identification and azimuth estimation for one or two concurrent speakers in simultaneous utterances. This method is applicable to human-machine interaction and robot audition. Identification and localization have been rarely mutually addressed and related works rely on time-frequency exploitation strategies to extract and treat each source's contribution to the received signal. The presented method relies on a training made with one speaker at a time, but it can exploit a speech segment to identify and localize two speakers. A cochlear filtering-based binaural front-end allows to extract equivalent rectangular bandwidth frequency cepstral coefficients (ERBFCC) and interaural level difference (ILD) features. Artificial neural networks (ANNs) exploit ERBFCCs to provide identity information, and a histogram-based exploitation of ILDs provides azimuth angle information. The method was evaluated in contexts including overlapping segments in the presence of noises and sound reflections and its efficiency was demonstrated. Even with fully overlapping utterances, we reached an 83% identification rate of both speakers, an 82% estimation accuracy of both azimuths and an 68% correct mutual identity and azimuth estimation rate. At least one speaker was correctly identified and localized in more than 99% of the tests for utterances lasting near 5s. Karim Youssef, Katsutoshi Itoyama, Kazuyoshi Yoshii |
SMC | 1 |
| 2013 | A learning-based approach to robust binaural sound localizationabstractSound source localization is an important feature designed and implemented on robots and intelligent systems. Like other artificial audition tasks, it is constrained to multiple problems, notably sound reflections and noises. This paper presents a sound source azimuth estimation approach in reverberant environments. It exploits binaural signals in a humanoid robotic context. Interaural Time and Level Differences (ITD and ILD) are extracted on multiple frequency bands and combined with a neural network-based learning scheme. A cue filtering process is used to reduce the reverberations effects. The system has been evaluated with simulation and real data, in multiple aspects covering realistic robot operating conditions, and was proven satisfying and effective as will be shown and discussed in the paper. Karim Youssef, Sylvain Argentieri, Jean-Luc Zarader |
IROS | 1 |
| 2012 | A binaural sound source localization method using auditive cues and visionabstractA fundamental task for a robotic audition system is sound source localization. This paper addresses the localization problem in a robotic humanoid context, providing a novel learning algorithm that uses binaural cues to determine the sound source's position. Sound signals are extracted from a humanoid robot's ears. Binaural cues are then computed to provide inputs for a neural network. The neural network uses pixel coordinates of a sound source in a camera image as outputs. This learning approach provides good localization performances as it reaches very small errors for azimuth and elevation angles estimates. Karim Youssef, Sylvain Argentieri, Jean-Luc Zarader |
ICASSP | 1 |
| 2012 | Towards a systematic study of binaural cuesabstractSound source localization is a need for robotic systems interacting with acoustically-active environments. In this domain, numerous binaural localization studies have been conducted within the last few decades. This paper provides an overview of a number of binaural localization cue extraction techniques. These are carefully addressed and applied on a simulated binaural database. Cues are evaluated in azimuth estimation and their discriminatory effectiveness is studied as a function of the reverberation time with statistical data analysis techniques. Results show that big differences exist between the discriminatory abilities of multiple types of cue extraction methods. Thus a careful cue selection must be performed before establishing a sound localization system. Karim Youssef, Sylvain Argentieri, Jean-Luc Zarader |
IROS | 1 |
| 2011 | Approaches for Automatic Speaker Recognition in a Binaural Humanoid Context
Karim Youssef, Bastien Breteau, Sylvain Argentieri, Jean-Luc Zarader |
ESANN | 1 |
| 2010 | Binaural speaker recognition for humanoid robotsabstractIn this paper, an original study of a binaural speaker identification system is presented. The state of the art shows that, contrarily to monaural and multi-microphone approaches, binaural systems are not so much studied in the specific task of automatic speaker recognition. Indeed, these systems are mostly used for speech recognition, or speaker localization. This study will focus on the benefits of the binaural context in comparison with monaural techniques. It demonstrates the interest of the binaural systems typically used in humanoid robotics. The system is first tested with monaural signals, and then with a binaural sensor, in many signal to noise ratios, speech durations and speaker directions. Up to 11 percent of improvement in recognition ratios of 23 ms frames can be obtained. The used database is a set of audio tracks recorded for 10 speakers, and filtered by HRTFs to obtain binaural signals in the directions of interest, for the binaural training and testing steps. This way, we study the sensitivity of the system to the speaker's location in an environment where a maximum of 10 speakers is present. Karim Youssef, Sylvain Argentieri, Jean-Luc Zarader |
ICARCV | 1 |