EDBT 2026 Demo / reviewers in the wild / expert
Jophin John
dblp:263/5593
· DBLP profile ↗
8ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-6670-5343ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Temporal Robustness in Hate Speech Detection: Updating German Classifiers with Advanced AI InfrastructuresabstractOver the past two decades, hate speech on social media has surged, causing significant harm and threatening democracies. Initially, research focused on English hate speech, but recent years have seen the development of non-English datasets and corresponding classifiers. However, these datasets are often small in size, with a specific topic of interest, and temporally misaligned, causing classifiers to quickly become outdated. Consequently, there is a growing need to understand how to efficiently update these outdated classifiers. Motivated by the AI4Dignity project, this study addresses the issue of diachronicity in hate speech detection, specifically focusing on German. The study conducts a series of experiments to explore multiple fine-tuning strategies for updating and improving outdated hate speech classifiers. Additionally, it examines the impact of different AI infrastructures (state-of-the-art GPU systems, Intel’s Gaudi 2 and a Cerebras CS-2 system) on training times. The findings provide insights into the temporal robustness of different classifiers and respective updating approaches, as well as which AI infrastructure works best with each update approach. These insights may not only help efficiently maintain hate speech detection systems for social media but can also be useful for other areas, such as continuously supporting the reliable detection of hate speech in online gaming environments. Michael Peter Hoffmann, Jan Fillies, Jophin John, Antonis Maronikolakis, Ajay Navilarekal, Sahana Udupa, Nicolay Hammer |
ECAI | 3 |
| 2025 | Towards a Holistic Evaluation of Novel AI Accelerators: Theory, Practice and EthicsabstractThe rapid advancement of artificial intelligence has driven significant diversification of specialized hardware accelerators, including TPUs, GPUs, FPGAs, ASICs, and emerging neuromorphic architectures. Traditionally, evaluation methods have centered narrowly on performance metrics like FLOPS, throughput, latency, and computational efficiency. However, such metrics inadequately capture the complex nature of real-world AI accelerator deployment, overlooking integration complexity, researcher adoption barriers, security, vendor relationships, and evolving hardware-software dynamics. This study proposes shifting from a pure benchmark-centric approach towards a holistic evaluative framework, introducing the Evaluation Circle that addresses six interconnected dimensions: installation and maintenance, experimental performance, community response, security assessment, co-design assessment, and energy consumption. We validated our approach by applying the Evaluation Circle to a Cerebras CS-2 system at a major European supercomputing center. Our analysis revealed insights invisible to conventional evaluation, including the CS-2's managementintensive nature and a paradox within the local AI community: strong interest coupled with skepticism about new platforms. The framework illuminated how collaborative vendor-datacenter projects addressed these issues. Particularly significant was the transformative potential of structured vendor co-design, creating a continuous improvement cycle benefiting both the institution and the broader AI ecosystem. This study contributes to AI infrastructure literature by reframing accelerators as components within sociotechnical systems, developing a comprehensive evaluation methodology, and emphasizing ethics, collaboration, and conversational learning for institutions integrating novel AI hardware. Michael Peter Hoffmann, Jophin John, Hoi-Fong Mak, Nicolay Hammer |
ICTAI | 2 |
| 2024 | Leveraging Resource-Aware Application-Level Checkpointing and RDMA for Fault Tolerance and Data Distribution in Malleable MPI ApplicationsabstractDynamic resource management and application malleability present numerous opportunities in High-Performance Computing (HPC), enhancing both system-level services and application performance. Recent trends in malleability research, encompassing both application and system dynamism, are building a new era in HPC. Dynamic applications, particularly those leveraging malleable resources, require adaptive checkpointing systems to enhance performance and resource use. Applications can significantly benefit from checkpointing systems becoming dynamic, especially in handling data redistribution during resource changes. Consequently, checkpointing services should also become malleable (or adaptive). Therefore, we propose iCheck, an adaptive application-level checkpoint management system that caters to malleable MPI applications. iCheck aids these applications by dynamically reconfiguring checkpointing resources and offering robust checkpointing and data redistribution services. By leveraging Remote Direct Memory Access (RDMA) to support malleable applications, iCheck facilitates faster data transfers (up to 40 times improvement over the PFS-based solution), ensuring efficient performance. The system can dynamically adjust checkpointing processes based on metrics such as available memory, checkpoint frequency, and number of processes, maintaining or improving checkpoint performance amid resource changes. This adaptive approach supports fault tolerance as well as simplifies the development of malleable applications by effectively managing resource redistribution during resource changes. Jophin John, Michael Gerndt |
HPCC | 1 |
| 2024 | Exploring the Suitability of the Cerebras Wafer Scale Engine for the Fast Prototyping of a Multilingual Hate Speech Detection SystemabstractThe era of digital communication has brought about a concerning rise of online hate speech. In response, researchers have focused on developing automated systems to detect and monitor such harmful content. While much attention has been given to monolingual detection systems, recent years have seen the emergence of new approaches for multilingual hate speech detection. However, there remains a limited understanding of how the underlying computational infrastructure impacts the training and development times of such systems. This study presents an innovative experimental design aimed at investigating the relationship between accelerator infrastructure and multilingual hate speech detection. It begins by constructing a prototype system of different classification algorithms for detecting hate speech in English, German, Italian and Spanish text-based social media content. The study then evaluates the fine-tuning times of these classifiers using both conventional GPU-based accelerators and cutting-edge AI-accelerator hardware, such as the Cerebras CS-2 system. The latter claims to speed up the development and fine-tuning of large language models significantly. Furthermore, the study compares the fine-tuning times of the same classifiers on two separate AI Accelerator machines. The study shows that the Cerebras AI Accelerator quickens training times by factor 4 compared to traditional setups, with little variation across high-performance computing infrastructures. However, the technology is in its early stages, with drawbacks including substantial upstart and compilation times, a limited number of models portable to the CS-2, and a need for refinement to improve accessibility for non-experts. Nonetheless, the Cerebras CS-2 technology heralds a new era for the rapid development of effective hate speech models. Michael Peter Hoffmann, Jophin John, Nicolay Hammer |
ICTAI | 2 |
| 2023 | Exploring the Use of WebAssembly in HPCabstractContainerization approaches based on namespaces offered by the Linux kernel have seen an increasing popularity in the HPC community both as a means to isolate applications and as a format to package and distribute them. However, their adoption and usage in HPC systems faces several challenges. These include difficulties in unprivileged running and building of scientific application container images directly on HPC resources, increasing heterogeneity of HPC architectures, and access to specialized networking libraries available only on HPC systems. These challenges of container-based HPC application development closely align with the several advantages that a new universal intermediate binary format called WebAssembly (Wasm) has to offer. These include a lightweight userspace isolation mechanism and portability across operating systems and processor architectures. In this paper, we explore the usage of Wasm as a distribution format for MPI-based HPC applications. To this end, we present MPIWasm, a novel Wasm embedder for MPI-based HPC applications that enables high-performance execution of Wasm code, has low-overhead for MPI calls, and supports high-performance networking interconnects present on HPC systems. We evaluate the performance and overhead of MPIWasm on a production HPC system and AWS Graviton2 nodes using standardized HPC benchmarks. Results from our experiments demonstrate that MPIWasm delivers competitive native application performance across all scenarios. Moreover, we observe that Wasm binaries are 139.5x smaller on average as compared to the statically-linked binaries for the different standardized benchmarks. Mohak Chadha, Nils Krueger, Jophin John, Anshul Jindal, Michael Gerndt, Shajulin Benedict |
PPoPP | 3 |
| 2022 | iCheck: Leveraging RDMA and Malleability for Application-Level Checkpointing in HPC SystemsabstractThe estimate that the mean time between failures will be in minutes in exascale supercomputers should be alarming for application developers. The inherent system’s complexity, millions of components, and susceptibility to failures make checkpointing more relevant than ever. Since most high performance scientific applications contain an in-house checkpoint restart mechanism, their performance can be impacted by the contention of parallel file system resources. A shift in checkpointing strategies is needed to thwart this behavior. With iCheck, we present a novel checkpointing framework that supports malleable multilevel application-level checkpointing. We employ an RDMA enabled configurable multi-agent-based checkpoint transfer mechanism where minimal application resources are utilized for checkpointing. The high-level API of iCheck facilitates easy integration and malleability. We have added the iCheck library into the Is1 mardyn application providing performance improvement up to five thousand times over the in-house checkpointing mechanism. LULESH, Jacobi 2D heat simulation, and a synthetic application were also used for extensive analysis. Jophin John, Isaac David Núñez Araya, Michael Gerndt |
ICPADS | 1 |
| 2020 | Toward an End-to-End Auto-tuning Framework in HPC PowerStackabstractEfficiently utilizing procured power and optimizing performance of scientific applications under power and energy constraints are challenging. The HPC PowerStack defines a software stack to manage power and energy of high-performance computing systems and standardizes the interfaces between different components of the stack. This survey paper presents the findings of a working group focused on the end-to-end tuning of the PowerStack. First, we provide a background on the PowerStack layer-specific tuning efforts in terms of their high-level objectives, the constraints and optimization goals, layer-specific telemetry, and control parameters, and we list the existing software solutions that address those challenges. Second, we propose the PowerStack end-to-end auto-tuning framework, identify the opportunities in co-tuning different layers in the PowerStack, and present specific use cases and solutions. Third, we discuss the research opportunities and challenges for collective auto-tuning of two or more management layers (or domains) in the PowerStack. This paper takes the first steps in identifying and aggregating the important R&D challenges in streamlining the optimization efforts across the layers of the PowerStack. Xingfu Wu, Aniruddha Marathe, Siddhartha Jana, Ondrej Vysocky, Jophin John, Andrea Bartolini, Lubomir Riha, Michael Gerndt, Valerie Taylor 0001, Sridutt Bhalachandra |
CLUSTER | 5 |
| 2020 | Extending SLURM for Dynamic Resource-Aware Adaptive Batch SchedulingabstractWith the growing constraints on power budget and increasing hardware failure rates, the operation of future exascale systems faces several challenges. Towards this, resource awareness and adaptivity by enabling malleable jobs has been actively researched in the HPC community. Malleable jobs can change their computing resources at runtime and can significantly improve HPC system performance. However, due to the rigid nature of popular parallel programming paradigms such as MPI and lack of support for dynamic resource management in batch systems, malleable jobs have been largely unrealized. In this paper, we extend the SLURM batch system to support the execution and batch scheduling of malleable jobs. The malleable applications are written using a new adaptive parallel paradigm called Invasive MPI which extends the MPI standard to support resource-adaptivity at runtime. We propose two malleable job scheduling strategies to support performance-aware and power-aware dynamic reconfiguration decisions at runtime. We implement the strategies in SLURM and evaluate them on a production HPC system. Results for our performance-aware scheduling strategy show improvements in makespan, average system utilization, average response, and waiting times as compared to other scheduling strategies. Moreover, we demonstrate dynamic power corridor management using our power-aware strategy. Mohak Chadha, Jophin John, Michael Gerndt |
HiPC | 2 |