Ryan Marinelli

dblp:336/2035 · DBLP profile ↗
← Back
5ranked-venue papers in the field
4as first author
5since 2021 · last 2026
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 4 (3 first)Information Retrieval & Web Search · 1 (1 first)
YearPublicationVenuePosition
2026 Adaptive Filtering for Large Language Models
Ryan Marinelli
NLDB1
2025 Position Paper: An Argument for Applications of Reinforcement in Image Segmentation to Mitigate Adversarial Attacks
Ryan Marinelli, Markus Lammle, Robert Andrew Chetwyn
IEEE Big Data1
2024 Causal Tracing to Identify Hacking Knowledge in Large Language Models
abstract
This research seeks to identify where hacking knowledge is stored in models. SQL Injection is used as a base case considering the risks involved with production systems. Through using causal tracing techniques, it is determined that SQL Injection knowledge capabilities are most prevalent in the last layers of the network. It is also established that there may be some shared representation of SQL Injection knowledge across models through training a KNN classifier. When training the classifier on GPT-Neo’s activations and attempting to classify on GPT-2, it was found to be 52% accurate.
Ryan Marinelli
IEEE Big Data1
2024 Optimizing Deployment of Homomorphic Encryption and SQL using Reinforcement Learning
abstract
This research explores optimization strategies for the deployment of SQL and homomorphic encryption using reinforcement learning. Marginal gains are established and suggests greater inquiry into optimization techniques as a means of introducing more secure compute to production environments.
Ryan Marinelli, Åvald Åslaugson Sommervoll, Laszlo Erdodi
IEEE Big Data1
2024 Dynamic Intelligence Assessment: Benchmarking LLMs on the Road to AGI with a Focus on Model Confidence
abstract
As machine intelligence evolves, the need to test and compare the problem-solving abilities of different AI models grows. However, current benchmarks are often simplistic, allowing models to perform uniformly well and making it difficult to distinguish their capabilities. Additionally, benchmarks typically rely on static question-answer pairs that the models might memorize or guess. To address these limitations, we introduce Dynamic Intelligence Assessment (DIA), a novel methodology for testing AI models using dynamic question templates and improved metrics across multiple disciplines such as mathematics, cryptography, cybersecurity, and computer science. The accompanying dataset, DIA-Bench, contains a diverse collection of challenge templates with mutable parameters presented in various formats, including text, PDFs, compiled binaries, visual puzzles, and CTF-style cybersecurity challenges. Our framework introduces four new metrics to assess a model’s reliability and confidence across multiple attempts. These metrics revealed that even simple questions are frequently answered incorrectly when posed in varying forms, highlighting significant gaps in models’ reliability. Notably, API models like GPT-4o often overestimated their mathematical capabilities, while ChatGPT-4o demonstrated better performance due to effective tool usage. In self-assessment OpenAI’s o1-mini proved to have the best judgement on what tasks it should attempt to solve. We evaluated 25 state-of-the-art LLMs using DIA-Bench, showing that current models struggle with complex tasks and often display unexpectedly low confidence, even with simpler questions. The DIA framework sets a new standard for assessing not only problem-solving, but also a model’s adaptive intelligence and ability to assess its limitations. The dataset is publicly available on the project’s page: https://github.com/DIA-Bench.
Norbert Tihanyi, Tamás Bisztray, Richard A. Dubniczky, Rebeka Tóth, Bertalan Borsos, Bilel Cherif, Ridhi Jain, Lajos Muzsai, Mohamed Amine Ferrag, Ryan Marinelli, Lucas C. Cordeiro, Mérouane Debbah, Vasileios Mavroeidis, Audun Jøsang
IEEE Big Data10