Roland Daynauth

dblp:347/8696 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0008-1149-0179ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 50% Learning theory · 50%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat · ACL (1) 2025
Machine learning › Learning theory
ranking
0.912025
Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat · ACL (1) 2025
YearPublicationVenuePosition
2025 Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat
abstract
Roland Daynauth, Christopher Clarke, Krisztian Flautner, Lingjia Tang, Jason Mars. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Roland Daynauth, Christopher Clarke, Krisztián Flautner, Lingjia Tang, Jason Mars
ACL (1)1
2024 Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production
abstract
Many companies use large language models (LLMs) offered as a service, like OpenAl's GPT-4, to create AI-enabled product experiences. Along with the benefits of ease-of-use and shortened time-to-solution, this reliance on proprietary services has downsides in model control, performance reliability, uptime predictability, and cost. At the same time, a flurry of open-source small language models (SLMs) has been made avail-able for commercial use. However, their readiness to replace existing capabilities remains unclear, and a systematic approach to holistically evaluate these SLMs is not readily available. This paper presents a systematic evaluation methodology and a characterization of modern open-source SLMs and their trade-offs when replacing proprietary LLMs for a real-world product feature. We have designed SLaM, an open-source automated analysis tool that enables the quantitative and qualitative testing of product features utilizing arbitrary SLMs. Using SLaM, we examine the quality and performance characteristics of modern SLMs relative to an existing customer-facing implementation using the OpenAI GPT-4 API. Across 9 SLMs and their 29 variants, we observe that SLMs provide competitive results, significant performance consistency improvements, and a cost reduction of 5xrv29x when compared to GPT-4.
Chandra Irugalbandara, Ashish Mahendra, Roland Daynauth, Tharuka Kasthuri Arachchige, Jayanaka L. Dantanarayana, Krisztián Flautner, Lingjia Tang, Yiping Kang, Jason Mars
ISPASS3