Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Krystal Kallarackal

dblp:344/8807 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0002-2337-0114ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Human-computer interaction and pervasive computing
1 paper
Accessibility and assistive technology · 77% Human-AI interaction · 23%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics
interactive visualization
0.912025
LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language Models · IEEE Trans. Vis. Comput. Graph. 2025
Visualization and visual analytics › visual analytics
visual analytics for machine learning
0.912025
LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language Models · IEEE Trans. Vis. Comput. Graph. 2025
Accessibility and assistive technology
augmentative and alternative communication
0.712023
"The less I type, the better": How AI Language Models can Enhance or Impede Communication for AAC Users · CHI 2023
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
0.312025
LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language Models · IEEE Trans. Vis. Comput. Graph. 2025

Methods — techniques the papers use, named apart from their topics

visual analytics workflow · 1.7large language model · 0.7
YearPublicationVenuePosition
2025 LLM Comparator: Interactive Analysis of Side-by-Side Evaluation of Large Language Models
abstract
Evaluating large language models (LLMs) presents unique challenges. While automatic side-by-side evaluation, also known as LLM-as-a-judge, has become a promising solution, model developers and researchers face difficulties with scalability and interpretability when analyzing these evaluation outcomes. To address these challenges, we introduce LLM Comparator, a new visual analytics tool designed for side-by-side evaluations of LLMs. This tool provides analytical workflows that help users understand when and why one LLM outperforms or underperforms another, and how their responses differ. Through close collaboration with practitioners developing LLMs at Google, we have iteratively designed, developed, and refined the tool. Qualitative feedback from these users highlights that the tool facilitates in-depth analysis of individual examples while enabling users to visually overview and flexibly slice data. This empowers users to identify undesirable patterns, formulate hypotheses about model behavior, and gain insights for model improvement. LLM Comparator has been integrated into Google's LLM evaluation platforms and open-sourced.
Minsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu, James Wexler, Emily Reif, Krystal Kallarackal, Minsuk Chang, Michael Terry, Lucas Dixon
IEEE Trans. Vis. Comput. Graph.7
2023 "The less I type, the better": How AI Language Models can Enhance or Impede Communication for AAC Users
abstract
Users of augmentative and alternative communication (AAC) devices sometimes find it difficult to communicate in real time with others due to the time it takes to compose messages. AI technologies such as large language models (LLMs) provide an opportunity to support AAC users by improving the quality and variety of text suggestions. However, these technologies may fundamentally change how users interact with AAC devices as users transition from typing their own phrases to prompting and selecting AI-generated phrases. We conducted a study in which 12 AAC users tested live suggestions from a language model across three usage scenarios: extending short replies, answering biographical questions, and requesting assistance. Our study participants believed that AI-generated phrases could save time, physical and cognitive effort when communicating, but felt it was important that these phrases reflect their own communication style and preferences. This work identifies opportunities and challenges for future AI-enhanced AAC devices.
Stephanie Valencia, Richard Cave, Krystal Kallarackal, Katie Seaver, Michael Terry, Shaun K. Kane
CHI3