Ivana Kajic

dblp:135/3514 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 44% Trustworthy machine learning · 20% Deep learning architectures and training · 14%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
concept erasure
0.912025
Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts · ICCV 2025
Machine learning › Trustworthy machine learning
generative model safety
0.912025
Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts · ICCV 2025
Visual content generation and editing › image generation
text-to-image generation
0.912025
Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human rating · ICLR 2025
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning
0.812024
Evaluating Numerical Reasoning in Text-to-Image Models · NeurIPS 2024
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image evaluation
0.812024
Evaluating Numerical Reasoning in Text-to-Image Models · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
Evaluating Numerical Reasoning in Text-to-Image Models · NeurIPS 2024
Machine learning › Deep learning architectures and training › memory-augmented neural networks
legendre memory unit
0.412019
Legendre Memory Units: Continuous-Time Representation in Recurrent Neural Networks · NeurIPS 2019
Machine learning › Deep learning architectures and training
recurrent neural network
0.412019
Legendre Memory Units: Continuous-Time Representation in Recurrent Neural Networks · NeurIPS 2019
Emerging computing paradigms › neuromorphic computing
neuromorphic circuits
0.412019
Legendre Memory Units: Continuous-Time Representation in Recurrent Neural Networks · NeurIPS 2019
Emerging computing paradigms
neuromorphic computing
0.412019
Legendre Memory Units: Continuous-Time Representation in Recurrent Neural Networks · NeurIPS 2019
Computer vision › Vision and language › cross-modal alignment › image-text alignment
text-to-image alignment
0.312025
Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human rating · ICLR 2025
Human-robot interaction
hand-eye coordination
0.212014
Learning hand-eye coordination for a humanoid robot using SOMs · HRI 2014
Robotics › Motion planning and robot control › robot learning
sensorimotor learning
0.112014
Learning hand-eye coordination for a humanoid robot using SOMs · HRI 2014

Methods — techniques the papers use, named apart from their topics

human annotation · 2.5correlation analysis · 1.7backpropagation through ODE solver · 0.8ordinary differential equations · 0.4ordinary differential equation · 0.4legendre polynomials · 0.4legendre polynomial · 0.4random walk · 0.4body babbling · 0.4self-organizing maps · 0.2self-organizing map · 0.2
YearPublicationVenuePosition
2025 Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts
Ibtihel Amara, Ahmed Imtiaz Humayun, Ivana Kajic, Zarana Parekh, Natalie Harris, Sarah Young, Chirag Nagpal, Najoung Kim, Junfeng He, Cristina Nader Vasconcelos, Deepak Ramachandran, Golnoosh Farnadi, Katherine A. Heller, Mohammad Havaei, Negar Rostamzadeh
ICCV3
2025 Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human rating
abstract
While text-to-image (T2I) generative models have become ubiquitous, they do not necessarily generate images that align with a given prompt. While many metrics and benchmarks have been proposed to evaluate T2I models and alignment metrics, the impact of the evaluation components (prompt sets, human annotations, evaluation task) has not been systematically measured. We find that looking at only *one slice of data*, i.e. one set of capabilities or human annotations, is not enough to obtain stable conclusions that generalise to new conditions or slices when evaluating T2I models or alignment metrics. We address this by introducing an evaluation suite of $>$100K annotations across four human annotation templates that comprehensively evaluates models' capabilities across a range of methods for gathering human annotations and comparing models. In particular, we propose (1) a carefully curated set of prompts -- *Gecko2K*; (2) a statistically grounded method of comparing T2I models; and (3) how to systematically evaluate metrics under three *evaluation tasks* -- *model ordering, pair-wise instance scoring, point-wise instance scoring*. Using this evaluation suite, we evaluate a wide range of metrics and find that a metric may do better in one setting but worse in another. As a result, we introduce a new, interpretable auto-eval metric that is consistently better correlated with human ratings than such existing metrics on our evaluation suite--across different human templates and evaluation settings--and on TIFA160.
Olivia Wiles, Isabela Albuquerque, Ivana Kajic, Su Wang 0001, Emanuele Bugliarello, Yasumasa Onoe, Pinelopi Papalampidi, Ira Ktena, Christopher Knutsen, Cyrus Rashtchian, Anant Nawalgaria, Jordi Pont-Tuset, Aida Nematzadeh
ICLR4
2024 Evaluating Numerical Reasoning in Text-to-Image Models
abstract
Text-to-image generative models are capable of producing high-quality images that often faithfully depict concepts described using natural language. In this work, we comprehensively evaluate a range of text-to-image models on numerical reasoning tasks of varying difficulty, and show that even the most advanced models have only rudimentary numerical skills. Specifically, their ability to correctly generate an exact number of objects in an image is limited to small numbers, it is highly dependent on the context the number term appears in, and it deteriorates quickly with each successive number. We also demonstrate that models have poor understanding of linguistic quantifiers (such as “few” or “as many as”), the concept of zero, and struggle with more advanced concepts such as fractional representations. We bundle prompts, generated images and human annotations into GeckoNum, a novel benchmark for evaluation of numerical reasoning.
Ivana Kajic, Olivia Wiles, Isabela Albuquerque, Su Wang 0001, Jordi Pont-Tuset, Aida Nematzadeh
NeurIPS1
2023 Evaluating Visual Number Discrimination in Deep Neural Networks
Ivana Kajic, Aida Nematzadeh
CogSci1
2021 Biologically Constrained Large-Scale Model of the Wisconsin Card Sorting Test
Ivana Kajic, Terrence C. Stewart
CogSci1
2020 Learning to cooperate: Emergent communication in multi-agent navigation
Ivana Kajic, Eser Aygün, Doina Precup
CogSci1
2019 Legendre Memory Units: Continuous-Time Representation in Recurrent Neural Networks
abstract
We propose a novel memory cell for recurrent neural networks that dynamically maintains information across long windows of time using relatively few resources. The Legendre Memory Unit~(LMU) is mathematically derived to orthogonalize its continuous-time history -- doing so by solving $d$ coupled ordinary differential equations~(ODEs), whose phase space linearly maps onto sliding windows of time via the Legendre polynomials up to degree $d - 1$. Backpropagation across LMUs outperforms equivalently-sized LSTMs on a chaotic time-series prediction task, improves memory capacity by two orders of magnitude, and significantly reduces training and inference times. LMUs can efficiently handle temporal dependencies spanning $100\text{,}000$ time-steps, converge rapidly, and use few internal state-variables to learn complex functions spanning long windows of time -- exceeding state-of-the-art performance among RNNs on permuted sequential MNIST. These results are due to the network's disposition to learn scale-invariant features independently of step size. Backpropagation through the ODE solver allows each layer to adapt its internal time-step, enabling the network to learn task-relevant time-scales. We demonstrate that LMU memory cells can be implemented using $m$ recurrently-connected Poisson spiking neurons, $\mathcal{O}( m )$ time and memory, with error scaling as $\mathcal{O}( d / \sqrt{m} )$. We discuss implementations of LMUs on analog and digital neuromorphic hardware.
Aaron Voelker, Ivana Kajic, Chris Eliasmith
NeurIPS2
2017 A Biologically Constrained Model of Semantic Memory Search
Ivana Kajic, Jan Gosmann, Brent Komer, Ryan W. Orr, Terrence C. Stewart, Chris Eliasmith
CogSci1
2016 Towards a Cognitively Realistic Representation of Word Associations
Ivana Kajic, Jan Gosmann, Terrence C. Stewart, Thomas Wennekers, Chris Eliasmith
CogSci1
2014 Learning hand-eye coordination for a humanoid robot using SOMs
abstract
Hand-eye coordination is an important motor skill acquired in infancy which precedes pointing behavior. Pointing facilitates social interactions by directing attention of engaged participants. It is thus essential for the natural flow of human-robot interaction. Here, we attempt to explain how pointing emerges from sensorimotor learning of hand-eye coordination in a humanoid robot. During a body babbling phase with a random walk strategy, a robot learned mappings of joints for different arm postures. Arm joint configurations were used to train biologically inspired models consisting of SOMs. We show that such a model implemented on a robotic platform accounts for pointing behavior while humans present objects out of reach of the robot's hand.
Ivana Kajic, Guido Schillaci, Sasa Bodiroza, Verena V. Hafner
HRI1