Krista A. Ehinger

dblp:13/8654 · DBLP profile ↗
← Back
31ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0003-2247-3020ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Fast concept-based counterfactual explanations for image classification
Ruihan Zhang 0002, Tim Miller 0001, Krista A. Ehinger, Benjamin I. P. Rubinstein
Artif. Intell.3
2025 State-Based Disassembly Planning
abstract
It has been shown recently that physics-based simulation significantly enhances the disassembly capabilities of real-world assemblies with diverse 3D shapes and stringent motion constraints. However, the efficiency suffers when tackling intricate disassembly tasks that require numerous simulations and increased simulation time. In this work, we propose a State-Based Disassembly Planning (SBDP) approach, prioritizing physics-based simulation with translational motion over rotational motion to facilitate autonomy, reducing dependency on human input, while storing intermediate motion states to improve search scalability. We introduce two novel evaluation functions derived from new Directional Blocking Graphs (DBGs) enriched with state information to scale up the search. Our experiments show that SBDP with new evaluation functions and DBGs constraints outperforms the state-of-the-art in disassembly planning in terms of success rate and computational efficiency over benchmark datasets consisting of thousands of physically valid industrial assemblies.
Nir Lipovetzky, Krista A. Ehinger
AAAI3
2025 TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model
abstract
We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to prevent overfitting. Subsequently, it uses a triplane-aware cross-attention diffusion model to learn and integrate these features effectively. Furthermore, the features generated by the diffusion model can be rapidly transformed into 3D volumes using a pre-trained decoder module. Our experiments on three different scales of medical datasets, BrainTumour 128x128x128, Pancreas 256x256x256, and Colon 512x512x512, demonstrated outstanding results. We utilized MSE and SSIM to evaluate reconstruction quality and leveraged the Wasserstein Generative Adversarial Network (W-GAN) critic to assess generative quality. Comparisons to existing approaches show that our method gives better reconstruction and generation results than other encoder-decoder methods with similar-sized latent spaces.
Zhenkai Zhang 0001, Krista A. Ehinger, Tom Drummond
AAAI2
2025 Planning-Driven Programming: A Large Language Model Programming Workflow
abstract
The strong performance of large language models (LLMs) raises extensive discussion on their application to code generation.Recent research suggests continuous program refinements through visible tests to improve code generation accuracy in LLMs.However, these methods suffer from LLMs' inefficiency and limited reasoning capacity.In this work, we propose an LLM programming workflow (LPW) designed to improve both initial code generation and subsequent refinements within a structured two-phase workflow.Specifically, the solution generation phase formulates a solution plan, which is then verified through visible tests to specify the intended natural language solution.Subsequently, the code implementation phase drafts an initial code according to the solution plan and its verification.If the generated code fails the visible tests, the plan verification serves as the intended solution to consistently inform the refinement process for correcting bugs.Compared to state-of-the-art methods across various existing LLMs, LPW significantly improves the Pass@1 accuracy by up to 16.4% on well-established text-tocode generation benchmarks.LPW also sets new state-of-the-art Pass@1 accuracy, achieving 98.2% on HumanEval, 84.8% on MBPP, 59.3% on LiveCode, 62.6% on APPS, and 34.7% on CodeContests, using GPT-4o as the backbone.Our code is publicly available at
Yanchuan Chang, Nir Lipovetzky, Krista A. Ehinger
ACL (1)4
2025 Open-World Amodal Appearance Completion
abstract
Understanding and reconstructing occluded objects is a challenging problem, especially in open-world scenarios where categories and contexts are diverse and unpredictable. Traditional methods, however, are typically restricted to closed sets of object categories, limiting their use in complex, open-world scenes. We introduce Open-World Amodal Appearance Completion, a training-free framework that expands amodal completion capabilities by accepting flexible text queries as input. Our approach generalizes to arbitrary objects specified by both direct terms and abstract queries. We term this capability reasoning amodal completion, where the system reconstructs the full appearance of the queried object based on the provided image and language query. Our framework unifies segmentation, occlusion analysis, and inpainting to handle complex occlusions and generates completed objects as RGBA elements, enabling seamless integration into applications such as 3D reconstruction and image editing. Extensive evaluations demonstrate the effectiveness of our approach in generalizing to novel objects and occlusions, establishing a new benchmark for amodal completion in open-world settings. Code and datasets available: https://github.com/saraao/amodal.
Jiayang Ao, Yanbei Jiang, Qiuhong Ke, Krista A. Ehinger
CVPR4
2025 Do Explanations Expose Bias? How Saliency Maps Affect Judgements of Biased Face-Recognition Models
abstract
Saliency-map explanations are intended to make computer-vision models more transparent, but it is unclear whether they help people recognise biased behaviour. We conducted a controlled on-line study with 40 participants who compared Layer-wise Relevance Propagation maps from convolutional face-recognition models. A fair model was trained on a balanced synthetic dataset; two biased models were trained on data in which either light- or dark-skinned faces appeared only in frontal pose. Each participant completed 32 comparison trials. When the fair model was paired with the dark-skinned-pose-biased model, selections were near chance (52.8% favouring the fair model, binomial p = .36). When the fair model was paired with the light-skinned-pose-biased model, participants chose the biased model significantly more often (58.1%, p = .005). Confidence ratings varied with condition and did not systematically track model fairness. These results indicate that pixel-level attribution alone does not reliably expose training bias and can, in some settings, mislead non-expert users.
Justyn Rodrigues, Krista A. Ehinger, Oliver Obst, X. Rosalind Wang
ECAI2
2025 Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
abstract
Multimodal Large Language Models (MLLMs) have demonstrated exceptional performance across various objective multimodal perception tasks, yet their application to subjective, emotionally nuanced domains, such as psychological analysis, remains largely unexplored. In this paper, we introduce PICK, a multi-step framework designed for Psychoanalytical Image Comprehension through hierarchical analysis and Knowledge injection with MLLMs, specifically focusing on the House-Tree-Person (HTP) Test, a psychological assessment test. First, we decompose drawings containing multiple instances into semantically meaningful sub-drawings, constructing a hierarchical representation that captures spatial structure and content across three levels: single-object level, multi-object level, and whole level. Next, we analyze these sub-drawings at each level with a targeted focus, extracting psychological or emotional insights from their visual cues. We also introduce an HTP knowledge base and design a feature extraction module, trained with reinforcement learning, to generate a psychological profile for single-object level analysis. This profile captures both holistic stylistic features and dynamic object-specific features (such as those of the house, tree, or person), correlating them with psychological states. Finally, we integrate these multi-faceted information to produce a well-informed assessment that aligns with expert-level reasoning. Our approach bridges the gap between MLLMs and specialized expert domains, offering a structured and interpretable framework for understanding human mental states through visual expression. Experimental results demonstrate that the proposed PICK significantly enhances the capability of MLLMs in psychological analysis. It is further validated as a general framework through extensions to emotion understanding tasks. Codes are released at https://github.com/YanbeiJiang/PICK.
Xueqi Ma, Yanbei Jiang, Sarah M. Erfani, James Bailey 0001, Weifeng Liu 0001, Krista A. Ehinger, Jey Han Lau
ACM Multimedia6
2024 Generalized Planning for the Abstraction and Reasoning Corpus
abstract
The Abstraction and Reasoning Corpus (ARC) is a general artificial intelligence benchmark that poses difficulties for pure machine learning methods due to its requirement for fluid intelligence with a focus on reasoning and abstraction. In this work, we introduce an ARC solver, Generalized Planning for Abstract Reasoning (GPAR). It casts an ARC problem as a generalized planning (GP) problem, where a solution is formalized as a planning program with pointers. We express each ARC problem using the standard Planning Domain Definition Language (PDDL) coupled with external functions representing object-centric abstractions. We show how to scale up GP solvers via domain knowledge specific to ARC in the form of restrictions over the actions model, predicates, arguments and valid structure of planning programs. Our experiments demonstrate that GPAR outperforms the state-of-the-art solvers on the object-centric tasks of the ARC, showing the effectiveness of GP and the expressiveness of PDDL to model ARC problems. The challenges provided by the ARC benchmark motivate research to advance existing GP solvers and understand new relations with other planning computational models. Code is available at github.com/you68681/GPAR.
Nir Lipovetzky, Krista A. Ehinger
AAAI3
2024 Sequential Amodal Segmentation via Cumulative Occlusion Learning
Jiayang Ao, Qiuhong Ke, Krista A. Ehinger
BMVC3
2024 KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
Yanbei Jiang, Krista A. Ehinger, Jey Han Lau
IJCAI2
2024 Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers
abstract
We present a novel bi-directional Transformer architecture (BiXT) which scales linearly with input size in terms of computational cost and memory consumption, but does not suffer the drop in performance or limitation to only one input modality seen with other efficient Transformer-based approaches. BiXT is inspired by the Perceiver architectures but replaces iterative attention with an efficient bi-directional cross-attention module in which input tokens and latent variables attend to each other simultaneously, leveraging a naturally emerging attention-symmetry between the two. This approach unlocks a key bottleneck experienced by Perceiver-like architectures and enables the processing and interpretation of both semantics ('what') and location ('where') to develop alongside each other over multiple layers -- allowing its direct application to dense and instance-based tasks alike. By combining efficiency with the generality and performance of a full Transformer architecture, BiXT can process longer sequences like point clouds, text or images at higher feature resolutions and achieves competitive performance across a range of tasks like point cloud part segmentation, semantic image segmentation, image classification, hierarchical sequence modeling and document retrieval. Our experiments demonstrate that BiXT models outperform larger competitors by leveraging longer sequences more efficiently on vision tasks like classification and segmentation, and perform on par with full Transformer variants on sequence modeling and document retrieval -- but require 28\% fewer FLOPs and are up to $8.4\times$ faster.
Markus Hiller, Krista A. Ehinger, Tom Drummond
NeurIPS2
2024 Amodal Intra-class Instance Segmentation: Synthetic Datasets and Benchmark
abstract
Images of realistic scenes often contain intra-class objects that are heavily occluded from each other, making the amodal perception task that requires parsing the occluded parts of the objects challenging. Although important for downstream tasks such as robotic grasping systems, the lack of large-scale amodal datasets with detailed annotations makes it difficult to model intra-class occlusions explicitly. This paper introduces two new amodal datasets for image amodal completion tasks, which contain a total of over 267K images of intra-class occlusion scenarios, annotated with multiple masks, amodal bounding boxes, dual order relations and full appearance for instances and background. We also present a point-supervised scheme with layer priors for amodal instance segmentation specifically designed for intra-class occlusion scenarios1. Experiments show that our weakly supervised approach outperforms the SOTA fully supervised methods, while our layer priors design exhibits remarkable performance improvements in the case of intra-class occlusion in both synthetic and real images.
Jiayang Ao, Qiuhong Ke, Krista A. Ehinger
WACV3
2023 Improving Denoising Diffusion Models via Simultaneous Estimation of Image and Noise
Zhenkai Zhang 0001, Krista A. Ehinger, Tom Drummond
ACML2
2023 Unicode Analogies: An Anti-Objectivist Visual Reasoning Challenge
abstract
Analogical reasoning enables agents to extract relevant information from scenes, and efficiently navigate them in familiar ways. While progressive-matrix problems (PMPs) are becoming popular for the development and evaluation of analogical reasoning in computer vision, we argue that the dominant methodology in this area struggles to expose the lack of meaningful generalisation in solvers, and rein-forces an objectivist stance on perception - that objects can only be seen one way - which we believe to be counter-productive. In this paper, we introduce the Unicode Analogies challenge, consisting of polysemic, character-based PMPs to benchmark fluid conceptualisation ability in vision systems. Writing systems have evolved characters at multiple levels of abstraction, from iconic through to symbolic representations, producing both visually interrelated yet exceptionally diverse images when compared to those exhibited by existing PMP datasets. Our framework has been designed to challenge models by presenting tasks much harder to complete without robust feature extraction, while remaining largely solvable by human participants. We therefore argue that Unicode Analogies elegantly captures and tests for a facet of human visual reasoning that is severely lacking in current-generation AI.
Steven Spratley, Krista A. Ehinger, Tim Miller 0001
CVPR2
2023 Truck Speed Detection Through Video Streams
abstract
Accurately assessing the speed of vehicles is important for traffic management systems. This is especially the case for heavy goods vehicles such as lorries/trucks, since they cannot easily stop at short notice. Previous work has shown that deep learning can be used for identifying and distinguishing trucks on the road from other vehicles, e.g., [1], however accurately estimating their speed from roadside cameras remains a challenge. One solution we employ is using video data from the roadside cameras, then extracting the speeds of vehicles in the video from the Infra-Red Traffic Logger (TIRTL) systems, which are provided by the Department of Transport, Victoria. The TIRTL system is very accurate but expensive and only deployed at a few key locations around Melbourne. A solution that works at the edge and uses lightweight Internet-of-Things devices to produce accurate speed data is thus highly desirable. In this paper, we propose a Convolutional Neural Network (CNN) model using a light-weight Siamese backbone and associated feature correlations to track and detect the speed of trucks. We build a dataset that contains images with speed and bounding-box annotations to train the proposed model. To enable the model to maintain a high degree of accuracy with different camera setups, we train and test the proposed model using image augmentation. The results show our model has an average speed estimation error of 4.92% and an average Intersection over Union (IoU) of 75.8% whilst incorporating different intrinsic and extrinsic parameters based on image augmentation. Such a capability has the potential to change the way services are deployed across the road network to record vehicle types and speeds.
Zuo Huang, Richard O. Sinnott, Krista A. Ehinger
e-Science3
2023 Novelty and Lifted Helpful Actions in Generalized Planning
abstract
It has been shown recently that successful techniques in classical planning, such as goal-oriented heuristics and landmarks, can improve the ability to compute planning programs for generalized planning (GP) problems. In this work, we introduce the notion of action novelty rank, which computes novelty with respect to a planning program, and propose novelty-based generalized planning solvers, which prune a newly generated planning program if its most frequent action repetition is greater than a given bound v, implemented by novelty-based best-first search BFS(v) and its progressive variant PGP(v). Besides, we introduce lifted helpful actions in GP derived from action schemes, and propose new evaluation functions and structural program restrictions to scale up the search. Our experiments show that the new algorithms BFS(v) and PGP(v) outperform the state-of-the-art in GP over the standard generalized planning benchmarks. Practical findings on the above-mentioned methods in generalized planning are briefly discussed.
Nir Lipovetzky, Krista A. Ehinger
SOCS3
2023 Image amodal completion: A survey
Jiayang Ao, Qiuhong Ke, Krista A. Ehinger
Comput. Vis. Image Underst.3
2023 An active foveated gaze prediction algorithm based on a Bayesian ideal observer
Shima Rashidi, Weilun Xu, Dian Lin, Andrew Turpin, Lars Kulik, Krista A. Ehinger
Pattern Recognit.6
2022 Lightness constancy in reality, in virtual reality, and on flat-panel displays
Khushbu Patel, Laurie M. Wilcox, Laurence T. Maloney, Krista A. Ehinger, Jaykishan Y. Patel, Emma Wiedenmann, Richard F. Murray
CogSci4
2021 Invertible Concept-based Explanations for CNN Models with Non-negative Concept Activation Vectors
abstract
Convolutional neural network (CNN) models for computer vision are powerful but lack explainability in their most basic form. This deficiency remains a key challenge when applying CNNs in important domains. Recent work on explanations through feature importance of approximate linear models has moved from input-level features (pixels or segments) to features from mid-layer feature maps in the form of concept activation vectors (CAVs). CAVs contain concept-level information and could be learned via clustering. In this work, we rethink the ACE algorithm of Ghorbani et~al., proposing an alternative invertible concept-based explanation (ICE) framework to overcome its shortcomings. Based on the requirements of fidelity (approximate models to target models) and interpretability (being meaningful to people), we design measurements and evaluate a range of matrix factorization methods with our framework. We find that non-negative concept activation vectors (NCAVs) from non-negative matrix factorization provide superior performance in interpretability and fidelity based on computational and human subject experiments. Our framework provides both local and global concept-level explanations for pre-trained CNN models.
Ruihan Zhang 0002, Prashan Madumal, Tim Miller 0001, Krista A. Ehinger, Benjamin I. P. Rubinstein
AAAI4
2020 A Closer Look at Generalisation in RAVEN
Steven Spratley, Krista A. Ehinger, Tim Miller 0001
ECCV (27)2
2020 Optimal visual search based on a model of target detectability in natural images
abstract
To analyse visual systems, the concept of an ideal observer promises an optimal response for a given task. Bayesian ideal observers can provide optimal responses under uncertainty, if they are given the true distributions as input. In visual search tasks, prior studies have used signal to noise ratio (SNR) or psychophysics experiments to set the distributional parameters for simple targets on backgrounds with known patterns, however these methods do not easily translate to complex targets on natural scenes. Here, we develop a model of target detectability in natural images to estimate the parameters of target-present and target-absent distributions for a visual search task. We present a novel approach for approximating the foveated detectability of a known target in natural backgrounds based on biological aspects of human visual system. Our model considers both the uncertainty about target position and the visual system's variability due to its reduced performance in the periphery compared to the fovea. Our automated prediction algorithm uses trained logistic regression as a post processing phase of a pre-trained deep neural network. Eye tracking data from 12 observers detecting targets on natural image backgrounds are used as ground truth to tune foveation parameters and evaluate the model, using cross-validation. Finally, the model of target detectability is used in a Bayesian ideal observer model of visual search, and compared to human search performance.
Shima Rashidi, Krista A. Ehinger, Andrew Turpin, Lars Kulik
NeurIPS2
2017 A novel graph-based optimization framework for salient object detection
Jinxia Zhang, Krista A. Ehinger, Haikun Wei, Kan-Jian Zhang, Jing-Yu Yang 0001
Pattern Recognit.2
2017 Erratum to: A novel graph-based optimization framework for salient object detection [Pattern Recognition 64C (2017) 39-50]
Jinxia Zhang, Krista A. Ehinger, Haikun Wei, Kan-Jian Zhang, Jing-Yu Yang 0001
Pattern Recognit.2
2016 SUN Database: Exploring a Large Collection of Scene Categories
Jianxiong Xiao, Krista A. Ehinger, James Hays, Antonio Torralba 0001, Aude Oliva
Int. J. Comput. Vis.2
2014 A prior-based graph for salient object detection
abstract
Recently, various graph-based methods have be proposed for salient object detection. These algorithms represent image points and their similarity as nodes and edges in a graph. Although the edge structure and weighting are the heart of these methods, the graph construction has not been studied in detail. In this paper, we exploit image priors, including spatial priors, color priors, and a central bias prior, to construct the graph. We connect nodes which are spatially close in the image, nodes which have similar color features, and the boundary nodes along the borders of the image, while weighting edges according to both their color similarity and spatial proximity. Moreover, we propose a new sine spatial distance instead of the commonly-used Euclidean spatial distance, which better captures the central bias in scenes. Extensive experiments show that our method outperforms thirteen state-of-the-art methods on four different image databases.
Jinxia Zhang, Krista A. Ehinger, Jundi Ding, Jing-Yu Yang 0001
ICIP2
2012 Recognizing scene viewpoint using panoramic place representation
abstract
We introduce the problem of scene viewpoint recognition, the goal of which is to classify the type of place shown in a photo, and also recognize the observer's viewpoint within that category of place. We construct a database of 360° panoramic images organized into 26 place categories. For each category, our algorithm automatically aligns the panoramas to build a full-view representation of the surrounding place. We also study the symmetry properties and canonical viewpoint of each place category. At test time, given a photo of a scene, the model can recognize the place category, produce a compass-like indication of the observer's most likely viewpoint within that place, and use this information to extrapolate beyond the available view, filling in the probable visual layout that would appear beyond the boundary of the photo.
Jianxiong Xiao, Krista A. Ehinger, Aude Oliva, Antonio Torralba 0001
CVPR2
2011 Canonical views of scenes depend on the shape of the space
Krista A. Ehinger, Aude Oliva
CogSci1
2011 Estimating scene typicality from human ratings and image features
Krista A. Ehinger, Jianxiong Xiao, Antonio Torralba 0001, Aude Oliva
CogSci1
2010 SUN database: Large-scale scene recognition from abbey to zoo
abstract
Scene categorization is a fundamental problem in computer vision. However, scene understanding research has been constrained by the limited scope of currently-used databases which do not capture the full variety of scene categories. Whereas standard databases for object categorization contain hundreds of different classes of objects, the largest available dataset of scene categories contains only 15 classes. In this paper we propose the extensive Scene UNderstanding (SUN) database that contains 899 categories and 130,519 images. We use 397 well-sampled categories to evaluate numerous state-of-the-art algorithms for scene recognition and establish new bounds of performance. We measure human scene classification performance on the SUN database and compare this with computational methods. Additionally, we study a finer-grained scene representation to detect scenes embedded inside of larger scenes.
Jianxiong Xiao, James Hays, Krista A. Ehinger, Aude Oliva, Antonio Torralba 0001
CVPR3
2009 Learning to predict where humans look
abstract
For many applications in graphics, design, and human computer interaction, it is essential to understand where humans look in a scene. Where eye tracking devices are not a viable option, models of saliency can be used to predict fixation locations. Most saliency approaches are based on bottom-up computation that does not consider top-down image semantics and often does not match actual eye movements. To address this problem, we collected eye tracking data of 15 viewers on 1003 images and use this database as training and testing examples to learn a model of saliency based on low, middle and high-level image features. This large database of eye tracking data is publicly available with this paper.
Tilke Judd, Krista A. Ehinger, Frédo Durand, Antonio Torralba 0001
ICCV2