Tuomas P. Oikarinen

dblp:243/3532 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Interpretable Generative Models through Post-hoc Concept Bottlenecks
abstract
Concept bottleneck models (CBM) aim to produce inherently interpretable models that rely on human-understandable concepts for their predictions. However, existing approaches to design interpretable generative models based on CBMs are not yet efficient and scalable, as they require expensive generative model training from scratch as well as real images with labor-intensive concept supervision. To address these challenges, we present two novel and low-cost methods to build interpretable generative models through post-hoc techniques and we name our approaches: concept-bottleneck autoencoder (CB-AE) and concept controller (CC). Our proposed approaches enable efficient and scalable training without the need of real data and require only minimal to no concept supervision. Additionally, our methods generalize across modern generative model families including generative adversarial networks and diffusion models. We demonstrate the superior interpretability and steerability of our methods on numerous standard datasets like CelebA, CelebA-HQ, and CUB with large improvements (average ∼25%) over the prior work, while being 4-15× faster to train. Finally, a large-scale user study is performed to validate the inter-pretability and steerability of our methods1.
Akshay R. Kulkarni, Chung-En Sun, Tuomas P. Oikarinen, Tsui-Wei Weng
CVPR4
2025 Concept Bottleneck Language Models For Protein Design
abstract
We introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept. Our architecture offers three key benefits: i) Control: We can intervene on concept values to precisely control the properties of generated proteins, achieving a 3$\times$ larger change in desired concept values compared to baselines. ii) Interpretability: A linear mapping between concept values and predicted tokens allows transparent analysis of the model's decision-making process. iii) Debugging: This transparency facilitates easy debugging of trained models. Our models achieve pre-training perplexity and downstream task performance comparable to traditional masked protein language models, demonstrating that interpretability does not compromise performance. While adaptable to any language model, we focus on masked protein language models due to their importance in drug discovery and the ability to validate our model's capabilities through real-world experiments and expert knowledge. We scale our CB-pLM from 24 million to 3 billion parameters, making them the largest Concept Bottleneck Models trained and the first capable of generative language modeling.
Aya Abdelsalam Ismail, Tuomas P. Oikarinen, Amy Wang, Julius Adebayo, Samuel Stanton, Héctor Corrada Bravo, Kyunghyun Cho, Nathan C. Frey
ICLR2
2025 Concept Bottleneck Large Language Models
abstract
We introduce Concept Bottleneck Large Language Models (CB-LLMs), a novel framework for building inherently interpretable Large Language Models (LLMs). In contrast to traditional black-box LLMs that rely on limited post-hoc interpretations, CB-LLMs integrate intrinsic interpretability directly into the LLMs -- allowing accurate explanations with scalability and transparency. We build CB-LLMs for two essential NLP tasks: text classification and text generation. In text classification, CB-LLMs is competitive with, and at times outperforms, traditional black-box models while providing explicit and interpretable reasoning. For the more challenging task of text generation, interpretable neurons in CB-LLMs enable precise concept detection, controlled generation, and safer outputs. The embedded interpretability empowers users to transparently identify harmful content, steer model behavior, and unlearn undesired concepts -- significantly enhancing the safety, reliability, and trustworthiness of LLMs, which are critical capabilities notably absent in existing language models.
Chung-En Sun, Tuomas P. Oikarinen, Berk Ustun, Tsui-Wei Weng
ICLR2
2025 Evaluating Neuron Explanations: A Unified Framework with Sanity Checks
abstract
Understanding the function of individual units in a neural network is an important building block for mechanistic interpretability. This is often done by generating a simple text explanation of the behavior of individual neurons or units. For these explanations to be useful, we must understand how reliable and truthful they are. In this work we unify many existing explanation evaluation methods under one mathematical framework. This allows us to compare and contrast existing evaluation metrics, understand the evaluation pipeline with increased clarity and apply existing statistical concepts on the evaluation. In addition, we propose two simple sanity checks on the evaluation metrics and show that many commonly used metrics fail these tests and do not change their score after massive changes to the concept labels. Based on our experimental and theoretical results, we propose guidelines that future evaluations should follow and identify a set of reliable evaluation metrics.
Tuomas P. Oikarinen, Tsui-Wei Weng
ICML1
2025 SAND: Enhancing Open-Set Neuron Descriptions through Spatial Awareness
abstract
We propose Spatially-Aware open-set Network Dissection (SAND), a technique to identify and label the learned representation of the neurons of deep vision networks. Con-trary to earlier open-vocabulary neuron explanation meth-ods, we also leverage a neuron's spatial pattern of activation to guide our predictions towards more accurate and rel-evant concepts, while avoiding being misled by confounding visual information. We highlight important regions for a neuron through image masking, which has the advantage of being able to block out irrelevant concepts from an im-age, handling irregularly shaped activation regions, and re-vealing the visual concepts that a neuron learns in order to identify objects. We use CLIP to connect highly activating image regions with descriptive concepts, and measure the quality of our results through human evaluation. Further, since such manual evaluation can be highly time consuming, costly, and unscalable, we also propose an automated approach which uses image generation to get quantitative feedback on the generated concepts. Finally, as an application of our interpretability method, we demonstrate how it can be tuned to the medical domain. Our code is available at https://github.com/Trustworthy-ML-LabISAND
Anvita A. Srinivas, Tuomas P. Oikarinen, Divyansh Srivastava, Wei-Hung Weng, Tsui-Wei Weng
WACV2
2024 Linear Explanations for Individual Neurons
abstract
In recent years many methods have been developed to understand the internal workings of neural networks, often by describing the function of individual neurons in the model. However, these methods typically only focus on explaining the very highest activations of a neuron. In this paper we show this is not sufficient, and that the highest activation range is only responsible for a very small percentage of the neuron’s causal effect. In addition, inputs causing lower activations are often very different and can’t be reliably predicted by only looking at high activations. We propose that neurons should instead be understood as a linear combination of concepts, and develop an efficient method for producing these linear explanations. In addition, we show how to automatically evaluate description quality using simulation, i.e. predicting neuron activations on unseen inputs in vision setting.
Tuomas P. Oikarinen, Tsui-Wei Weng
ICML1
2023 Corrupting Neuron Explanations of Deep Visual Features
abstract
The inability of DNNs to explain their black-box behavior has led to a recent surge of explainability methods. However, there are growing concerns that these explainability methods are not robust and trustworthy. In this work, we perform the first robustness analysis of Neuron Explanation Methods under a unified pipeline and show that these explanations can be significantly corrupted by random noises and well-designed perturbations added to their probing data. We find that even adding small random noise with a standard deviation of 0.02 can already change the assigned concepts of up to 28% neurons in the deeper layers. Furthermore, we devise a novel corruption algorithm and show that our algorithm can manipulate the explanation of more than 80% neurons by poisoning less than 10% of probing data. This raises the concern of trusting Neuron Explanation Methods in real-life safety and fairness critical applications.
Divyansh Srivastava, Tuomas P. Oikarinen, Tsui-Wei Weng
ICCV2
2023 Label-free Concept Bottleneck Models
Tuomas P. Oikarinen, Subhro Das, Lam M. Nguyen, Tsui-Wei Weng
ICLR1
2023 CLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networks
Tuomas P. Oikarinen, Tsui-Wei Weng
ICLR1
2021 GraphMDN: Leveraging graph structure and deep learning to solve inverse problems
abstract
The recent introduction of Graph Neural Networks (GNNs) and their growing popularity in the past few years has enabled the application of deep learning algorithms to non-Euclidean, graph-structured data. GNNs have achieved state-of-the-art results across an impressive array of graph-based machine learning problems. Nevertheless, despite their rapid pace of development, much of the work on GNNs has focused on graph classification and embedding techniques, largely ignoring regression tasks over graph data. In this paper, we develop a Graph Mixture Density Network (GraphMDN), which combines graph neural networks with mixture density network (MDN) outputs. By combining these techniques, GraphMDNs have the advantage of naturally being able to incorporate graph structured information into a neural architecture, as well as the ability to model multi-modal regression targets. As such, GraphMDNs are designed to excel on regression tasks wherein the data are graph structured, and target statistics are better represented by mixtures of densities rather than singular values (so-called “inverse problems”). To demonstrate this, we extend an existing GNN architecture known as Semantic GCN (SemGCN) to a GraphMDN structure, and show results from the Human3.6M pose estimation task. The extended model consistently outperforms both GCN and MDN architectures on their own, with a comparable number of parameters.
Tuomas P. Oikarinen, Daniel C. Hannah, Sohrob Kazerounian
IJCNN1
2021 Robust Deep Reinforcement Learning through Adversarial Loss
abstract
Recent studies have shown that deep reinforcement learning agents are vulnerable to small adversarial perturbations on the agent's inputs, which raises concerns about deploying such agents in the real world. To address this issue, we propose RADIAL-RL, a principled framework to train reinforcement learning agents with improved robustness against $l_p$-norm bounded adversarial attacks. Our framework is compatible with popular deep reinforcement learning algorithms and we demonstrate its performance with deep Q-learning, A3C and PPO. We experiment on three deep RL benchmarks (Atari, MuJoCo and ProcGen) to show the effectiveness of our robust training algorithm. Our RADIAL-RL agents consistently outperform prior methods when tested against attacks of varying strength and are more computationally efficient to train. In addition, we propose a new evaluation method called Greedy-Worst-Case Reward (GWC) to measure attack agnostic robustness of deep RL agents. We show that GWC can be evaluated efficiently and is a good estimate of the reward under the worst possible sequence of adversarial attacks. All code used for our experiments is available at https://github.com/tuomaso/radial_rl_v2.
Tuomas P. Oikarinen, Alexandre Megretski, Luca Daniel, Tsui-Wei Weng
NeurIPS1
2019 Landslide Geohazard Assessment with Convolutional Neural Networks Using Sentinel-2 Imagery Data
abstract
In this paper, the authors aim to combine the latest state of the art models in image recognition with the best publicly available satellite images to create a system for landslide risk mitigation. We focus first on landslide detection and further propose a similar system to be used for prediction. Such models are valuable as they could easily be scaled up to provide data for hazard evaluation, as satellite imagery becomes increasingly available. The goal is to use satellite images and correlated data to enrich the public repository of data and guide disaster relief efforts for locating precise areas where landslides have occurred. Different image augmentation methods are used to increase diversity in the chosen dataset and create more robust classification. The resulting outputs are then fed into variants of 3-D convolutional neural networks. A review of the current literature indicates there is no research using CNNs (Convolutional Neural Networks) and freely available satellite imagery for classifying landslide risk. The model has shown to be ultimately able to achieve a significantly better than baseline accuracy.
Silvia Liberata Ullo, Maximillian S. Langenkamp, Tuomas P. Oikarinen, Maria P. del Rosso, Alessandro Sebastianelli, Federica Piccirillo, Stefania Sica
IGARSS3