VLDB 2026 Research / reviewers in the wild / expert
Irina Rish
dblp:r/IrinaRish
· DBLP profile ↗
73ranked-venue papers
9as first author
35since 2021 · last 2026
0000-0001-6856-5057ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 55 · 7 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Computer networks · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorTheory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation HistoryabstractLarge language models require consistent behavioral patterns for safe deployment, yet there are indications of large variability that may lead to an instable expression of personality traits in these models. We present PERSIST (PERsonality Stability in Synthetic Text), a comprehensive evaluation framework testing 25 open-source models (1B-685B parameters) across 2 million+ responses. Using traditional (BFI, SD3) and novel LLM-adapted personality questionnaires, we systematically vary model size, personas, reasoning modes, question order or paraphrasing, and conversation history. Our findings challenge fundamental assumptions: (1) Question reordering alone can introduce large shifts in personality measurements; (2) Scaling provides limited stability gains: even 400B+ models exhibit standard deviations >0.3 on 5-point scales; (3) Interventions expected to stabilize behavior, such as reasoning and inclusion of conversation history, can paradoxically increase variability; (4) Detailed persona instructions produce mixed effects, with misaligned personas showing significantly higher variability than the helpful assistant baseline; (5) The LLM-adapted questionnaires, despite their improved ecological validity, exhibit instability comparable to human-centric versions. This persistent instability across scales and mitigation strategies suggests that current LLMs lack the architectural foundations for genuine behavioral consistency. For safety-critical applications requiring predictable behavior, these findings indicate that current alignment strategies may be inadequate. Tommaso Tosato, Saskia Helbling, Yorguin José Mantilla Ramos, Mahmood Hegazy, Alberto Tosato, David John Lemay, Irina Rish, Guillaume Dumas |
AAAI | 7 |
| 2026 | GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version IncompatibilitiesabstractThe rapid evolution of software libraries poses a considerable hurdle for code generation, necessitating continuous adaptation to frequent version updates while preserving backward compatibility. While existing code evolution benchmarks provide valuable insights, they typically lack execution-based evaluation for generating code compliant with specific library versions. To address this, we introduce GitChameleon 2.0, a novel, meticulously curated dataset comprising 328 Python code completion problems, each conditioned on specific library versions and accompanied by executable unit tests. GitChameleon 2.0 rigorously evaluates the capacity of contemporary large language models (LLMs), LLM-powered agents, code assistants, and RAG systems to perform version-conditioned code generation that demonstrates functional accuracy through execution. Our extensive evaluations indicate that state-of-the-art systems encounter significant challenges with this task; enterprise models achieving baseline success rates in the 48-51% range, underscoring the intricacy of the problem. By offering an execution-based benchmark emphasizing the dynamic nature of code libraries, GitChameleon 2.0 enables a clearer understanding of this challenge and helps guide the development of more adaptable and dependable AI code generation methods. Diganta Misra, Nizar Islah, Victor May, Brice Rauby, Justine Gehring, Antonio Orvieto, Muawiz Chaudhary, Eilif B. Muller, Irina Rish, Samira Ebrahimi Kahou, Massimo Caccia |
ACL (1) | 10 |
| 2025 | Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum LearningabstractAndrei Mircea, Supriyo Chakraborty, Nima Chitsazan, Irina Rish, Ekaterina Lobacheva. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Andrei Mircea, Supriyo Chakraborty, Nima Chitsazan, Irina Rish, Ekaterina Lobacheva |
ACL (1) | 4 |
| 2025 | Scaling Laws and Efficient Inference for Ternary Language ModelsabstractTejas Vaidhya, Ayush Kaushal, Vineet Jain, Francis Couture-Harpin, Prashant Shishodia, Majid Behbahani, Yuriy Nevmyvaka, Irina Rish. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tejas Vaidhya, Ayush Kaushal, Vineet Jain, Francis Couture Harpin, Prashant Shishodia, Majid Behbahani, Yuriy Nevmyvaka, Irina Rish |
ACL (1) | 8 |
| 2025 | CAVE : Detecting and Explaining Commonsense Anomalies in Visual EnvironmentsabstractCAVE: Commonsense Anomalies in Visual Environment 🏠 Project Page📄 Paper (EMNLP 2025)💻 Code Dataset Details Dataset Description CAVE is the first benchmark of real-world visual anomalies for evaluating Vision-Language Models (VLMs). It is curated from images captured in real-life settings (photographs and screenshots taken by individuals), sourced from Reddit. The benchmark is grounded in cognitive science literature on how humans detect and resolve anomalies. Each image is annotated with rich, multi-task annotations that support three open-ended tasks (anomaly description, explanation, and justification), one visual grounding task (anomaly localization via bounding boxes), and classification along four dimensions (anomaly category, severity, surprisal, and complexity) that characterize the anomaly. CAVE reveals that state-of-the-art VLMs struggle substantially with visual anomaly perception and commonsense reasoning: the best model (GPT-4o) achieves only ~57% F1-score on anomaly detection even with advanced prompting strategies. Curated by: Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut Affiliations: EPFL, MILA Language: English License: CC-BY-4.0 Published at: EMNLP 2025 (Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing) Dataset Sources Project Page: https://smontariol.github.io/cave-visual-anomalies/ Paper: CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments Contact: [email protected], [email protected] Uses Intended Uses CAVE is designed to evaluate VLMs on their ability to: Detect real-world commonsense anomalies in images (anomaly description). Explain why a detected situation is anomalous (anomaly explanation). Justify how an anomaly might have occurred (anomaly justification). Localize anomalies within images via bounding boxes (anomaly localization). Classify anomalies by their visual manifestation type and numerical features (severity, surprisal, complexity). It also serves as a resource for studying the alignment between human and machine processing of visual anomalies, and for developing improved prompting strategies or fine-tuning approaches for anomaly-related tasks. Out-of-Scope Uses CAVE is a benchmark for evaluation purposes. Its small size (361 images) makes it unsuitable as a training set. It should not be used to deploy anomaly detection systems in safety-critical settings without additional validation. Dataset Structure Overview CAVE consists of 361 images: 309 anomalous and 52 normal (non-anomalous) images. Anomalous images contain up to 3 anomalies each, totaling 334 annotated anomalies. Each anomaly is paired with a unique bounding box. Annotation Fields Each sample includes the following fields: Field Description image The image (photograph or screenshot) image_description Short description of the image content (without describing the anomaly) anomaly_description Textual description of what is anomalous in the image anomaly_explanation Explanation of why the situation is anomalous (commonsense reasoning) anomaly_justification Plausible explanation of how the anomaly might have occurred anomaly_category Category of the anomaly's visual manifestation (see taxonomy below) bounding_box Coordinates of the bounding box demarcating the anomalous region severity 1–5 score: does the anomaly require immediate action? surprisal 1–5 score: how much does the situation deviate from expectations? complexity 1–5 score: how hard is the anomaly to detect? Anomaly Category Taxonomy Anomalies are categorized by how they visually manifest, inspired by MMBench's taxonomy of visual reasoning types: Category Description Example Entity Presence An object is present when it shouldn't be A black bear in an industrial building Entity Absence An expected object is missing A person using a cutter without protective gear Entity Attribute An object has an anomalous attribute (color, shape, label, orientation, usage) A snack packet opened from the wrong side Spatial Relation An object is incorrectly positioned relative to another Furniture blocking an emergency button Uniformity Breach A disruption in an expected uniform/symmetrical pattern One tile with a different orientation Textual Anomaly Text in the image conveys an unexpected or contradictory message A "KEEP RIGHT" sign with an arrow pointing left Dataset Creation Images were collected from four Reddit subreddits that specialize in content featuring unusual or uncommon situations: r/ocdtriggers r/mildlyconfusing r/mildlyinfuriating r/OSHA The top 1,000 posts from each subreddit were downloaded using the PRAW library. Images were filtered through both automatic and manual processes to remove: Unclear or ambiguous content Non-realistic images NSFW or sensitive content Images with text annotations, circles, or other overlaid marks Images below icon resolution Annotation proceeded in two rounds, with Amazon Mechanical Turk followed by Expert Verification & Consolidation, with 3 independent raters per anomaly for severity, surprisal, and complexity scores. Citation @inproceedings{bhagwatkar-etal-2025-cave, title = "{CAVE} : Detecting and Explaining Commonsense Anomalies in Visual Environments", author = "Bhagwatkar, Rishika and Montariol, Syrielle and Romanou, Angelika and Borges, Beatriz and Rish, Irina and Bosselut, Antoine", booktitle = "Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing", month = nov, year = "2025", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2025.emnlp-main.1379/", doi = "10.18653/v1/2025.emnlp-main.1379", pages = "27110--27151", } Acknowledgements The authors acknowledge support from Canada CIFAR AI Chair Program, Canada Excellence Research Chairs Program, Swiss National Science Foundation (No. 215390), Innosuisse (PFFS-21-29), EPFL Center for Imaging, Sony Group Corporation, and a Meta LLM Evaluation Research Grant. Computational resources were provided by MILA - Quebec AI Institute. Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut |
EMNLP | 5 |
| 2025 | Handling Delay in Real-Time Reinforcement LearningabstractReal-time reinforcement learning (RL) introduces several challenges. First, policies are constrained to a fixed number of actions per second due to hardware limitations. Second, the environment may change while the network is still computing an action, leading to observational delay. The first issue can partly be addressed with pipelining, leading to higher throughput and potentially better policies. However, the second issue remains: if each neuron operates in parallel with an execution time of $\tau$, an $N$-layer feed-forward network experiences observation delay of $\tau N$.
Reducing the number of layers can decrease this delay, but at the cost of the network's expressivity. In this work, we explore the trade-off between minimizing delay and network's expressivity. We present a theoretically motivated solution that leverages temporal skip connections combined with history-augmented observations. We evaluate several architectures and show that those incorporating temporal skip connections achieve strong performance across various neuron execution times, reinforcement learning algorithms, and environments, including four Mujoco tasks and all MinAtar games. Moreover, we demonstrate parallel neuron computation can accelerate inference by 6-350\% on standard hardware. Our investigation into temporal skip connections and parallel computations paves the way for more efficient RL agents in real-time setting. Ivan Anokhin, Rishav Rishav, Matthew Riemer, Stephen Chung, Irina Rish, Samira Ebrahimi Kahou |
ICLR | 5 |
| 2025 | Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced ReasoningabstractDecoder-only Transformers often struggle with complex reasoning tasks, particularly arithmetic reasoning requiring multiple sequential operations. In this work, we identify representation collapse in the model’s intermediate layers as a key factor limiting their reasoning capabilities. To address this, we propose Sequential Variance-Covariance Regularization (Seq-VCR), which enhances the entropy of intermediate representations and prevents collapse. Combined with dummy pause tokens as substitutes for chain-of-thought (CoT) tokens, our method significantly improves performance in arithmetic reasoning problems. In the challenging 5 × 5 integer multiplication task, our approach achieves 99.5% exact match accuracy, outperforming models of the same size (which yield 0% accuracy) and GPT-4 with five-shot CoT prompting (44%). We also demonstrate superior results on arithmetic expression and longest increasing subsequence (LIS) datasets. Our findings highlight the importance of preventing intermediate layer representation collapse to enhance the reasoning capabilities of Transformers and show that Seq-VCR offers an effective solution without requiring explicit CoT supervision. Md Rifat Arefin, Gopeshh Subbaraj, Nicolas Angelard-Gontier, Yann LeCun, Irina Rish, Ravid Shwartz-Ziv, Christopher Joseph Pal |
ICLR | 5 |
| 2025 | Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingabstractIn inverse reinforcement learning (IRL), an agent seeks to replicate expert demonstrations through interactions with the environment.
Traditionally, IRL is treated as an adversarial game, where an adversary searches over reward models, and a learner optimizes the reward through repeated RL procedures.
This game-solving approach is both computationally expensive and difficult to stabilize.
In this work, we propose a novel approach to IRL by _direct policy search_:
by exploiting a linear factorization of the return as the inner product of successor features and a reward vector, we design an IRL algorithm by policy gradient descent on the gap between the learner and expert features.
Our non-adversarial method does not require learning an explicit reward function and can be solved seamlessly with existing RL algorithms.
Remarkably, our approach works in state-only settings without expert action labels, a setting which behavior cloning (BC) cannot solve.
Empirical results demonstrate that our method learns from as few as a single expert demonstration and achieves improved performance on various control tasks. Arnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish, Glen Berseth, Sanjiban Choudhury |
ICLR | 4 |
| 2025 | Surprising Effectiveness of pretraining Ternary Language Model at ScaleabstractRapid advancements in GPU computational power has outpaced memory capacity and bandwidth growth, creating bottlenecks in Large Language Model (LLM) inference. Post-training quantization is the leading method for addressing memory-related bottlenecks in LLM inference, but it suffers from significant performance degradation below 4-bit precision. This paper addresses these challenges by investigating the pretraining of low-bitwidth models specifically Ternary Language Models (TriLMs) as an alternative to traditional floating-point models (FloatLMs) and their post-training quantized versions (QuantLMs). We present Spectra LLM suite, the first open suite of LLMs spanning multiple bit-widths, including FloatLMs, QuantLMs, and TriLMs, ranging from 99M to 3.9B parameters trained on 300B tokens. Our comprehensive evaluation demonstrates that TriLMs offer superior scaling behavior in terms of model size (in bits). Surprisingly, at scales exceeding one billion parameters, TriLMs consistently outperform their QuantLM and FloatLM counterparts for a given bit size across various benchmarks. Notably, the 3.9B parameter TriLM matches the performance of the FloatLM 3.9B across all benchmarks, despite having fewer bits than FloatLM 830M. Overall, this research provides valuable insights into the feasibility and scalability of low-bitwidth language models, paving the way for the development of more efficient LLMs. Ayush Kaushal, Tejas Vaidhya, Arnab Kumar Mondal, Tejas Pandey, Aaryan Bhagat, Irina Rish |
ICLR | 6 |
| 2025 | Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous InferenceabstractRealtime environments change even as agents perform action inference and learning, thus requiring high interaction frequencies to effectively minimize regret. However, recent advances in machine learning involve larger neural networks with longer inference times, raising questions about their applicability in realtime systems where reaction time is crucial. We present an analysis of lower bounds on regret in realtime reinforcement learning (RL) environments to show that minimizing long-term regret is generally impossible within the typical sequential interaction and learning paradigm, but often becomes possible when sufficient asynchronous compute is available. We propose novel algorithms for staggering asynchronous inference processes to ensure that actions are taken at consistent time intervals, and demonstrate that use of models with high action inference times is only constrained by the environment's effective stochasticity over the inference horizon, and not by action frequency. Our analysis shows that the number of inference processes needed scales linearly with increasing inference times while enabling use of models that are multiple orders of magnitude larger than existing approaches when learning from a realtime simulation of Game Boy games such as Pokemon and Tetris. Matthew Riemer, Gopeshh Subbaraj, Glen Berseth, Irina Rish |
ICLR | 4 |
| 2025 | Context is Key: A Benchmark for Forecasting with Essential Textual InformationabstractForecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable and accurate predictions. Human forecasters frequently rely on additional information, such as background knowledge and constraints, which can efficiently be communicated through natural language. However, in spite of recent progress with LLM-based forecasters, their ability to effectively integrate this textual information remains an open question. To address this, we introduce "Context is Key" (CiK), a time-series forecasting benchmark that pairs numerical data with diverse types of carefully crafted textual context, requiring models to integrate both modalities; crucially, every task in CiK requires understanding textual context to be solved successfully. We evaluate a range of approaches, including statistical models, time series foundation models, and LLM-based forecasters, and propose a simple yet effective LLM prompting method that outperforms all other tested methods on our benchmark. Our experiments highlight the importance of incorporating contextual information, demonstrate surprising performance when using LLM-based forecasting models, and also reveal some of their critical shortcomings. This benchmark aims to advance multimodal forecasting by promoting models that are both accurate and accessible to decision-makers with varied technical expertise. The benchmark can be visualized at https://servicenow.github.io/context-is-key-forecasting/v0. Andrew Robert Williams, Arjun Ashok, Étienne Marcotte, Valentina Zantedeschi, Jithendaraa Subramanian, Roland Riachi, James Requeima, Alexandre Lacoste, Irina Rish, Nicolas Chapados, Alexandre Drouin |
ICML | 9 |
| 2025 | AI for Global Climate Cooperation: Modeling Global Climate Negotiations, Agreements, and Long-Term Cooperation in RICE-NabstractGlobal cooperation on climate change mitigation is essential to limit temperature increases while supporting long-term, equitable economic growth and sustainable development. Achieving such cooperation among diverse regions, each with different incentives, in a dynamic environment shaped by complex geopolitical and economic factors, without a central authority, is a profoundly challenging game-theoretic problem. This article introduces RICE-N, a multi-region integrated assessment model that simulates the global climate, economy, and climate negotiations and agreements. RICE-N uses multi-agent reinforcement learning (MARL) to encourage agents to develop strategic behaviors based on the environmental dynamics and the actions of the others. We present two negotiation protocols: (1) Bilateral Negotiation, an exemplary protocol and (2) Basic Club, inspired from Climate Clubs and the carbon border adjustment mechanism (Nordhaus, 2015; Comissions, 2022). We compare their impact against a no-negotiation baseline with various mitigation strategies, showing that both protocols significantly reduce temperature growth at the cost of a minor drop in production while ensuring a more equitable distribution of the emission reduction costs. Andrew Robert Williams, Phillip Wozny, Kai-Hendrik Cohrs, Koen Ponse, Marco Jiralerspong, Soham R. Phade, Sunil Srinivasa, Prateek Gupta, Erman Acar, Irina Rish, Yoshua Bengio, Stephan Zheng |
ICML | 13 |
| 2024 | Decision-Making Paradoxes in Humans vs Machines: The case of the Allais and Ellsberg Paradoxes
Ardavan Salehi Nobandegani, Irina Rish, Thomas R. Shultz |
CogSci | 2 |
| 2024 | Unsupervised Concept Discovery Mitigates Spurious CorrelationsabstractModels prone to spurious correlations in training data often produce brittle predictions and introduce unintended biases. Addressing this challenge typically involves methods relying on prior knowledge and group annotation to remove spurious correlations, which may not be readily available in many applications. In this paper, we establish a novel connection between unsupervised object-centric learning and mitigation of spurious correlations. Instead of directly inferring subgroups with varying correlations with labels, our approach focuses on discovering concepts: discrete ideas that are shared across input samples. Leveraging existing object-centric representation learning, we introduce CoBalT: a concept balancing technique that effectively mitigates spurious correlations without requiring human labeling of subgroups. Evaluation across the benchmark datasets for sub-population shifts demonstrate superior or competitive performance compared state-of-the-art baselines, without the need for group annotation. Code is available at https://github.com/rarefin/CoBalT Md Rifat Arefin, Aristide Baratin, Francesco Locatello, Irina Rish, Dianbo Liu, Kenji Kawaguchi |
ICML | 5 |
| 2024 | Knowledge Distillation in Federated Learning: A Practical Guide
Alessio Mora, Irene Tenison, Paolo Bellavista, Irina Rish |
IJCAI | 4 |
| 2024 | Using Unity to Help Solve Reinforcement LearningabstractLeveraging the depth and flexibility of XLand as well as the rapid prototyping features of the Unity engine, we present the United Unity Universe — an open-source toolkit designed to accelerate the creation of innovative reinforcement learning environments. This toolkit includes a robust implementation of XLand 2.0 complemented by a user-friendly interface which allows users to modify the details of procedurally generated terrains and task rules with ease. Additionally, we provide a curated selection of terrains and rule sets, accompanied by implementations of reinforcement learning baselines to facilitate quick experimentation with novel architectural designs for adaptive agents. Furthermore, we illustrate how the United Unity Universe serves as a high-level language that enables researchers to develop diverse and endlessly variable 3D environments within a unified framework. This functionality establishes the United Unity Universe (U3) as an essential tool for advancing the field of reinforcement learning, especially in the development of adaptive and generalizable learning systems. Connor Brennan, Andrew Robert Williams, Omar G. Younis, Vedant Vyas, Daria Yasafova, Irina Rish |
NeurIPS | 6 |
| 2024 | RedPajama: an Open Dataset for Training Large Language ModelsabstractLarge language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset composition and filtering remain largely elusive. Many of the top-performing models lack transparency in their dataset curation and model development processes, posing an obstacle to the development of fully open language models. In this paper, we identify three core data-related challenges that must be addressed to advance open-source language models. These include (1) transparency in model development, including the data curation process, (2) access to large quantities of high-quality data, and (3) availability of artifacts and metadata for dataset curation and analysis. To address these challenges, we release RedPajama-V1, an open reproduction of the LLaMA training dataset. In addition, we release RedPajama-V2, a massive web-only dataset consisting of raw, unfiltered text data together with quality signals and metadata.Together, the RedPajama datasets comprise over 100 trillion tokens spanning multiple domains and with their quality signals facilitate the filtering of data, aiming to inspire the development of numerous new datasets. To date, these datasets have already been used in the training of strong language models used in production, such as Snowflake Arctic, Salesforce's XGen and AI2's OLMo. To provide insight into the quality of RedPajama, we present a series of analyses and ablation studies with decoder-only language models with up to 1.6B parameters. Our findings demonstrate how quality signals for web data can be effectively leveraged to curate high-quality subsets of the dataset, underscoring the potential of RedPajama to advance the development of transparent and high-performing language models at scale. Maurice Weber, Daniel Y. Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, Ce Zhang 0001 |
NeurIPS | 18 |
| 2023 | AI Agents Learn to Trust
Ardavan Salehi Nobandegani, Irina Rish, Thomas R. Shultz |
CogSci | 2 |
| 2023 | Dialogue System with Missing ObservationabstractWithin the domain of dialogue, the ability to orchestrate multiple independently trained dialogue agents to create a unified system is of particular importance. Where we define orchestration as the task of selecting a subset of skills which most appropriately answer a user input using features extracted from both the user input and the individual skills. In this work, we study the task of online dialogue orchestration where the user feedback associated with the dialogue agent may not always be observed. In order to address the missing feedback setting, we propose to combine the attentive contextual bandit approach with an unsupervised learning mechanism such as clustering. By leveraging clustering to estimate missing reward, we are able to learn from each incoming event, even those with missing rewards. Promising empirical results are obtained on proprietary conversational datasets. Djallel Bouneffouf 0001, Mayank Agarwal, Irina Rish |
ICASSP | 3 |
| 2023 | Broken Neural Scaling Laws
Ethan Caballero, Irina Rish, David Krueger 0001 |
ICLR | 3 |
| 2023 | Maximum State Entropy Exploration using Predecessor and Successor RepresentationsabstractAnimals have a developed ability to explore that aids them in important tasks such as locating food, exploring for shelter, and finding misplaced items. These exploration skills necessarily track where they have been so that they can plan for finding items with relative efficiency. Contemporary exploration algorithms often learn a less efficient exploration strategy because they either condition only on the current state or simply rely on making random open-loop exploratory moves. In this work, we propose $\eta\psi$-Learning, a method to learn efficient exploratory policies by conditioning on past episodic experience to make the next exploratory move. Specifically, $\eta\psi$-Learning learns an exploration policy that maximizes the entropy of the state visitation distribution of a single trajectory. Furthermore, we demonstrate how variants of the predecessor representation and successor representations can be combined to predict the state visitation entropy. Our experiments demonstrate the efficacy of $\eta\psi$-Learning to strategically explore the environment and maximize the state coverage with limited samples. Arnav Kumar Jain, Lucas Lehnert, Irina Rish, Glen Berseth |
NeurIPS | 3 |
| 2022 | Cognitive Models as Simulators: The Case of Moral Decision-Making
Ardavan Salehi Nobandegani, Thomas R. Shultz, Irina Rish |
CogSci | 3 |
| 2022 | Parametric Scattering NetworksabstractThe wavelet scattering transform creates geometric in-variants and deformation stability. In multiple signal do-mains, it has been shown to yield more discriminative rep-resentations compared to other non-learned representations and to outperform learned representations in certain tasks, particularly on limited labeled data and highly structured signals. The wavelet filters used in the scattering trans-form are typically selected to create a tight frame via a pa-rameterized mother wavelet. In this work, we investigate whether this standard wavelet filterbank construction is op-timal. Focusing on Morlet wavelets, we propose to learn the scales, orientations, and aspect ratios of the filters to produce problem-specific parameterizations of the scattering transform. We show that our learned versions of the scattering transform yield significant performance gains in small-sample classification settings over the standard scat-tering transform. Moreover, our empirical results suggest that traditional filterbank constructions may not always be necessary for scattering transforms to extract effective rep-resentations. Shanel Gauthier, Benjamin Thérien, Laurent Alsène-Racicot, Muawiz Chaudhary, Irina Rish, Eugene Belilovsky, Michael Eickenberg, Guy Wolf |
CVPR | 5 |
| 2022 | A Remedy For Distributional Shifts Through Expected Domain TranslationabstractMachine learning models often fail to generalize to unseen domains due to the distributional shifts. A family of such shifts, “correlation shifts,” is caused by spurious correlations in the data. It is studied under the overarching topic of “domain generalization.” In this work, we employ multi-modal translation networks to tackle the correlation shifts that appear when data is sampled out-of-distribution. Learning a generative model from training domains enables us to translate each training sample under the special characteristics of other possible domains. We show that by training a predictor solely on the generated samples, the spurious correlations in training domains average out, and the invariant features corresponding to true correlations emerge. Our proposed technique, Expected Domain Translation (EDT), is benchmarked on the Colored MNIST dataset and drastically improves the state-of-the-art classification accuracy by 38% with train-domain validation model selection. Jean-Christophe Gagnon-Audet, Soroosh Shahtalebi, Frank Rudzicz, Irina Rish |
ICASSP | 4 |
| 2022 | Compositional Attention: Disentangling Search and Retrieval
Sarthak Mittal, Sharath Chandra, Irina Rish, Yoshua Bengio, Guillaume Lajoie |
ICLR | 3 |
| 2022 | Towards Scaling Difference Target Propagation by Learning Backprop TargetsabstractThe development of biologically-plausible learning algorithms is important for understanding learning in the brain, but most of them fail to scale-up to real-world tasks, limiting their potential as explanations for learning by real brains. As such, it is important to explore learning algorithms that come with strong theoretical guarantees and can match the performance of backpropagation (BP) on complex tasks. One such algorithm is Difference Target Propagation (DTP), a biologically-plausible learning algorithm whose close relation with Gauss-Newton (GN) optimization has been recently established. However, the conditions under which this connection rigorously holds preclude layer-wise training of the feedback pathway synaptic weights (which is more biologically plausible). Moreover, good alignment between DTP weight updates and loss gradients is only loosely guaranteed and under very specific conditions for the architecture being trained. In this paper, we propose a novel feedback weight training scheme that ensures both that DTP approximates BP and that layer-wise feedback weight training can be restored without sacrificing any theoretical guarantees. Our theory is corroborated by experimental results and we report the best performance ever achieved by DTP on CIFAR-10 and ImageNet 32x32. Maxence Ernoult, Fabrice Normandin, Abhinav Moudgil, Sean Spinney, Eugene Belilovsky, Irina Rish, Blake A. Richards, Yoshua Bengio |
ICML | 6 |
| 2022 | Continual Learning In Environments With Polynomial Mixing TimesabstractThe mixing time of the Markov chain induced by a policy limits performance in real-world continual learning scenarios. Yet, the effect of mixing times on learning in continual reinforcement learning (RL) remains underexplored. In this paper, we characterize problems that are of long-term interest to the development of continual RL, which we call scalable MDPs, through the lens of mixing times. In particular, we theoretically establish that scalable MDPs have mixing times that scale polynomially with the size of the problem. We go on to demonstrate that polynomial mixing times present significant difficulties for existing approaches that suffer from myopic bias and stale bootstrapped estimates. To validate the proposed theory, we study the empirical scaling behavior of mixing times with respect to the number of tasks and task switching frequency for pretrained high performing policies on seven Atari games. Our analysis demonstrates both that polynomial mixing times do emerge in practice and how their existence may lead to unstable learning behavior like catastrophic forgetting in continual learning settings. Matthew Riemer, Sharath Chandra, Ignacio Cases, Gopeshh Subbaraj, Maximilian Puelma Touzel, Irina Rish |
NeurIPS | 6 |
| 2022 | Towards Continual Reinforcement Learning: A Review and PerspectivesabstractIn this article, we aim to provide a literature review of different formulations and approaches to continual reinforcement learning (RL), also known as lifelong or non-stationary RL. We begin by discussing our perspective on why RL is a natural fit for studying continual learning. We then provide a taxonomy of different continual RL formulations by mathematically characterizing two key properties of non-stationarity, namely, the scope and driver non-stationarity. This offers a unified view of various formulations. Next, we review and present a taxonomy of continual RL approaches. We go on to discuss evaluation of continual RL agents, providing an overview of benchmarks used in the literature and important metrics for understanding agent performance. Finally, we highlight open problems and challenges in bridging the gap between the current state of continual RL and findings in neuroscience. While still in its early days, the study of continual RL has the promise to develop better incremental reinforcement learners that can function in increasingly realistic applications where non-stationarity plays a vital role. These include applications such as those in the fields of healthcare, education, logistics, and robotics. Khimya Khetarpal, Matthew Riemer, Irina Rish, Doina Precup |
J. Artif. Intell. Res. | 3 |
| 2021 | Toward Skills Dialog Orchestration with Online LearningabstractBuilding multi-domain AI agents is a challenging task and an open problem in the area of AI. Within the domain of dialog, the ability to orchestrate multiple independently trained dialog agents, or skills, to create a unified system is of particular significance. In this work, we study the task of online posterior dialog orchestration, where we define posterior orchestration as the task of selecting a subset of skills which most appropriately answer a user input using features extracted from both the user input and the individual skills. To account for the various costs associated with extracting skill features, we consider online posterior orchestration under a skill execution budget. We formalize this setting as Context Attentive Bandit with Observations (CABO), a variant of context attentive bandits, and evaluate it on proprietary conversational datasets. Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Mayank Agarwal, Yasaman Khazaeni, Irina Rish |
ICASSP | 6 |
| 2021 | Double-Linear Thompson Sampling for Context-Attentive BanditsabstractIn this paper, we analyze and extend an online learning frame-work known as Context-Attentive Bandit, motivated by various practical applications, from medical diagnosis to dialog systems, where due to observation costs only a small subset of a potentially large number of context variables can be observed at each iteration; however, the agent has a freedom to choose which variables to observe. We derive a novel algorithm, called Context-Attentive Thompson Sampling (CATS), which builds upon the Linear Thompson Sampling approach, adapting it to Context-Attentive Bandit setting. We provide a theoretical regret analysis and an extensive empirical evaluation demonstrating advantages of the proposed approach over several baseline methods on a variety of real-life datasets. Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Yasaman Khazaeni, Irina Rish |
ICASSP | 5 |
| 2021 | Predicting Infectiousness for Proactive Contact Tracing
Yoshua Bengio, Prateek Gupta, Tegan Maharaj, Nasim Rahaman, Martin Weiss, Tristan Deleu, Eilif B. Muller, Meng Qu, Victor Schmidt, Pierre-Luc St-Charles, Hannah Alsdurf, Olexa Bilaniuk, David L. Buckeridge, Gaétan Marceau-Caron, Pierre Luc Carrier, Joumana Ghosn, Satya Ortiz-Gagne, Christopher Joseph Pal, Irina Rish, Bernhard Schölkopf, Jian Tang 0005, Andrew Robert Williams |
ICLR | 19 |
| 2021 | Toward Optimal Solution for the Context-Attentive Bandit ProblemabstractIn various recommender system applications, from medical diagnosis to dialog systems, due to observation costs only a small subset of a potentially large number of context variables can be observed at each iteration; however, the agent has a freedom to choose which variables to observe. In this paper, we analyze and extend an online learning framework known as Context-Attentive Bandit, We derive a novel algorithm, called Context-Attentive Thompson Sampling (CATS), which builds upon the Linear Thompson Sampling approach, adapting it to Context-Attentive Bandit setting. We provide a theoretical regret analysis and an extensive empirical evaluation demonstrating advantages of the proposed approach over several baseline methods on a variety of real-life datasets. Djallel Bouneffouf 0001, Raphaël Féraud, Sohini Upadhyay, Irina Rish, Yasaman Khazaeni |
IJCAI | 4 |
| 2021 | Invariance Principle Meets Information Bottleneck for Out-of-Distribution GeneralizationabstractThe invariance principle from causality is at the heart of notable approaches such as invariant risk minimization (IRM) that seek to address out-of-distribution (OOD) generalization failures. Despite the promising theory, invariance principle-based approaches fail in common classification tasks, where invariant (causal) features capture all the information about the label. Are these failures due to the methods failing to capture the invariance? Or is the invariance principle itself insufficient? To answer these questions, we revisit the fundamental assumptions in linear regression tasks, where invariance-based approaches were shown to provably generalize OOD. In contrast to the linear regression tasks, we show that for linear classification tasks we need much stronger restrictions on the distribution shifts, or otherwise OOD generalization is impossible. Furthermore, even with appropriate restrictions on distribution shifts in place, we show that the invariance principle alone is insufficient. We prove that a form of the information bottleneck constraint along with invariance helps address the key failures when invariant features capture all the information about the label and also retains the existing success when they do not. We propose an approach that incorporates both of these principles and demonstrate its effectiveness in several experiments. Kartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet, Yoshua Bengio, Ioannis Mitliagkas, Irina Rish |
NeurIPS | 7 |
| 2021 | Adversarial Feature DesensitizationabstractNeural networks are known to be vulnerable to adversarial attacks -- slight but carefully constructed perturbations of the inputs which can drastically impair the network's performance. Many defense methods have been proposed for improving robustness of deep networks by training them on adversarially perturbed inputs. However, these models often remain vulnerable to new types of attacks not seen during training, and even to slightly stronger versions of previously seen attacks. In this work, we propose a novel approach to adversarial robustness, which builds upon the insights from the domain adaptation field. Our method, called Adversarial Feature Desensitization (AFD), aims at learning features that are invariant towards adversarial perturbations of the inputs. This is achieved through a game where we learn features that are both predictive and robust (insensitive to adversarial attacks), i.e. cannot be used to discriminate between natural and adversarial data. Empirical results on several benchmarks demonstrate the effectiveness of the proposed approach against a wide range of attack types and attack strengths. Our code is available at https://github.com/BashivanLab/afd. Pouya Bashivan, Reza Bayat, Adam Ibrahim, Kartik Ahuja, Mojtaba Faramarzi, Touraj Laleh, Blake A. Richards, Irina Rish |
NeurIPS | 8 |
| 2021 | Learning Brain Dynamics With Coupled Low-Dimensional Nonlinear Oscillators and Deep Recurrent NetworksabstractMany natural systems, especially biological ones, exhibit complex multivariate nonlinear dynamical behaviors that can be hard to capture by linear autoregressive models. On the other hand, generic nonlinear models such as deep recurrent neural networks often require large amounts of training data, not always available in domains such as brain imaging; also, they often lack interpretability. Domain knowledge about the types of dynamics typically observed in such systems, such as a certain type of dynamical systems models, could complement purely data-driven techniques by providing a good prior. In this work, we consider a class of ordinary differential equation (ODE) models known as van der Pol (VDP) oscil lators and evaluate their ability to capture a low-dimensional representation of neural activity measured by different brain imaging modalities, such as calcium imaging (CaI) and fMRI, in different living organisms: larval zebrafish, rat, and human. We develop a novel and efficient approach to the nontrivial problem of parameters estimation for a network of coupled dynamical systems from multivariate data and demonstrate that the resulting VDP models are both accurate and interpretable, as VDP's coupling matrix reveals anatomically meaningful excitatory and inhibitory interactions across different brain subsystems. VDP outperforms linear autoregressive models (VAR) in terms of both the data fit accuracy and the quality of insight provided by the coupling matrices and often tends to generalize better to unseen data when predicting future brain activity, being comparable to and sometimes better than the recurrent neural networks (LSTMs). Finally, we demonstrate that our (generative) VDP model can also serve as a data-augmentation tool leading to marked improvements in predictive accuracy of recurrent neural networks. Thus, our work contributes to both basic and applied dimensions of neuroimaging: gaining scientific insights and improving brain-based predictive models, an area of potentially high practical importance in clinical diagnosis and neurotechnology. Germán Abrevaya, Guillaume Dumas, Aleksandr Y. Aravkin, Peng Zheng 0002, Jean-Christophe Gagnon-Audet, James R. Kozloski, Pablo Polosecki, Guillaume Lajoie, David D. Cox, Silvina Ponce Dawson, Guillermo A. Cecchi, Irina Rish |
Neural Comput. | 12 |
| 2020 | Modeling Dialogues with Hashcode Representations: A Nonparametric ApproachabstractWe propose a novel dialogue modeling framework, the first-ever nonparametric kernel functions based approach for dialogue modeling, which learns hashcodes as text representations; unlike traditional deep learning models, it handles well relatively small datasets, while also scaling to large ones. We also derive a novel lower bound on mutual information, used as a model-selection criterion favoring representations with better alignment between the utterances of participants in a collaborative dialogue setting, as well as higher predictability of the generated responses. As demonstrated on three real-life datasets, including prominently psychotherapy sessions, the proposed approach significantly outperforms several state-of-art neural network based dialogue systems, both in terms of computational efficiency, reducing training time from days or weeks to hours, and the response quality, achieving an order of magnitude improvement over competitors in frequency of being chosen as the best model by human evaluators. Sahil Garg, Irina Rish, Guillermo A. Cecchi, Palash Goyal, Sarik Ghazarian, Shuyang Gao, Greg Ver Steeg, Aram Galstyan |
AAAI | 2 |
| 2020 | Survey on Applications of Multi-Armed and Contextual BanditsabstractIn recent years, the multi-armed bandit (MAB) framework has attracted a lot of attention in various applications, from recommender systems and information retrieval to healthcare and finance. This success is due to its stellar performance combined with attractive properties, such as learning from less feedback. The multiarmed bandit field is currently experiencing a renaissance, as novel problem settings and algorithms motivated by various practical applications are being introduced, building on top of the classical bandit problem. This article aims to provide a comprehensive review of top recent developments in multiple real-life applications of the multi-armed bandit. Specifically, we introduce a taxonomy of common MAB-based applications and summarize the state-of-the-art for each of those domains. Furthermore, we identify important current trends and provide new perspectives pertaining to the future of this burgeoning field. Djallel Bouneffouf 0001, Irina Rish, Charu C. Aggarwal |
CEC | 2 |
| 2020 | Online Fast Adaptation and Knowledge Accumulation (OSAKA): a New Approach to Continual LearningabstractContinual learning agents experience a stream of (related) tasks. The main challenge is that the agent must not forget previous tasks and also adapt to novel tasks in the stream. We are interested in the intersection of two recent continual-learning scenarios. In meta-continual learning, the model is pre-trained using meta-learning to minimize catastrophic forgetting of previous tasks. In continual-meta learning, the aim is to train agents for faster remembering of previous tasks through adaptation. In their original formulations, both methods have limitations. We stand on their shoulders to propose a more general scenario, OSAKA, where an agent must quickly solve new (out-of-distribution) tasks, while also requiring fast remembering. We show that current continual learning, meta-learning, meta-continual learning, and continual-meta learning techniques fail in this new scenario. We propose Continual-MAML, an online extension of the popular MAML algorithm as a strong baseline for this scenario. We show in an empirical study that Continual-MAML is better suited to the new scenario than the aforementioned methodologies including standard continual learning and meta-learning approaches. Massimo Caccia, Pau Rodríguez, Oleksiy Ostapenko, Fabrice Normandin, Lucas Caccia, Issam H. Laradji, Irina Rish, Alexandre Lacoste, David Vázquez 0001, Laurent Charlin |
NeurIPS | 8 |
| 2019 | Kernelized Hashcode Representations for Relation ExtractionabstractKernel methods have produced state-of-the-art results for a number of NLP tasks such as relation extraction, but suffer from poor scalability due to the high cost of computing kernel similarities between natural language structures. A recently proposed technique, kernelized locality-sensitive hashing (KLSH), can significantly reduce the computational cost, but is only applicable to classifiers operating on kNN graphs. Here we propose to use random subspaces of KLSH codes for efficiently constructing an explicit representation of NLP structures suitable for general classification methods. Further, we propose an approach for optimizing the KLSH model for classification problems by maximizing an approximation of mutual information between the KLSH codes (feature vectors) and the class labels. We evaluate the proposed approach on biomedical relation extraction datasets, and observe significant and robust improvements in accuracy w.r.t. state-ofthe-art classifiers, along with drastic (orders-of-magnitude) speedup compared to conventional kernel methods. Sahil Garg, Aram Galstyan, Greg Ver Steeg, Irina Rish, Guillermo A. Cecchi, Shuyang Gao |
AAAI | 4 |
| 2019 | Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference
Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu 0001, Irina Rish, Yuhai Tu, Gerald Tesauro |
ICLR (Poster) | 5 |
| 2019 | Beyond Backprop: Online Alternating Minimization with Auxiliary VariablesabstractDespite significant recent advances in deep neural networks, training them remains a challenge due to the highly non-convex nature of the objective function. State-of-the-art methods rely on error backpropagation, which suffers from several well-known issues, such as vanishing and exploding gradients, inability to handle non-differentiable nonlinearities and to parallelize weight-updates across layers, and biological implausibility. These limitations continue to motivate exploration of alternative training algorithms, including several recently proposed auxiliary-variable methods which break the complex nested objective function into local subproblems. However, those techniques are mainly offline (batch), which limits their applicability to extremely large datasets, as well as to online, continual or reinforcement learning. The main contribution of our work is a novel online (stochastic/mini-batch) alternating minimization (AM) approach for training deep neural networks, together with the first theoretical convergence guarantees for AM in stochastic settings and promising empirical results on a variety of architectures and datasets. Anna Choromanska, Benjamin Cowen, Sadhana Kumaravel, Ronny Luss, Mattia Rigotti, Irina Rish, Paolo Diachille, Viatcheslav Gurev, Brian Kingsbury, Ravi Tejwani, Djallel Bouneffouf 0001 |
ICML | 6 |
| 2017 | Context Attentive Bandits: Contextual Bandit with Restricted ContextabstractWe consider a novel formulation of the multi-armed bandit model, which we call the contextual bandit with restricted context, where only a limited number of features can be accessed by the learner at every iteration. This novel formulation is motivated by different online problems arising in clinical trials, recommender systems and attention modeling.Herein, we adapt the standard multi-armed bandit algorithm known as Thompson Sampling to take advantage of our restricted context setting, and propose two novel algorithms, called the Thompson Sampling with Restricted Context (TSRC) and the Windows Thompson Sampling with Restricted Context (WTSRC), for handling stationary and nonstationary environments, respectively. Our empirical results demonstrate advantages of the proposed approaches on several real-life datasets. Djallel Bouneffouf 0001, Irina Rish, Guillermo A. Cecchi, Raphaël Féraud |
IJCAI | 2 |
| 2017 | Neurogenesis-Inspired Dictionary Learning: Online Model Adaption in a Changing WorldabstractWe address the problem of online model adaptation when learning representations from non-stationary data streams. Specifically, we focus here on online dictionary learning (i.e. sparse linear autoencoder), and propose a simple but effective online model selection approach involving “birth” (addition) and “death” (removal) of hidden units representing dictionary elements, in response to changing inputs; we draw inspiration from the adult neurogenesis phenomenon in the dentate gyrus of the hippocampus, known to be associated with better adaptation to new environments. Empirical evaluation on real-life datasets (images and text), as well as on synthetic data, demonstrates that the proposed approach can considerably outperform the state-of-art non-adaptive online sparse coding of [Mairal et al., 2009] in the presence of non-stationary data. Moreover, we identify certain data- and model properties associated with such improvements. Sahil Garg, Irina Rish, Guillermo A. Cecchi, Aurélie C. Lozano |
IJCAI | 2 |
| 2016 | MINT: Mutual Information Based Transductive Feature Selection for Genetic Trait PredictionabstractWhole genome prediction of complex phenotypic traits using high-density genotyping arrays has attracted a lot of attention, as it is relevant to the fields of plant and animal breeding and genetic epidemiology. Since the number of genotypes is generally much bigger than the number of samples, predictive models suffer from the curse of dimensionality. The curse of dimensionality problem not only affects the computational efficiency of a particular genomic selection method, but can also lead to a poor performance, mainly due to possible overfitting, or un-informative features. In this work, we propose a novel transductive feature selection method, called MINT, which is based on the MRMR (Max-Relevance and Min-Redundancy) criterion. We apply MINT on genetic trait prediction problems and show that, in general, MINT is a better feature selection method than the state-of-the-art inductive method MRMR. Dan He 0001, Irina Rish, David Haws, Laxmi Parida |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2014 | Transductive HSIC LassoabstractSparse regression methods such as l1-regularized linear regression, or Lasso [18], are commonly used for analysis of high-dimensional, small-sample datasets, due to their good generalization and feature-selection properties. However, predictive accuracy of sparse regression can be further improved by incorporating more realistic data-modeling assumptions (e.g., non-linearity) and by fully exploiting all available data, as suggested by the transductive approach [6], which makes instance-specific predictions based on both labeled (training) data and unlabeled (test) data, instead of learning a single fixed model from training data. Based on these ideas, we develop a novel method, called Transductive HSIC Lasso, that incorporates transduction into a nonlinear sparse regression approach known as HSIC Lasso [19]. Unlike the existing transductive Lasso algorithm of [1], our approach does not rely on imputation, i.e., on estimation of unknown labels using a predictor built on training data; the latter may sometimes result in poor overall performance due to unreliable label estimates. Instead, our method exploits the structure of the HSIC Lasso, which maximizes the relevance between the selected features and the label, while minimizing the redundancy between the selected features; transduction is achieved by including unlabeled samples into the redundancy computation. Our experiments demonstrate advantages of the proposed method over the state-of-the-art approaches, both on simulated and real-life data, such as prediction of phenotypic traits from genomic data, and prediction of subject's pain level from his/her functional MRI data. Dan He 0001, Irina Rish, Laxmi Parida |
SDM | 2 |
| 2013 | Functional MRI Analysis with Sparse Models
Irina Rish |
ECML/PKDD (3) | 1 |
| 2012 | Predictive Dynamics of Human Pain PerceptionabstractWhile the static magnitude of thermal pain perception has been shown to follow a power-law function of the temperature, its dynamical features have been largely overlooked. Due to the slow temporal experience of pain, multiple studies now show that the time evolution of its magnitude can be captured with continuous online ratings. Here we use such ratings to model quantitatively the temporal dynamics of thermal pain perception. We show that a differential equation captures the details of the temporal evolution in pain ratings in individual subjects for different stimulus pattern complexities, and also demonstrates strong predictive power to infer pain ratings, including readouts based only on brain functional images. Guillermo A. Cecchi, Lejian Huang, Javeria Ali Hashmi, Marwan N. Baliki, María V. Centeno, Irina Rish, Apkar Vania Apkarian |
PLoS Comput. Biol. | 6 |
| 2010 | Learning Sparse Gaussian Markov Networks Using a Greedy Coordinate Ascent Approach
Katya Scheinberg, Irina Rish |
ECML/PKDD (3) | 2 |
| 2009 | Map approach to learning sparse Gaussian Markov networksabstractRecently proposed l1-regularized maximum-likelihood optimization methods for learning sparse Markov networks result into convex problems that can be solved optimally and efficiently. However, the accuracy of such methods can be very sensitive to the choice of regularization parameter, and optimal selection of this parameter remains an open problem. Herein, we propose a maximum a posteriori probability (MAP) approach that investigates different priors on the regularization parameter and yields promising empirical results on both synthetic data and real-life application such as brain imaging data (fMRI). Narges Bani Asadi, Irina Rish, Katya Scheinberg, Dimitri Kanevsky, Bhuvana Ramabhadran |
ICASSP | 2 |
| 2009 | Discriminative Network Models of SchizophreniaabstractSchizophrenia is a complex psychiatric disorder that has eluded a characterization in terms of local abnormalities of brain activity, and is hypothesized to affect the collective, ``emergent working of the brain. We propose a novel data-driven approach to capture emergent features using functional brain networks [Eguiluzet al] extracted from fMRI data, and demonstrate its advantage over traditional region-of-interest (ROI) and local, task-specific linear activation analyzes. Our results suggest that schizophrenia is indeed associated with disruption of global, emergent brain properties related to its functioning as a network, which cannot be explained by alteration of local activation patterns. Moreover, further exploitation of interactions by sparse Markov Random Field classifiers shows clear gain over linear methods, such as Gaussian Naive Bayes and SVM, allowing to reach 86% accuracy (over 50% baseline - random guess), which is quite remarkable given that it is based on a single fMRI experiment using a simple auditory task. Guillermo A. Cecchi, Irina Rish, Benjamin Thyreau, Bertrand Thirion, Marion Plaze, Marie-Laure Paillère-Martinot, Catherine Martelli, Jean-Luc Martinot, Jean-Baptiste Poline |
NIPS | 2 |
| 2008 | Closed-form supervised dimensionality reduction with generalized linear modelsabstractWe propose a family of supervised dimensionality reduction (SDR) algorithms that combine feature extraction (dimensionality reduction) with learning a predictive model in a unified optimization framework, using data- and class-appropriate generalized linear models (GLMs), and handling both classification and regression problems. Our approach uses simple closed-form update rules and is provably convergent. Promising empirical results are demonstrated on a variety of high-dimensional datasets. Irina Rish, Genady Grabarnik, Guillermo A. Cecchi, Francisco Pereira 0001, Geoffrey J. Gordon |
ICML | 1 |
| 2007 | Estimating End-to-End Performance by Collaborative Prediction with Active SamplingabstractAccurately estimating end-to-end performance in distributed systems is essential both for monitoring compliance with service-level agreements (SLAs) and for performance optimization (e.g., choosing the highest-bandwidth server for a download request in a content-distribution system). Due to infeasibility of exhaustive pairwise measurements, a natural alternative is to predict unobserved end-to-end performances from available historic data, with minimal additional measurements. In this paper we present an approach to this based on Collaborative Prediction (CP), an estimation method designed to work with sparse data, that has enjoyed much success in other domains (e.g. product recommendation systems), and obviates the need for landmark nodes commonly assumed in other approaches. Specifically, we use Max-Margin Matrix Factorization (MMMF), a linear factor model for CP that has outperformed state- of-art CP techniques. Moreover, our approach readily admits active sampling based on prediction confidence, and we further propose a novel active-sampling CP approach yielding even higher predictive accuracy, while allowing a flexible trade-off between "exploration" (choosing suboptimal samples to improve estimation accuracy) and "exploitation" (choosing node with best estimated performance). We demonstrate successful empirical results on a variety of practical problems, including network latency prediction (NLANR-AMP, P2PSim and PlanetLab datasets) and bandwidth prediction in content-distribution systems (IBM's downloadGrid data). Irina Rish, Gerald Tesauro |
Integrated Network Management | 1 |
| 2007 | Blind source separation approach to performance diagnosis and dependency discoveryabstractWe consider the problem of diagnosing performance problems in distributed system and networks given end-to-end performance measurements provided by test transactions, or probes. Common techniques for problem diagnosis such as, for example, codebook and network tomography usually assume a known dependency (e.g., routing) matrix that describes how each probe depends on the systems components. However, collecting full information about routing and/or probe dependencies on all systems components can be very costly, if not impossible, in large-scale, dynamic networks and distributed systems. We propose an approach to problem diagnosis and dependency discovery from end-to-end performance measurements in cases when the dependency/routing information is unknown or partially known. Our method is based on Blind Source Separation (BSS) approach that aims at reconstructing unobserved input signals and the mixing-weights matrix from the observed mixtures of signals. Particularly, we apply sparse non-negative matrix factorization techniques that appear particularly fitted to the problem of recovering network bottlenecks and dependency (routing) matrix, and show promising experimental results on several realistic network topologies. Gaurav Chandalia, Irina Rish |
Internet Measurement Conference | 2 |
| 2007 | Empirical Study of Topology Effects on Diagnosis in Computer NetworksabstractIn this paper, we compare the efficiency of fault detection and diagnosis in networks having different topological properties, such as scale-free networks and Erdos-Renyi random graphs. Efficiency measures include both the number of tests (e.g., end-to-end network probes) necessary for diagnosis and the computational complexity of diagnosis. We observe that diagnosis in scale-free networks typically requires significantly larger number of tests than diagnosis in random networks. However, the computational complexity of diagnosis appears to be much lower for scale-free networks since the corresponding Bayesian network models used for probabilistic diagnosis tend to have much lower induced width - a topological parameter controlling the complexity of inference in Bayesian networks. We believe that our observations provide important insights for design and deployment of cost-efficient diagnostic methods in computer networks and distributed systems. Natalia Odintsova, Irina Rish |
MASS | 2 |
| 2006 | Bayesian Learning of Markov Network Structure
Aleks Jakulin, Irina Rish |
ECML | 2 |
| 2005 | Test-based diagnosis: tree and matrix representationsabstractA common problem encountered in many application scenarios is how to represent some prior knowledge about a system in order to determine its true state as efficiently as possible. The information is typically in the form of tests, or questions about the system. Each test can potentially reduce our uncertainty about the system's state. The problem is to represent the information capturing the dependence between tests, their outcomes, and possible states in an efficiently navigable way to aid diagnosis. The most common such representation is a flowchart with leaf nodes corresponding to possible states, and non-leaf nodes corresponding to tests about the state. The problem with flowcharts is that they are notoriously difficult to maintain. Additional knowledge often has to be manually integrated as the system changes, making it impossible to keep track of all possible decision paths, let alone optimize the flow to maximize performance. We propose an efficient method for optimizing an existing flowchart based on a conversion to an auxiliary matrix representation. The main goal of the paper is show a synergy between the two representations in the hope that this will help practitioners choose a better strategy for their applications. We show that such a conversion suggests ways to improve both representations - ways that were not envisioned when using each representation alone. Finally, we show that the two representations are informationally equivalent in the sense that one can be transformed into the other so that if both are used as black-boxes, one would not be able to tell them apart, regardless of which state the system is in. Alina Beygelzimer, Mark Brodie, Sheng Ma, Irina Rish |
Integrated Network Management | 4 |
| 2005 | Statictical Models for Unequally Spaced Time SeriesabstractIrregularly observed time series and their analysis are fundamental for any application in which data are collected in a distributed or asynchronous manor. We propose a theoretical framework for analyzing both stationary and non-stationary irregularly spaced time series. Our models can be viewed as extensions of the well known autoregression (AR) model. We provide experiments suggesting that, in practice, the proposed approach performs well in computing the basic statistics and doing prediction. We also develop a resampling strategy that uses the proposed models to reduce irregular time series to regular time series. This enables us to take advantage of the vast number of approaches developed for analyzing regular time series. Alina Beygelzimer, Emre Erdogan, Sheng Ma, Irina Rish |
SDM | 4 |
| 2005 | Efficient Test Selection in Active Diagnosis via Entropy Approximation
Alice X. Zheng, Irina Rish, Alina Beygelzimer |
UAI | 2 |
| 2005 | Adaptive diagnosis in distributed systemsabstractReal-time problem diagnosis in large distributed computer systems and networks is a challenging task that requires fast and accurate inferences from potentially huge data volumes. In this paper, we propose a cost-efficient, adaptive diagnostic technique called active probing. Probes are end-to-end test transactions that collect information about the performance of a distributed system. Active probing uses probabilistic reasoning techniques combined with information-theoretic approach, and allows a fast online inference about the current system state via active selection of only a small number of most-informative tests. We demonstrate empirically that the active probing scheme greatly reduces both the number of probes (from 60% to 75% in most of our real-life applications), and the time needed for localizing the problem when compared with nonadaptive (preplanned) probing schemes. We also provide some theoretical results on the complexity of probe selection, and the effect of "noisy" probes on the accuracy of diagnosis. Finally, we discuss how to model the system's dynamics using dynamic Bayesian networks (DBNs), and an efficient approximate approach called sequential multifault; empirical results demonstrate clear advantage of such approaches over "static" techniques that do not handle system's changes. Irina Rish, Mark Brodie, Sheng Ma, Natalia Odintsova, Alina Beygelzimer, Genady Grabarnik, Karina Hernandez |
IEEE Trans. Neural Networks | 1 |
| 2004 | Real-time problem determination in distributed systems using active probingabstractWe describe algorithms and an architecture for a real-time problem determination system that uses online selection of most-informative measurements - the approach called herein active probing. Probes are end-to-end test transactions which gather information about system components. Active probing allows probes to be selected and sent on-demand, in response to one's belief about the state of the system. At each step the most informative next probe is computed and sent. As probe results are received, belief about the system state is updated using probabilistic inference. This process continues until the problem is diagnosed. We demonstrate through both analysis and simulation that the active probing scheme greatly reduces both the number of probes and the time needed for localizing the problem when compared with non-active probing schemes. Irina Rish, Mark Brodie, Natalia Odintsova, Sheng Ma, Genady Grabarnik |
NOMS (1) | 1 |
| 2003 | A Decomposition of Classes via Clustering to Explain and Improve Naive Bayes
Ricardo Vilalta, Irina Rish |
ECML | 2 |
| 2003 | Active Probing Strategies for Problem Diagnosis in Distributed Systems
Mark Brodie, Irina Rish, Sheng Ma, Natalia Odintsova |
IJCAI | 2 |
| 2003 | Critical event prediction for proactive management in large-scale computer clustersabstractAs the complexity of distributed computing systems increases, systems management tasks require significantly higher levels of automation; examples include diagnosis and prediction based on real-time streams of computer events, setting alarms, and performing continuous monitoring. The core of autonomic computing, a recently proposed initiative towards next-generation IT-systems capable of 'self-healing', is the ability to analyze data in real-time and to predict potential problems. The goal is to avoid catastrophic failures through prompt execution of remedial actions.This paper describes an attempt to build a proactive prediction and control system for large clusters. We collected event logs containing various system reliability, availability and serviceability (RAS) events, and system activity reports (SARs) from a 350-node cluster system for a period of one year. The 'raw' system health measurements contain a great deal of redundant event data, which is either repetitive in nature or misaligned with respect to time. We applied a filtering technique and modeled the data into a set of primary and derived variables. These variables used probabilistic networks for establishing event correlations through prediction algorithms. We also evaluated the role of time-series methods, rule-based classification algorithms and Bayesian network models in event prediction.Based on historical data, our results suggest that it is feasible to predict system performance parameters (SARs) with a high degree of accuracy using time-series models. Rule-based classification techniques can be used to extract machine-event signatures to predict critical events with up to 70% accuracy. Ramendra K. Sahoo, Adam J. Oliner, Irina Rish, Manish Gupta 0002, José E. Moreira, Sheng Ma, Ricardo Vilalta, Anand Sivasubramaniam |
KDD | 3 |
| 2003 | Approximability of Probability DistributionsabstractWe consider the question of how well a given distribution can be approx- imated with probabilistic graphical models. We introduce a new param- eter, effective treewidth, that captures the degree of approximability as a tradeoff between the accuracy and the complexity of approximation. We present a simple approach to analyzing achievable tradeoffs that ex- ploits the threshold behavior of monotone graph properties, and provide experimental results that support the approach. Alina Beygelzimer, Irina Rish |
NIPS | 2 |
| 2003 | Mini-buckets: A general scheme for bounded inferenceabstractThis article presents a class of approximation algorithms that extend the idea of bounded-complexity inference, inspired by successful constraint propagation algorithms, to probabilistic inference and combinatorial optimization. The idea is to bound the dimensionality of dependencies created by inference algorithms. This yields a parameterized scheme, called mini-buckets , that offers adjustable trade-off between accuracy and efficiency. The mini-bucket approach to optimization problems, such as finding the most probable explanation (MPE) in Bayesian networks, generates both an approximate solution and bounds on the solution quality. We present empirical results demonstrating successful performance of the proposed approximation scheme for the MPE task, both on randomly generated problems and on realistic domains such as medical diagnosis and probabilistic decoding. Rina Dechter, Irina Rish |
J. ACM | 2 |
| 2002 | Inference Complexity as a Model-Selection Criterion for Learning Bayesian Networks
Alina Beygelzimer, Irina Rish |
KR | 2 |
| 2001 | A Unified Framework for Evaluation Metrics in Classification Using Decision Trees
Ricardo Vilalta, Mark Brodie, Daniel Oblinger, Irina Rish |
ECML | 4 |
| 2000 | Resolution versus Search: Two Strategies for SAT
Irina Rish, Rina Dechter |
J. Autom. Reason. | 1 |
| 1998 | Empirical Evaluation of Approximation Algorithms for Probabilistic Decoding
Irina Rish, Kalev Kask, Rina Dechter |
UAI | 1 |
| 1997 | Statistical Analysis of Backtracking on Inconsistent CSPs
Irina Rish, Daniel Frost |
CP | 1 |
| 1997 | A Scheme for Approximating Probabilistic Inference
Rina Dechter, Irina Rish |
UAI | 2 |
| 1996 | To Guess or to Think? Hybrid Algorithms for SAT (Extended Abstract)
Irina Rish, Rina Dechter |
CP | 1 |
| 1994 | Directional Resolution: The Davis-Putnam Procedure, Revisited
Rina Dechter, Irina Rish |
KR | 2 |