VLDB 2026 Research / reviewers in the wild / expert
Aram Galstyan
dblp:16/3411
· DBLP profile ↗
98ranked-venue papers
8as first author
42since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 77 · 4 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Systems, architecture and hardware · 3 · 1 first-authorTheory of computation · 3 · 1 first-author · 1 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward SystemabstractJiacheng Liang, Yao Ma, Tharindu Kumarage, Satyapriya Krishna, Rahul Gupta, Kai-Wei Chang, Aram Galstyan, Charith Peris. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiacheng Liang, Tharindu Kumarage, Satyapriya Krishna, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan, Charith Peris |
ACL (1) | 7 |
| 2026 | SWAN: Semantic Watermarking with Abstract Meaning RepresentationabstractZiping Ye, Gourab Dey, Christos Christodoulopoulos, Charith Peris, Anil Ramakrishna, Weitong Ruan, Aram Galstyan, Kai-Wei Chang, Rahul Gupta, Ninareh Mehrabi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ziping Ye, Gourab Dey, Christos Christodoulopoulos 0001, Charith Peris, Anil Ramakrishna, Weitong Ruan, Aram Galstyan, Kai-Wei Chang 0001, Rahul Gupta 0001, Ninareh Mehrabi |
ACL (1) | 7 |
| 2025 | Accelerated Test-Time Scaling with Model-Free Speculative SamplingabstractWoomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin, Aram Galstyan, Sravan Babu Bodapati. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin, Aram Galstyan, Sravan Babu Bodapati |
EMNLP | 6 |
| 2025 | Knowledge Enhanced Multi-Domain Recommendations in an AI Assistant ApplicationabstractThis work explores unifying knowledge enhanced recommendation with multi-domain recommendation systems in a conversational AI assistant application. Multi-domain recommendation leverages users’ interactions in previous domains to improve recommendations in a new one. Knowledge graph enhancement seeks to use external knowledge graphs to improve recommendations within a single domain. Both research threads incorporate related information to improve the recommendation task. We propose to unify these approaches: using information from interactions in other domains as well as external knowledge graphs to make predictions in a new domain that would not be possible with either information source alone. We develop a new model and demonstrate the additive benefit of these approaches on a dataset derived from millions of users’ queries for content across three domains (videos, music, and books) in a live virtual assistant application. We demonstrate significant improvement on overall recommendations as well as on recommendations for new users of a domain. Elan Markowitz, Ziyan Jiang, Fan Yang 0155, Zheng Chen 0010, Greg Ver Steeg, Aram Galstyan |
ICASSP | 7 |
| 2025 | SeRA: Self-Reviewing and Alignment of LLMs using Implicit Reward MarginsabstractDirect alignment algorithms (DAAs), such as direct preference optimization (DPO), have become popular alternatives to Reinforcement Learning from Human Feedback (RLHF) due to their simplicity, efficiency, and stability. However, the preferences used by DAAs are usually collected before alignment training begins and remain unchanged (off-policy). This design leads to two problems where the policy model (1) picks up on spurious correlations in the dataset (as opposed to only learning alignment to human preferences), and (2) overfits to feedback on off-policy trajectories that have less likelihood of being generated by the updated policy model. To address these issues, we introduce Self-Reviewing and Alignment (SeRA), a cost-efficient and effective method that can be readily combined with existing DAAs. SeRA comprises of two components: (1) sample selection using implicit reward margin to alleviate over-optimization on such undesired features, and (2) preference bootstrapping using implicit rewards to augment preference data with updated policy models in a cost-efficient manner. Extensive experiments, including on instruction-following tasks, demonstrate the effectiveness and generality of SeRA in training LLMs with diverse offline preference datasets and and DAAs. Jongwoo Ko, Saket Dingliwal, Bhavana Ganesh, Sailik Sengupta, Sravan Babu Bodapati, Aram Galstyan |
ICLR | 6 |
| 2025 | Compress, Gather, and Recompute: REFORMing Long-Context Processing in TransformersabstractAs large language models increasingly gain popularity in real-world applications, processing extremely long contexts, often exceeding the model’s pre-trained context limits, has emerged as a critical challenge. While existing approaches to efficient long-context processing show promise, recurrent compression-based methods struggle with information preservation, whereas random access approaches require substantial memory resources. We introduce REFORM, a novel inference framework that efficiently handles long contexts through a two-phase approach. First, it incrementally processes input chunks while maintaining a compressed KV cache, constructs cross-layer context embeddings, and utilizes early exit strategy for improved efficiency. Second, it identifies and gathers essential tokens via similarity matching and selectively recomputes the KV cache. Compared to baselines, REFORM achieves over 50% and 27% performance gains on RULER and BABILong respectively at 1M context length. It also outperforms baselines on ∞-Bench, RepoEval, and MM-NIAH, demonstrating flexibility across diverse tasks and domains. Additionally, REFORM reduces inference time by 30% and peak memory usage by 5%, achieving both efficiency and superior performance. Woomin Song, Sai Muralidhar Jayanthi, Srikanth Ronanki, Kanthashree Mysore Sathyendra, Jinwoo Shin, Aram Galstyan, Shubham Katiyar, Sravan Babu Bodapati |
NeurIPS | 6 |
| 2024 | Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge GraphsabstractElan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta, Kai-Wei Chang, Aram Galstyan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Elan Markowitz, Anil Ramakrishna, Jwala Dhamala, Ninareh Mehrabi, Charith Peris, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
ACL (1) | 8 |
| 2024 | Policy Learning for Localized Interventions from Observational DataabstractA largely unaddressed problem in causal inference is that of learning reliable policies in continuous, high-dimensional treatment variables from observational data. Especially in the presence of strong confounding, it can be infeasible to learn the entire heterogeneous response surface from treatment to outcome. It is also not particularly useful, when there are practical constraints on the size of the interventions altering the observational treatments. Since it tends to be easier to learn the outcome for treatments near existing observations, we propose a new framework for evaluating and optimizing the effect of small, tailored, and localized interventions that nudge the observed treatment assignments. Our doubly robust effect estimator plugs into a policy learner that stays within the interventional scope by optimal transport. Consequently, the error of the total policy effect is restricted to prediction errors nearby the observational distribution, rather than the whole response surface. Myrl G. Marmarelis, Fred Morstatter, Aram Galstyan, Greg Ver Steeg |
AISTATS | 3 |
| 2024 | Agenda-Driven Question Generation: A Case Study in the Courtroom DomainabstractThis paper introduces a novel problem of automated question generation for courtroom examinations, CourtQG. While question generation has been studied in domains such as educational testing and product description, CourtQG poses several unique challenges owing to its non-cooperative and agenda-driven nature. Specifically, not only the generated questions need to be relevant to the case and underlying context, they also have to achieve certain objectives such as challenging the opponent’s arguments and/or revealing potential inconsistencies in their answers. We propose to leverage large language models (LLM) for CourtQG by fine-tuning them on two auxiliary tasks, agenda explanation (i.e., uncovering the underlying intents) and question type prediction. We additionally propose cold-start generation of questions from background documents without relying on examination history. We construct a dataset to evaluate our proposed method and show that it generates better questions according to standard metrics when compared to several baselines. Yi R. Fung 0001, Aram Galstyan, Heng Ji 0001, Premkumar Natarajan |
LREC/COLING | 3 |
| 2024 | FLIRT: Feedback Loop In-context Red TeamingabstractNinareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard Zemel, Kai-Wei Chang, Aram Galstyan, Rahul Gupta. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Shalini Ghosh, Richard S. Zemel, Kai-Wei Chang 0001, Aram Galstyan, Rahul Gupta 0001 |
EMNLP | 8 |
| 2024 | Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language ModelsabstractData is a crucial element in large language model (LLM) alignment.Recent studies have explored using LLMs for efficient data collection.However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints.To address these problems, we propose DATA ADVISOR, an enhanced LLMbased method for generating data that takes into account the characteristics of the desired dataset.Starting from a set of pre-defined principles in hand, DATA ADVISOR monitors the status of the generated data, identifies weaknesses in the current dataset, and advises the next iteration of data generation accordingly.DATA ADVISOR can be easily integrated into existing data generation methods to enhance data quality and coverage.Experiments on safety alignment of three representative LLMs (i.e., Mistral, Llama2, and Falcon) demonstrate the effectiveness of DATA ADVISOR in enhancing model safety against various fine-grained safety issues without sacrificing model utility.Warning: this paper contains example data that may be offensive or harmful. Fei Wang 0060, Ninareh Mehrabi, Palash Goyal, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
EMNLP | 6 |
| 2024 | The steerability of large language models toward data-driven personasabstractJunyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, Rahul Gupta. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Junyi Li 0002, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang 0001, Aram Galstyan, Richard S. Zemel, Rahul Gupta 0001 |
NAACL-HLT | 6 |
| 2023 | Overcoming Concept Shift in Domain-Aware Settings through Consolidated Internal DistributionsabstractWe develop an algorithm to improve the predictive performance of a pre-trained model under \textit{concept shift} without retraining the model from scratch when only unannotated samples of initial concepts are accessible. We model this problem as a domain adaptation problem, where the source domain data is inaccessible during model adaptation. The core idea is based on consolidating the intermediate internal distribution, learned to represent the source domain data, after adapting the model. We provide theoretical analysis and conduct extensive experiments on five benchmark datasets to demonstrate that the proposed method is effective. Aram Galstyan |
AAAI | 2 |
| 2023 | ACCENT: An Automatic Event Commonsense Evaluation Metric for Open-Domain Dialogue SystemsabstractCommonsense reasoning is omnipresent in human communications and thus is an important feature for open-domain dialogue systems.However, evaluating commonsense in dialogue systems is still an open challenge.We take the first step by focusing on event commonsense that considers events and their relations, and is crucial in both dialogues and general commonsense reasoning.We propose AC-CENT, an event commonsense evaluation metric empowered by commonsense knowledge bases (CSKBs).ACCENT first extracts eventrelation tuples from a dialogue, and then evaluates the response by scoring the tuples in terms of their compatibility with the CSKB.To evaluate ACCENT, we construct the first public event commonsense evaluation dataset for open-domain dialogues.Our experiments show that ACCENT is an efficient metric for event commonsense evaluation, which achieves higher correlations with human judgments than existing baselines. Sarik Ghazarian, Yijia Shao, Rujun Han, Aram Galstyan, Nanyun Peng 0001 |
ACL (1) | 4 |
| 2023 | ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-TranslationabstractKuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar, Kai-Wei Chang, Aram Galstyan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kuan-Hao Huang, Varun Iyer, I-Hung Hsu, Kai-Wei Chang 0001, Aram Galstyan |
ACL (1) | 6 |
| 2023 | Resolving Ambiguities in Text-to-Image Generative ModelsabstractNinareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Rahul Gupta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Kai-Wei Chang 0001, Richard S. Zemel, Aram Galstyan, Rahul Gupta 0001 |
ACL (1) | 9 |
| 2023 | The First Workshop on Personalized Generative AI @ CIKM 2023: Personalization Meets Large Language ModelsabstractThe First Workshop on Personalized Generative AI1 aims to be a cornerstone event fostering innovation and collaboration in the dynamic field of personalized AI. Leveraging the potent capabilities of Large Language Models (LLMs) to enhance user experiences with tailored responses and recommendations, the workshop is designed to address a range of pressing challenges including knowledge gap bridging, hallucination mitigation, and efficiency optimization in handling extensive user profiles. As a nexus for academics and industry professionals, the event promises rich discussions on a plethora of topics such as the development and fine-tuning of foundational models, strategies for multi-modal personalization, and the imperative ethical and privacy considerations in LLM deployment. Through a curated series of keynote speeches, insightful panel discussions, and hands-on sessions, the workshop aspires to be a catalyst in the development of more precise, contextually relevant, and user-centric AI systems. It aims to foster a landscape where generative AI systems are not only responsive but also anticipatory of individual user needs, marking a significant stride in personalized experiences. Zheng Chen 0010, Ziyan Jiang, Fan Yang 0155, Zhankui He, Yupeng Hou, Eunah Cho, Julian J. McAuley, Aram Galstyan, Xiaohua Hu 0001, Jie Yang 0028 |
CIKM | 8 |
| 2023 | Cognitively Inspired Learning of Incremental Drifting ConceptsabstractHumans continually expand their learned knowledge to new domains and learn new concepts without any interference with past learned experiences. In contrast, machine learning models perform poorly in a continual learning setting, where input data distribution changes over time. Inspired by the nervous system learning mechanisms, we develop a computational model that enables a deep neural network to learn new concepts and expand its learned knowledge to new domains incrementally in a continual learning setting. We rely on the Parallel Distributed Processing theory to encode abstract concepts in an embedding space in terms of a multimodal distribution. This embedding space is modeled by internal data representations in a hidden network layer. We also leverage the Complementary Learning Systems theory to equip the model with a memory mechanism to overcome catastrophic forgetting through implementing pseudo-rehearsal. Our model can generate pseudo-data points for experience replay and accumulate new experiences to past learned experiences without causing cross-task interference. Aram Galstyan |
IJCAI | 2 |
| 2023 | Partial identification of dose responses with hidden confoundersabstractInferring causal effects of continuous-valued treatments from observational data is a crucial task promising to better inform policy- and decision-makers. A critical assumption needed to identify these effects is that all confounding variables—causal parents of both the treatment and the outcome—are included as covariates. Unfortunately, given observational data alone, we cannot know with certainty that this criterion is satisfied. Sensitivity analyses provide principled ways to give bounds on causal estimates when confounding variables are hidden. While much attention is focused on sensitivity analyses for discrete-valued treatments, much less is paid to continuous-valued treatments. We present novel methodology to bound both average and conditional average continuous-valued treatment-effect estimates when they cannot be point identified due to hidden confounding. A semi-synthetic benchmark on multiple datasets shows our method giving tighter coverage of the true dose-response curve than a recently proposed continuous sensitivity model and baselines. Finally, we apply our method to a real-world observational case study to demonstrate the value of identifying dose-dependent causal effects. Myrl G. Marmarelis, Elizabeth Haddad, Andrew Jesson, Neda Jahanshad, Aram Galstyan, Greg Ver Steeg |
UAI | 5 |
| 2023 | Incorporating Fairness in Large Scale NLU SystemsabstractNLU models power several user facing experiences such as conversations agents and chat bots. Building NLU models typically consist of 3 stages: a) building or finetuning a pre-trained model b) distilling or fine-tuning the pre-trained model to build task specific models and, c) deploying the task-specific model to production. In this presentation, we will identify fairness considerations that can be incorporated in the aforementioned three stages in the life-cycle of NLU model building: (i) selection/building of a large scale language model, (ii) distillation/fine-tuning the large model into task specific model and, (iii) deployment of the task specific model. We will present select metrics that can be used to quantify fairness in NLU models and fairness enhancement techniques that can be deployed in each of these stages. Finally, we will share some recommendations to successfully implement fairness considerations when building an industrial scale NLU system. Rahul Gupta 0001, Lisa Bauer, Kai-Wei Chang 0001, Jwala Dhamala, Aram Galstyan, Palash Goyal, Avni Khatri, Rohit Parimi, Charith Peris, Apurv Verma, Richard S. Zemel, Premkumar Natarajan |
WSDM | 5 |
| 2022 | DEAM: Dialogue Coherence Evaluation using AMR-based Semantic ManipulationsabstractAutomatic evaluation metrics are essential for the rapid development of open-domain dialogue systems as they facilitate hyperparameter tuning and comparison between models.Although recently proposed trainable conversation-level metrics have shown encouraging results, the quality of the metrics is strongly dependent on the quality of training data.Prior works mainly resort to heuristic textlevel manipulations (e.g.utterances shuffling) to bootstrap incoherent conversations (negative examples) from coherent dialogues (positive examples).Such approaches are insufficient to appropriately reflect the incoherence that occurs in interactions between advanced dialogue models and humans.To tackle this problem, we propose DEAM, a Dialogue coherence Evaluation metric that relies on Abstract Meaning Representation (AMR) to apply semanticlevel Manipulations for incoherent (negative) data generation.AMRs naturally facilitate the injection of various types of incoherence sources, such as coreference inconsistency, irrelevancy, contradictions, and decrease engagement, at the semantic level, thus resulting in more natural incoherent samples.Our experiments show that DEAM 1 achieves higher correlations with human judgments compared to baseline methods on several dialog datasets by significant margins.We also show that DEAM can distinguish between coherent and incoherent dialogues generated by baseline manipulations, whereas those baseline models cannot detect incoherent examples generated by DEAM.Our results demonstrate the potential of AMRbased semantic manipulations for natural negative example generation. Sarik Ghazarian, Nuan Wen, Aram Galstyan, Nanyun Peng 0001 |
ACL (1) | 3 |
| 2022 | Failure Modes of Domain Generalization AlgorithmsabstractDomain generalization algorithms use training data from multiple domains to learn models that generalize well to unseen domains. While recently proposed benchmarks demon-strate that most of the existing algorithms do not outperform simple baselines, the established evaluation methods fail to expose the impact of various factors that contribute to the poor performance. In this paper we propose an evaluation framework for domain generalization algorithms that allows decomposition of the error into components capturing distinct aspects of generalization. Inspired by the prevalence of algorithms based on the idea of domain-invariant representation learning, we extend the evaluation framework to capture various types of failures in achieving invariance. We show that the largest contributor to the generalization error varies across methods, datasets, regularization strengths and even training lengths. We observe two problems associated with the strategy of learning domain-invariant representations. On Colored MNIST, most domain generalization algorithms fail because they reach domain-invariance only on the training domains. On Camelyon-17, domain-invariance degrades the quality of representations on unseen domains. We hypothesize that focusing instead on tuning the classifier on top of a rich representation can be a promising direction. Tigran Galstyan, Hrayr Harutyunyan, Hrant Khachatrian, Greg Ver Steeg, Aram Galstyan |
CVPR | 5 |
| 2022 | Learning Under Label Noise for Robust Spoken Language Understanding systems
Aravind Illa, Sriram Venkatapathy, Subhrangshu Nandi, Pritam Varma, Anurag Dwarakanath, Aram Galstyan |
INTERSPEECH | 8 |
| 2022 | Formal limitations of sample-wise information-theoretic generalization boundsabstractSome of the tightest information-theoretic generalization bounds depend on the average information between the learned hypothesis and a single training example. However, these sample-wise bounds were derived only for expected generalization gap. We show that even for expected squared generalization gap no such sample-wise information-theoretic bounds exist. The same is true for PAC-Bayes and single-draw bounds. Remarkably, PAC-Bayes, single-draw and expected squared generalization gap bounds that depend on information in pairs of examples exist. Hrayr Harutyunyan, Greg Ver Steeg, Aram Galstyan |
ITW | 3 |
| 2022 | Robust Conversational Agents against Imperceptible Toxicity TriggersabstractNinareh Mehrabi, Ahmad Beirami, Fred Morstatter, Aram Galstyan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ninareh Mehrabi, Ahmad Beirami, Fred Morstatter, Aram Galstyan |
NAACL-HLT | 4 |
| 2022 | An Analysis of The Effects of Decoding Algorithms on Fairness in Open-Ended Language GenerationabstractSeveral prior works have shown that language models (LMs) can generate text containing harmful social biases and stereotypes. While decoding algorithms play a central role in determining properties of LM generated text, their impact on the fairness of the generations has not been studied. We present a systematic analysis of the impact of decoding algorithms on LM fairness, and analyze the trade-off between fairness, diversity and quality. Our experiments with top-p, top-k and temperature decoding algorithms, in open-ended language generation, show that fairness across demographic groups changes significantly with change in decoding algorithm's hyper-parameters. Notably, decoding algorithms that output more diverse text also output more texts with negative sentiment and regard. We present several findings and provide recommendations on standardized reporting of decoding details in fairness evaluations and optimization of decoding algorithms for fairness alongside quality and diversity. Jwala Dhamala, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
SLT | 5 |
| 2022 | A Metric Space for Point Process ExcitationsabstractA multivariate Hawkes process enables self- and cross-excitations through a triggering matrix that behaves like an asymmetrical covariance structure, characterizing pairwise interactions between the event types. Full-rank estimation of all interactions is often infeasible in empirical settings. Models that specialize on a spatiotemporal application alleviate this obstacle by exploiting spatial locality, allowing the dyadic relationships between events to depend only on separation in time and relative distances in real Euclidean space. Here we generalize this framework to any multivariate Hawkes process, and harness it as a vessel for embedding arbitrary event types in a hidden metric space. Specifically, we propose a Hidden Hawkes Geometry (HHG) model to uncover the hidden geometry between event excitations in a multivariate point process. The low dimensionality of the embedding regularizes the structure of the inferred interactions. We develop a number of estimators and validate the model by conducting several experiments. In particular, we investigate regional infectivity dynamics of COVID-19 in an early South Korean record and recent Los Angeles confirmed cases. By additionally performing synthetic experiments on short records as well as explorations into options markets and the Ebola epidemic, we demonstrate that learning the embedding alongside a point process uncovers salient interactions in a broad range of applications. Myrl G. Marmarelis, Greg Ver Steeg, Aram Galstyan |
J. Artif. Intell. Res. | 3 |
| 2021 | Exacerbating Algorithmic Bias through Fairness AttacksabstractAlgorithmic fairness has attracted significant attention in recent years, with many quantitative measures suggested for characterizing the fairness of different machine learning algorithms. Despite this interest, the robustness of those fairness measures with respect to an intentional adversarial attack has not been properly addressed. Indeed, most adversarial machine learning has focused on the impact of malicious attacks on the accuracy of the system, without any regard to the system's fairness. We propose new types of data poisoning attacks where an adversary intentionally targets the fairness of a system. Specifically, we propose two families of attacks that target fairness measures. In the anchoring attack, we skew the decision boundary by placing poisoned points near specific target points to bias the outcome. In the influence attack on fairness, we aim to maximize the covariance between the sensitive attributes and the decision outcome and affect the fairness of the model. We conduct extensive experiments that indicate the effectiveness of our proposed attacks. Ninareh Mehrabi, Fred Morstatter, Aram Galstyan |
AAAI | 4 |
| 2021 | ForecastQA: A Question Answering Challenge for Event Forecasting with Temporal Text DataabstractWoojeong Jin, Rahul Khanna, Suji Kim, Dong-Ho Lee, Fred Morstatter, Aram Galstyan, Xiang Ren. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Woojeong Jin 0001, Rahul Khanna, Fred Morstatter, Aram Galstyan, Xiang Ren 0001 |
ACL/IJCNLP (1) | 6 |
| 2021 | Layer-Wise Neural Network Compression via Layer FusionabstractThis paper proposes \textit{layer fusion} - a model compression technique that discovers which weights to combine and then fuses weights of similar fully-connected, convolutional and attention layers. Layer fusion can significantly reduce the number of layers of the original network with little additional computation overhead, while maintaining competitive performance. From experiments on CIFAR-10, we find that various deep convolution neural networks can remain within 2% accuracy points of the original networks up to a compression ratio of 3.33 when iteratively retrained with layer fusion. For experiments on the WikiText-2 language modelling dataset, we compress Transformer models to 20% of their original size while being within 5 perplexity points of the original network. We also find that other well-established compression techniques can achieve competitive performance when compared to their original networks given a sufficient number of retraining steps. Generally, we observe a clear inflection point in performance as the amount of compression increases, suggesting a bound on the amount of compression that can be achieved before an exponential degradation in performance. James O'Neill, Greg Ver Steeg, Aram Galstyan |
ACML | 3 |
| 2021 | Influence Decompositions For Neural Network AttributionabstractMethods of neural network attribution have emerged out of a necessity for explanation and accountability in the predictions of black-box neural models. Most approaches use a variation of sensitivity analysis, where individual input variables are perturbed and the downstream effects on some output metric are measured. We demonstrate that a number of critical functional properties are not revealed when only considering lower-order perturbations. Motivated by these shortcomings, we propose a general framework for decomposing the orders of influence that a collection of input variables has on an output classification. These orders are based on the cardinality of input subsets which are perturbed to yield a change in classification. This decomposition can be naturally applied to attribute which input variables rely on higher-order coordination to impact the classification decision. We demonstrate that our approach correctly identifies higher-order attribution on a number of synthetic examples. Additionally, we showcase the differences between attribution in our approach and existing approaches on benchmark networks for MNIST and ImageNet. Kyle Reing, Greg Ver Steeg, Aram Galstyan |
AISTATS | 3 |
| 2021 | Lawyers are Dishonest? Quantifying Representational Harms in Commonsense Knowledge ResourcesabstractWarning: this paper contains content that may be offensive or upsetting.Commonsense knowledge bases (CSKB) are increasingly used for various natural language processing tasks.Since CSKBs are mostly human-generated and may reflect societal biases, it is important to ensure that such biases are not conflated with the notion of commonsense.Here we focus on two widely used CSKBs, ConceptNet and GenericsKB, and establish the presence of bias in the form of two types of representational harms, overgeneralization of polarized perceptions and representation disparity across different demographic groups in both CSKBs.Next, we find similar representational harms for downstream models that use ConceptNet.Finally, we propose a filtering-based approach for mitigating such harms, and observe that our filtered-based approach can reduce the issues in both resources and models but leads to a performance drop, leaving room for future work to build fairer and stronger commonsense models. Ninareh Mehrabi, Fred Morstatter, Jay Pujara, Xiang Ren 0001, Aram Galstyan |
EMNLP (1) | 6 |
| 2021 | Partner-Assisted Learning for Few-Shot Image ClassificationabstractFew-shot Learning has been studied to mimic human visual capabilities and learn effective models without the need of exhaustive human annotation. Even though the idea of meta-learning for adaptation has dominated the few-shot learning methods, how to train a feature extractor is still a challenge. In this paper, we focus on the design of training strategy to obtain an elemental representation such that the prototype of each novel class can be estimated from a few labeled samples. We propose a two-stage training scheme, Partner-Assisted Learning (PAL), which first trains a Partner Encoder to model pair-wise similarities and extract features serving as soft-anchors, and then trains a Main Encoder by aligning its outputs with soft-anchors while attempting to maximize classification performance. Two alignment constraints from logit-level and feature-level are designed individually. For each few-shot task, we perform prototype classification. Our method consistently outperforms the state-of-the-art methods on four benchmarks. Detailed ablation studies of PAL are provided to justify the selection of each component involved in training. Jiawei Ma, Hanchen Xie, Guangxing Han, Shih-Fu Chang, Aram Galstyan, Wael Abd-Almageed |
ICCV | 5 |
| 2021 | Graph Traversal with Tensor Functionals: A Meta-Algorithm for Scalable Learning
Elan Markowitz, Keshav Balasubramanian, Mehrnoosh Mirtaheri, Sami Abu-El-Haija, Bryan Perozzi, Greg Ver Steeg, Aram Galstyan |
ICLR | 7 |
| 2021 | Plot-guided Adversarial Example Construction for Evaluating Open-domain Story GenerationabstractSarik Ghazarian, Zixi Liu, Akash S M, Ralph Weischedel, Aram Galstyan, Nanyun Peng. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Sarik Ghazarian, Akash SM, Ralph M. Weischedel, Aram Galstyan, Nanyun Peng 0001 |
NAACL-HLT | 5 |
| 2021 | Implicit SVD for Graph Representation LearningabstractRecent improvements in the performance of state-of-the-art (SOTA) methods for Graph Representational Learning (GRL) have come at the cost of significant computational resource requirements for training, e.g., for calculating gradients via backprop over many data epochs. Meanwhile, Singular Value Decomposition (SVD) can find closed-form solutions to convex problems, using merely a handful of epochs. In this paper, we make GRL more computationally tractable for those with modest hardware. We design a framework that computes SVD of *implicitly* defined matrices, and apply this framework to several GRL tasks. For each task, we derive first-order approximation of a SOTA model, where we design (expensive-to-store) matrix $\mathbf{M}$ and train the model, in closed-form, via SVD of $\mathbf{M}$, without calculating entries of $\mathbf{M}$. By converging to a unique point in one step, and without calculating gradients, our models show competitive empirical test performance over various graphs such as article citation and biological interaction networks. More importantly, SVD can initialize a deeper model, that is architected to be non-linear almost everywhere, though behaves linearly when its parameters reside on a hyperplane, onto which SVD initializes. The deeper model can then be fine-tuned within only a few epochs. Overall, our algorithm trains hundreds of times faster than state-of-the-art methods, while competing on test empirical performance. We open-source our implementation at: https://github.com/samihaija/isvd Sami Abu-El-Haija, Hesham Mostafa, Marcel Nassar, Valentino Crespi, Greg Ver Steeg, Aram Galstyan |
NeurIPS | 6 |
| 2021 | Information-theoretic generalization bounds for black-box learning algorithmsabstractWe derive information-theoretic generalization bounds for supervised learning algorithms based on the information contained in predictions rather than in the output of the training algorithm. These bounds improve over the existing information-theoretic bounds, are applicable to a wider range of algorithms, and solve two key challenges: (a) they give meaningful results for deterministic algorithms and (b) they are significantly easier to estimate. We show experimentally that the proposed bounds closely follow the generalization gap in practical scenarios for deep learning. Hrayr Harutyunyan, Maxim Raginsky, Greg Ver Steeg, Aram Galstyan |
NeurIPS | 4 |
| 2021 | Hamiltonian Dynamics with Non-Newtonian Momentum for Rapid SamplingabstractSampling from an unnormalized probability distribution is a fundamental problem in machine learning with applications including Bayesian modeling, latent factor inference, and energy-based model training. After decades of research, variations of MCMC remain the default approach to sampling despite slow convergence. Auxiliary neural models can learn to speed up MCMC, but the overhead for training the extra model can be prohibitive. We propose a fundamentally different approach to this problem via a new Hamiltonian dynamics with a non-Newtonian momentum. In contrast to MCMC approaches like Hamiltonian Monte Carlo, no stochastic step is required. Instead, the proposed deterministic dynamics in an extended state space exactly sample the target distribution, specified by an energy function, under an assumption of ergodicity. Alternatively, the dynamics can be interpreted as a normalizing flow that samples a specified energy model without training. The proposed Energy Sampling Hamiltonian (ESH) dynamics have a simple form that can be solved with existing ODE solvers, but we derive a specialized solver that exhibits much better performance. ESH dynamics converge faster than their MCMC competitors enabling faster, more stable training of neural network energy models. Greg Ver Steeg, Aram Galstyan |
NeurIPS | 2 |
| 2021 | q-Paths: Generalizing the geometric annealing path using power meansabstractMany common machine learning methods involve the geometric annealing path, a sequence of intermediate densities between two distributions of interest constructed using the geometric average. While alternatives such as the moment-averaging path have demonstrated performance gains in some settings, their practical applicability remains limited by exponential family endpoint assumptions and a lack of closed form energy function. In this work, we introduce $q$-paths, a family of paths which is derived from a generalized notion of the mean, includes the geometric and arithmetic mixtures as special cases, and admits a simple closed form involving the deformed logarithm function from nonextensive thermodynamics. Following previous analysis of the geometric path, we interpret our $q$-paths as corresponding to a $q$-exponential family of distributions, and provide a variational representation of intermediate densities as minimizing a mixture of $\alpha$-divergences to the endpoints. We show that small deviations away from the geometric path yield empirical gains for Bayesian inference using Sequential Monte Carlo and generative model evaluation using Annealed Importance Sampling. Vaden Masrani, Rob Brekelmans, Thang Bui, Frank Nielsen, Aram Galstyan, Greg Ver Steeg, Frank D. Wood |
UAI | 5 |
| 2021 | MUSCLE: Strengthening Semi-Supervised Learning Via Concurrent Unsupervised Learning Using Mutual Information MaximizationabstractDeep neural networks are powerful, massively parameterized machine learning models that have been shown to perform well in supervised learning tasks. However, very large amounts of labeled data are usually needed to train deep neural networks. Several semi-supervised learning approaches have been proposed to train neural networks using smaller amounts of labeled data with a large amount of unlabeled data. The performance of these semisupervised methods significantly degrades as the size of labeled data decreases. We introduce Mutual-information-based Unsupervised & Semi-supervised Concurrent LEarning (MUSCLE), a hybrid learning approach that uses mutual information to combine both unsupervised and semisupervised learning. MUSCLE can be used as a standalone training scheme for neural networks, and can also be incorporated into other learning approaches. We show that the proposed hybrid model outperforms state of the art on several standard benchmarks, including CIFAR-10, CIFAR-100, and Mini-Imagenet. Furthermore, the performance gain consistently increases with the reduction in the amount of labeled data, as well as in the presence of bias. We also show that MUSCLE has the potential to boost the classification performance when used in the fine-tuning phase for a model pre-trained only on unlabeled data. Hanchen Xie, Mohamed E. Hussein 0001, Aram Galstyan, Wael Abd-Almageed |
WACV | 3 |
| 2021 | Bin2vec: learning representations of binary executable programs for security tasksabstractAbstract Tackling binary program analysis problems has traditionally implied manually defining rules and heuristics, a tedious and time consuming task for human analysts. In order to improve automation and scalability, we propose an alternative direction based on distributed representations of binary programs with applicability to a number of downstream tasks. We introduce Bin2vec, a new approach leveraging Graph Convolutional Networks (GCN) along with computational program graphs in order to learn a high dimensional representation of binary executable programs. We demonstrate the versatility of this approach by using our representations to solve two semantically different binary analysis tasks – functional algorithm classification and vulnerability discovery. We compare the proposed approach to our own strong baseline as well as published results, and demonstrate improvement over state-of-the-art methods for both tasks. We evaluated Bin2vec on 49191 binaries for the functional algorithm classification task, and on 30 different CWE-IDs including at least 100 CVE entries each for the vulnerability discovery task. We set a new state-of-the-art result by reducing the classification error by 40% compared to the source-code based inst2vec approach, while working on binary code. For almost every vulnerability class in our dataset, our prediction accuracy is over 80% (and over 90% in multiple classes). Shushan Arakelyan, Sima Arasteh, Christophe Hauser, Erik Kline, Aram Galstyan |
Cybersecur. | 5 |
| 2021 | Identifying and Analyzing Cryptocurrency Manipulations in Social MediaabstractInterest surrounding cryptocurrencies, digital or virtual currencies that are used as a medium for financial transactions, has grown tremendously in the recent years. The anonymity surrounding these currencies makes investors particularly susceptible to fraudity-such as “pump and dump” scams-where the goal is to artificially inflate the perceived worth of a currency, luring victims into investing before the fraudsters can sell their holdings. Because of the speed and relative anonymity offered by social platforms such as Twitter and Telegram, social media has become a preferred platform for scammers who wish to spread false hype about the cryptocurrency they are trying to pump. In this work, we propose and evaluate a computational approach that can automatically identify pump and dump scams as they unfold by combining information across social media platforms. We also develop a multi-modal approach for predicting whether a particular pump attempt will succeed or not. Finally, we analyze the prevalence of bots in cryptocurrency related tweets, and observe a significant increase in bot activity during the pump attempts. Mehrnoosh Mirtaheri, Sami Abu-El-Haija, Fred Morstatter, Greg Ver Steeg, Aram Galstyan |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2020 | Modeling Dialogues with Hashcode Representations: A Nonparametric ApproachabstractWe propose a novel dialogue modeling framework, the first-ever nonparametric kernel functions based approach for dialogue modeling, which learns hashcodes as text representations; unlike traditional deep learning models, it handles well relatively small datasets, while also scaling to large ones. We also derive a novel lower bound on mutual information, used as a model-selection criterion favoring representations with better alignment between the utterances of participants in a collaborative dialogue setting, as well as higher predictability of the generated responses. As demonstrated on three real-life datasets, including prominently psychotherapy sessions, the proposed approach significantly outperforms several state-of-art neural network based dialogue systems, both in terms of computational efficiency, reducing training time from days or weeks to hours, and the response quality, achieving an order of magnitude improvement over competitors in frequency of being chosen as the best model by human evaluators. Sahil Garg, Irina Rish, Guillermo A. Cecchi, Palash Goyal, Sarik Ghazarian, Shuyang Gao, Greg Ver Steeg, Aram Galstyan |
AAAI | 8 |
| 2020 | Predictive Engagement: An Efficient Metric for Automatic Evaluation of Open-Domain Dialogue SystemsabstractUser engagement is a critical metric for evaluating the quality of open-domain dialogue systems. Prior work has focused on conversation-level engagement by using heuristically constructed features such as the number of turns and the total time of the conversation. In this paper, we investigate the possibility and efficacy of estimating utterance-level engagement and define a novel metric, predictive engagement, for automatic evaluation of open-domain dialogue systems. Our experiments demonstrate that (1) human annotators have high agreement on assessing utterance-level engagement scores; (2) conversation-level engagement scores can be predicted from properly aggregated utterance-level engagement scores. Furthermore, we show that the utterance-level engagement scores can be learned from data. These scores can be incorporated into automatic evaluation metrics for open-domain dialogue systems to improve the correlation with human judgements. This suggests that predictive engagement can be used as a real-time feedback for training better dialogue models. Sarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, Nanyun Peng 0001 |
AAAI | 3 |
| 2020 | All in the Exponential Family: Bregman Duality in Thermodynamic Variational InferenceabstractThe recently proposed Thermodynamic Variational Objective (TVO) leverages thermodynamic integration to provide a family of variational inference objectives, which both tighten and generalize the ubiquitous Evidence Lower Bound (ELBO). However, the tightness of TVO bounds was not previously known, an expensive grid search was used to choose a “schedule” of intermediate distributions, and model learning suffered with ostensibly tighter bounds. In this work, we propose an exponential family interpretation of the geometric mixture curve underlying the TVO and various path sampling methods, which allows us to characterize the gap in TVO likelihood bounds as a sum of KL divergences. We propose to choose intermediate distributions using equal spacing in the moment parameters of our exponential family, which matches grid search performance and allows the schedule to adaptively update over the course of training. Finally, we derive a doubly reparameterized gradient estimator which improves model learning and allows the TVO to benefit from more refined bounds. To further contextualize our contributions, we provide a unified framework for understanding thermodynamic integration and the TVO using Taylor series remainders. Rob Brekelmans, Vaden Masrani, Frank D. Wood, Greg Ver Steeg, Aram Galstyan |
ICML | 5 |
| 2020 | Improving generalization by controlling label-noise information in neural network weightsabstractIn the presence of noisy or incorrect labels, neural networks have the undesirable tendency to memorize information about the noise. Standard regularization techniques such as dropout, weight decay or data augmentation sometimes help, but do not prevent this behavior. If one considers neural network weights as random variables that depend on the data and stochasticity of training, the amount of memorized information can be quantified with the Shannon mutual information between weights and the vector of all training labels given inputs, $I(w; \mathbf{y} \mid \mathbf{x})$. We show that for any training algorithm, low values of this term correspond to reduction in memorization of label-noise and better generalization bounds. To obtain these low values, we propose training algorithms that employ an auxiliary network that predicts gradients in the final layers of a classifier without accessing labels. We illustrate the effectiveness of our approach on versions of MNIST, CIFAR-10, and CIFAR-100 corrupted with various noise models, and on a large-scale dataset Clothing1M that has noisy labels. Hrayr Harutyunyan, Kyle Reing, Greg Ver Steeg, Aram Galstyan |
ICML | 4 |
| 2020 | Maximizing Multivariate Information With Error-Correcting CodesabstractMultivariate mutual information provides a conceptual framework for characterizing higher-order interactions in complex systems. Two well-known measures of multivariate information-total correlation and dual total correlation-admit a spectrum of measures with varying sensitivity to intermediate orders of dependence. Unfortunately, these intermediate measures have not received much attention due to their opaque representation of information. Here we draw on results from matroid theory to show that these measures are closely related to error-correcting codes. This connection allows us to derive the class of global maximizers for each measure, which coincide with maximum distance separable codes of order k. In addition to deepening the understanding of these measures and multivariate information more generally, we use these results to show that previously proposed bounds on information geometric quantities are met with equality for the global min and max. Kyle Reing, Greg Ver Steeg, Aram Galstyan |
IEEE Trans. Inf. Theory | 3 |
| 2019 | Kernelized Hashcode Representations for Relation ExtractionabstractKernel methods have produced state-of-the-art results for a number of NLP tasks such as relation extraction, but suffer from poor scalability due to the high cost of computing kernel similarities between natural language structures. A recently proposed technique, kernelized locality-sensitive hashing (KLSH), can significantly reduce the computational cost, but is only applicable to classifiers operating on kNN graphs. Here we propose to use random subspaces of KLSH codes for efficiently constructing an explicit representation of NLP structures suitable for general classification methods. Further, we propose an approach for optimizing the KLSH model for classification problems by maximizing an approximation of mutual information between the KLSH codes (feature vectors) and the class labels. We evaluate the proposed approach on biomedical relation extraction datasets, and observe significant and robust improvements in accuracy w.r.t. state-ofthe-art classifiers, along with drastic (orders-of-magnitude) speedup compared to conventional kernel methods. Sahil Garg, Aram Galstyan, Greg Ver Steeg, Irina Rish, Guillermo A. Cecchi, Shuyang Gao |
AAAI | 2 |
| 2019 | Auto-Encoding Total Correlation ExplanationabstractAdvances in unsupervised learning enable reconstruction and generation of samples from complex distributions, but this success is marred by the inscrutability of the representations learned. We propose an information-theoretic approach to characterizing disentanglement and dependence in representation learning using multivariate mutual information, also called total correlation. The principle of Total Cor-relation Ex-planation (CorEx) has motivated successful unsupervised learning applications across a variety of domains but under some restrictive assumptions. Here we relax those restrictions by introducing a flexible variational lower bound to CorEx. Surprisingly, we find this lower bound is equivalent to the one in variational autoencoders (VAE) under certain conditions. This information-theoretic view of VAE deepens our understanding of hierarchical VAE and motivates a new algorithm, AnchorVAE, that makes latent codes more interpretable through information maximization and enables generation of richer and more realistic samples. Shuyang Gao, Rob Brekelmans, Greg Ver Steeg, Aram Galstyan |
AISTATS | 4 |
| 2019 | Debiasing community detection: the importance of lowly connected nodesabstractCommunity detection is an important task in social network analysis, allowing us to identify and understand the communities within the social structures provided by the network. However, many community detection approaches either fail to assign low-degree (or lowly connected) users to communities, or assign them to trivially small communities that prevent them from being included in analysis. In this work we investigate how excluding these users can bias analysis results. We then introduce an approach that is more inclusive for lowly connected users by incorporating them into larger groups. Experiments show that our approach outperforms the existing state-of-the-art in terms of F1 and Jaccard similarity scores while reducing the bias towards low-degree users. Ninareh Mehrabi, Fred Morstatter, Nanyun Peng 0001, Aram Galstyan |
ASONAM | 4 |
| 2019 | Deep Structured Neural Network for Event Temporal Relation ExtractionabstractWe propose a novel deep structured learning framework for event temporal relation extraction.The model consists of 1) a recurrent neural network (RNN) to learn scoring functions for pair-wise relations, and 2) a structured support vector machine (SSVM) to make joint predictions.The neural network automatically learns representations that account for long-term contexts to provide robust features for the structured model, while the SSVM incorporates domain knowledge such as transitive closure of temporal relations as constraints to make better globally consistent decisions.By jointly training the two components, our model combines the benefits of both data-driven learning and knowledge exploitation.Experimental results on three highquality event temporal relation datasets (TCR, MATRES, and TB-Dense) demonstrate that incorporated with pre-trained contextualized embeddings, the proposed model achieves significantly better performances than the stateof-the-art methods on all three datasets.We also provide thorough ablation studies to investigate our model. Rujun Han, I-Hung Hsu, Mu Yang, Aram Galstyan, Ralph M. Weischedel, Nanyun Peng 0001 |
CoNLL | 4 |
| 2019 | Nearly-Unsupervised Hashcode Representations for Biomedical Relation ExtractionabstractSahil Garg, Aram Galstyan, Greg Ver Steeg, Guillermo Cecchi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sahil Garg, Aram Galstyan, Greg Ver Steeg, Guillermo A. Cecchi |
EMNLP/IJCNLP (1) | 2 |
| 2019 | MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood MixingabstractExisting popular methods for semi-supervised learning with Graph Neural Networks (such as the Graph Convolutional Network) provably cannot learn a general class of neighborhood mixing relationships. To address this weakness, we propose a new model, MixHop, that can learn these relationships, including difference operators, by repeatedly mixing feature representations of neighbors at various distances. MixHop requires no additional memory or computational complexity, and outperforms on challenging baselines. In addition, we propose sparsity regularization that allows us to visualize how the network prioritizes neighborhood information across different graph datasets. Our analysis of the learned architectures reveals that neighborhood mixing varies per datasets. Sami Abu-El-Haija, Bryan Perozzi, Amol Kapoor, Nazanin Alipourfard, Kristina Lerman, Hrayr Harutyunyan, Greg Ver Steeg, Aram Galstyan |
ICML | 8 |
| 2019 | SAGE: A Hybrid Geopolitical Event Forecasting SystemabstractForecasting of geopolitical events is a notoriously difficult task, with experts failing to significantly outperform a random baseline across many types of forecasting events. One successful way to increase the performance of forecasting tasks is to turn to crowdsourcing: leveraging many forecasts from non-expert users. Simultaneously, advances in machine learning have led to models that can produce reasonable, although not perfect, forecasts for many tasks. Recent efforts have shown that forecasts can be further improved by ``hybridizing'' human forecasters: pairing them with the machine models in an effort to combine the unique advantages of both. In this demonstration, we present Synergistic Anticipation of Geopolitical Events (SAGE), a platform for human/computer interaction that facilitates human reasoning with machine models. Fred Morstatter, Aram Galstyan, Gleb Satyukov, Daniel Benjamin, Andrés Abeliuk, Mehrnoosh Mirtaheri, K. S. M. Tozammel Hossain, Pedro A. Szekely, Emilio Ferrara, Akira Matsui, Mark Steyvers, Stephen Bennett, David V. Budescu, Mark Himmelstein, Michael D. Ward, Andreas Beger, Michele Catasta, Rok Sosic, Jure Leskovec, Pavel Atanasov, Regina Joseph, Rajiv Sethi, Ali E. Abbas |
IJCAI | 2 |
| 2019 | Exact Rate-Distortion in Autoencoders via Echo NoiseabstractCompression is at the heart of effective representation learning. However, lossy compression is typically achieved through simple parametric models like Gaussian noise to preserve analytic tractability, and the limitations this imposes on learning are largely unexplored. Further, the Gaussian prior assumptions in models such as variational autoencoders (VAEs) provide only an upper bound on the compression rate in general. We introduce a new noise channel, Echo noise, that admits a simple, exact expression for mutual information for arbitrary input distributions. The noise is constructed in a data-driven fashion that does not require restrictive distributional assumptions. With its complex encoding mechanism and exact rate regularization, Echo leads to improved bounds on log-likelihood and dominates beta-VAEs across the achievable range of rate-distortion trade-offs. Further, we show that Echo noise can outperform flow-based methods without the need to train additional distributional transformations. Rob Brekelmans, Daniel Moyer, Aram Galstyan, Greg Ver Steeg |
NeurIPS | 3 |
| 2019 | Fast structure learning with modular regularizationabstractEstimating graphical model structure from high-dimensional and undersampled data is a fundamental problem in many scientific fields. Existing approaches, such as GLASSO, latent variable GLASSO, and latent tree models, suffer from high computational complexity and may impose unrealistic sparsity priors in some cases. We introduce a novel method that leverages a newly discovered connection between information-theoretic measures and structured latent factor models to derive an optimization objective which encourages modular structures where each observed variable has a single latent parent. The proposed method has linear stepwise computational complexity w.r.t. the number of observed variables. Our experiments on synthetic data demonstrate that our approach is the only method that recovers modular structure better as the dimensionality increases. We also use our approach for estimating covariance structure for a number of real-world datasets and show that it consistently outperforms state-of-the-art estimators at a fraction of the computational cost. Finally, we apply the proposed method to high-resolution fMRI data (with more than 10^5 voxels) and show that it is capable of extracting meaningful patterns. Greg Ver Steeg, Hrayr Harutyunyan, Daniel Moyer, Aram Galstyan |
NeurIPS | 4 |
| 2019 | Coupled Clustering of Time-Series and NetworksabstractMotivated by the problem of human-trafficking, where it is often observed that criminal organizations are linked and behave similarly over time, we introduce the problem of Coupled Clustering of Time-series and their underlying Network. The goal is to find tightly connected subgroups of nodes that also have similar node-specific time series (temporal—not necessarily structural—behavior). We formulate the problem as a coupled matrix factorization for the time series, combined with regularization for network smoothness. We propose CCTN, and an incrementally-updated counterpart, CCTN-inc, which efficiently handles network updates. Extensive experiments show that CCTN is up to 4x more accurate than baselines that consider graph structure or time series alone, and CCTN-inc is up to 55x faster than CCTN. As an application, we explore an exclusive database with millions of online ads on human trafficking, and successfully deploy our technique to detect criminal organizations. Linhong Zhu, Pedro A. Szekely, Aram Galstyan, Danai Koutra |
SDM | 4 |
| 2018 | Invariant Representations without Adversarial TrainingabstractRepresentations of data that are invariant to changes in specified factors are useful for a wide range of problems: removing potential biases in prediction problems, controlling the effects of covariates, and disentangling meaningful factors of variation. Unfortunately, learning representations that exhibit invariance to arbitrary nuisance factors yet remain useful for other tasks is challenging. Existing approaches cast the trade-off between task performance and invariance in an adversarial way, using an iterative minimax optimization. We show that adversarial training is unnecessary and sometimes counter-productive; we instead cast invariant representation learning as a single information-theoretic objective that can be directly optimized. We demonstrate that this approach matches or exceeds performance of state-of-the-art adversarial approaches for learning fair representations and for generative modeling with controllable transformations. Daniel Moyer, Shuyang Gao, Rob Brekelmans, Aram Galstyan, Greg Ver Steeg |
NeurIPS | 4 |
| 2018 | A Forest Mixture Bound for Block-Free Parallel Inference
Neal Lawton, Greg Ver Steeg, Aram Galstyan |
UAI | 3 |
| 2018 | Adaptive decision making via entropy minimization
Armen E. Allahverdyan, Aram Galstyan, Ali E. Abbas, Zbigniew R. Struzik |
Int. J. Approx. Reason. | 2 |
| 2018 | Capturing Edge Attributes via Network EmbeddingabstractNetwork embedding, which aims to learn low-dimensional representations of nodes, has been used for various graph related tasks including visualization, link prediction, and node classification. Most existing embedding methods rely solely on network structure. However, in practice, we often have auxiliary information about the nodes and/or their interactions, e.g., the content of scientific papers in coauthorship networks, or topics of communication in Twitter mention networks. Here, we propose a novel embedding method that uses both network structure and edge attributes to learn better network representations. Our method jointly minimizes the reconstruction error for higher order node neighborhood, social roles, and edge attributes using a deep architecture that can adequately capture highly nonlinear interactions. We demonstrate the efficacy of our model over existing state-of-the-art methods on a variety of real-world networks including collaboration networks and social networks. We also observe that using edge attributes to inform network embedding yields better performance in downstream tasks such as link prediction and node classification. Palash Goyal, Homa Hosseinmardi, Emilio Ferrara, Aram Galstyan |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2017 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social Networks (Extended Abstract)abstractWe propose to model dependence within a network view using the temporal latent space model, which uses a time-dependent low-dimensional geometric projections to represent the high-dimensional dependence structure in time-varying networks. Once we obtain the lowdimensional temporal latent space representation for graphs from time 1 to t, we can accurately predict future links in time t + 1 (i.e., Gt+1). We present a global optimization algorithm to effectively infer the temporal latent space using block coordinate gradient descent (BCGD). We further introduce two new variants of BCGD: a local BCGD algorithm and an incremental BCGD algorithm, to scale the inference algorithm to massive networks. Linhong Zhu, Junming Yin, Greg Ver Steeg, Aram Galstyan |
ICDE | 5 |
| 2017 | Sifting Common Information from Many VariablesabstractMeasuring the relationship between any pair of variables is a rich and active area of research that is central to scientific practice. In contrast, characterizing the common information among any group of variables is typically a theoretical exercise with few practical methods for high-dimensional data. A promising solution would be a multivariate generalization of the famous Wyner common information, but this approach relies on solving an apparently intractable optimization problem. We leverage the recently introduced information sieve decomposition to formulate an incremental version of the common information problem that admits a simple fixed point solution, fast convergence, and complexity that is linear in the number of variables. This scalable approach allows us to demonstrate the usefulness of common information in high-dimensional learning problems.The sieve outperforms standard methods on dimensionality reduction tasks, solves a blind source separation problem that cannot be solved with ICA, and accurately recovers structure in brain imaging data. Greg Ver Steeg, Shuyang Gao, Kyle Reing, Aram Galstyan |
IJCAI | 4 |
| 2016 | Extracting Biomolecular Interactions Using Semantic Parsing of Biomedical TextabstractWe advance the state of the art in biomolecular interaction extraction with three contributions: (i) We show that deep, Abstract Meaning Representations (AMR) significantly improve the accuracy of a biomolecular interaction extraction system when compared to a baseline that relies solely on surface- and syntax-based features; (ii) In contrast with previous approaches that infer relations on a sentence-by-sentence basis, we expand our framework to enable consistent predictions over sets of sentences (documents); (iii) We further modify and expand a graph kernel learning framework to enable concurrent exploitation of automatically induced AMR (semantic) and dependency structure (syntactic) representations. Our experiments show that our approach yields interaction extraction systems that are more robust in environments where there is a significant mismatch between training and test conditions. Sahil Garg, Aram Galstyan, Ulf Hermjakob, Daniel Marcu |
AAAI | 2 |
| 2016 | Modeling Concept Dependencies in a Scientific CorpusabstractOur goal is to generate reading lists for students that help them optimally learn technical material.Existing retrieval algorithms return items directly relevant to a query but do not return results to help users read about the concepts supporting their query.This is because the dependency structure of concepts that must be understood before reading material pertaining to a given query is never considered.Here we formulate an information-theoretic view of concept dependency and present methods to construct a "concept graph" automatically from a text corpus.We perform the first human evaluation of concept dependency edges (to be published as open data), and the results verify the feasibility of automatic approaches for inferring concepts and their dependency relations.This result can support search capabilities that may be tuned to help users learn a subject rather than retrieve documents based on a single query. Jonathan Gordon 0001, Linhong Zhu, Aram Galstyan, Premkumar Natarajan, Gully A. P. C. Burns |
ACL (1) | 3 |
| 2016 | The Information SieveabstractWe introduce a new framework for unsupervised learning of representations based on a novel hierarchical decomposition of information. Intuitively, data is passed through a series of progressively fine-grained sieves. Each layer of the sieve recovers a single latent factor that is maximally informative about multivariate dependence in the data. The data is transformed after each pass so that the remaining unexplained information trickles down to the next layer. Ultimately, we are left with a set of latent factors explaining all the dependence in the original data and remainder information consisting of independent noise. We present a practical implementation of this framework for discrete variables and apply it to a variety of fundamental tasks in unsupervised learning including independent component analysis, lossy and lossless compression, and predicting missing values in data. Greg Ver Steeg, Aram Galstyan |
ICML | 2 |
| 2016 | Variational Information Maximization for Feature SelectionabstractFeature selection is one of the most fundamental problems in machine learning. An extensive body of work on information-theoretic feature selection exists which is based on maximizing mutual information between subsets of features and class labels. Practical methods are forced to rely on approximations due to the difficulty of estimating mutual information. We demonstrate that approximations made by existing methods are based on unrealistic assumptions. We formulate a more flexible and general class of assumptions based on variational distributions and use them to tractably generate lower bounds for mutual information. These bounds define a novel information-theoretic framework for feature selection, which we prove to be optimal under tree graphical models with proper choice of variational distributions. Our experiments demonstrate that the proposed method strongly outperforms existing information-theoretic feature selection approaches. Shuyang Gao, Greg Ver Steeg, Aram Galstyan |
NIPS | 3 |
| 2016 | Unsupervised Entity Resolution on Multi-type Graphs
Linhong Zhu, Majid Ghasemi-Gol, Pedro A. Szekely, Aram Galstyan, Craig A. Knoblock |
ISWC (1) | 4 |
| 2016 | Latent Space Model for Multi-Modal Social DataabstractWith the emergence of social networking services, researchers enjoy the increasing availability of large-scale heterogenous datasets capturing online user interactions and behaviors. Traditional analysis of techno-social systems data has focused mainly on describing either the dynamics of social interactions, or the attributes and behaviors of the users. However, overwhelming empirical evidence suggests that the two dimensions affect one another, and therefore they should be jointly modeled and analyzed in a multi-modal framework. The benefits of such an approach include the ability to build better predictive models, leveraging social network information as well as user behavioral signals. To this purpose, here we propose the Constrained Latent Space Model (CLSM), a generalized framework that combines Mixed Membership Stochastic Blockmodels (MMSB) and Latent Dirichlet Allocation (LDA) incorporating a constraint that forces the latent space to concurrently describe the multiple data modalities. We derive an efficient inference algorithm based on Variational Expectation Maximization that has a computational cost linear in the size of the network, thus making it feasible to analyze massive social datasets. We validate the proposed framework on two problems: prediction of social interactions from user attributes and behaviors, and behavior prediction exploiting network information. We perform experiments with a variety of multi-modal social systems, spanning location-based social networks (Gowalla), social media services (Instagram, Orkut), e-commerce and review sites (Amazon, Ciao), and finally citation networks (Cora). The results indicate significant improvement in prediction accuracy over state of the art methods, and demonstrate the flexibility of the proposed approach for addressing a variety of different learning problems commonly occurring with multi-modal social data. Yoon-Sik Cho, Greg Ver Steeg, Emilio Ferrara, Aram Galstyan |
WWW | 4 |
| 2016 | Scalable Temporal Latent Space Inference for Link Prediction in Dynamic Social NetworksabstractWe propose a temporal latent space model for link prediction in dynamic social networks, where the goal is to predict links over time based on a sequence of previous graph snapshots. The model assumes that each user lies in an unobserved latent space, and interactions are more likely to occur between similar users in the latent space representation. In addition, the model allows each user to gradually move its position in the latent space as the network structure evolves over time. We present a global optimization algorithm to effectively infer the temporal latent space. Two alternative optimization algorithms with local and incremental updates are also proposed, allowing the model to scale to larger networks without compromising prediction accuracy. Empirically, we demonstrate that our model, when evaluated on a number of real-world dynamic networks, significantly outperforms existing approaches for temporal link prediction in terms of both scalability and predictive power. Linhong Zhu, Junming Yin, Greg Ver Steeg, Aram Galstyan |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2015 | Efficient Estimation of Mutual Information for Strongly Dependent VariablesabstractWe demonstrate that a popular class of non-parametric mutual information (MI) estimators based on k-nearest-neighbor graphs requires number of samples that scales exponentially with the true MI. Consequently, accurate estimation of MI between two strongly dependent variables is possible only for prohibitively large sample size. This important yet overlooked shortcoming of the existing estimators is due to their implicit reliance on local uniformity of the underlying joint distribution. We introduce a new estimator that is robust to local non-uniformity, works well with limited data, and is able to capture relationship strengths over many orders of magnitude. We demonstrate the superior performance of the proposed estimator on both synthetic and real-world data. Shuyang Gao, Greg Ver Steeg, Aram Galstyan |
AISTATS | 3 |
| 2015 | Maximally Informative Hierarchical Representations of High-Dimensional DataabstractWe consider a set of probabilistic functions of some input variables as a representation of the inputs. We present bounds on how informative a representation is about input data. We extend these bounds to hierarchical representations so that we can quantify the contribution of each layer towards capturing the information in the original data. The special form of these bounds leads to a simple, bottom-up optimization procedure to construct hierarchical representations that are also maximally informative about the data. This optimization has linear computational complexity and constant sample complexity in the number of variables. These results establish a new approach to unsupervised learning of deep representations that is both principled and practical. We demonstrate the usefulness of the approach on both synthetic and real-world data. Greg Ver Steeg, Aram Galstyan |
AISTATS | 2 |
| 2015 | Estimating Mutual Information by Local Gaussian Approximation
Shuyang Gao, Greg Ver Steeg, Aram Galstyan |
UAI | 3 |
| 2014 | Where and Why Users "Check In"abstractThe emergence of location based social network (LBSN) services makes it possible to study individuals’ mobility patterns at a fine-grained level and to see how they are impacted by social factors. In this study we analyze the check-in patterns in LBSN and observe significant temporal clustering of check-in activities. We explore how self-reinforcing behaviors, social factors, and exogenous effects contribute to this clustering and introduce a framework to distinguish these effects at the level of individual check-ins for both users and venues. Using check-in data from three major cities, we show not only that our model can improve prediction of future check-ins, but also that disentangling of different factors allows us to infer meaningful properties of different venues. Yoon-Sik Cho, Greg Ver Steeg, Aram Galstyan |
AAAI | 3 |
| 2014 | Demystifying Information-Theoretic ClusteringabstractWe propose a novel method for clustering data which is grounded in information-theoretic principles and requires no parametric assumptions. Previous attempts to use information theory to define clusters in an assumption-free way are based on maximizing mutual information between data and cluster labels. We demonstrate that this intuition suffers from a fundamental conceptual flaw that causes clustering performance to deteriorate as the amount of data increases. Instead, we return to the axiomatic foundations of information theory to define a meaningful clustering measure based on the notion of consistency under coarse-graining for finite data. Greg Ver Steeg, Aram Galstyan, Fei Sha, Simon DeDeo |
ICML | 2 |
| 2014 | Discovering Structure in High-Dimensional Data Through Correlation Explanation
Greg Ver Steeg, Aram Galstyan |
NIPS | 2 |
| 2014 | Tripartite graph clustering for dynamic sentiment analysis on social mediaabstractThe growing popularity of social media (e.g., Twitter) allows users to easily share information with each other and influence others by expressing their own sentiments on various subjects. In this work, we propose an unsupervised tri-clustering framework, which analyzes both user-level and tweet-level sentiments through co-clustering of a tripartite graph. A compelling feature of the proposed framework is that the quality of sentiment clustering of tweets, users, and features can be mutually improved by joint clustering. We further investigate the evolution of user-level sentiments and latent feature vectors in an online framework and devise an efficient online algorithm to sequentially update the clustering of tweets, users and features with newly arrived data. The online framework not only provides better quality of both dynamic user-level and tweet-level sentiment analysis, but also improves the computational and storage efficiency. We verified the effectiveness and efficiency of the proposed approaches on the November 2012 California ballot Twitter data. Linhong Zhu, Aram Galstyan, James Cheng, Kristina Lerman |
SIGMOD Conference | 2 |
| 2014 | Modeling Temporal Activity Patterns in Dynamic Social NetworksabstractThe focus of this work is on developing probabilistic models for temporal activity of users in social networks (e.g., posting and tweeting) by incorporating the social network influence as perceived by the user. Although prior work in this area has developed sophisticated models for user activity, these models either ignore social network influence completely or incorporate it in an implicit manner. We overcome the nontransparency of the network in the model at the individual scale by proposing a coupled hidden Markov model (HMM), where each user's activity evolves according to a Markov chain with a hidden state that is influenced by the collective activity of the friends of the user. We develop generalized Baum-Welch and Viterbi algorithms for parameter learning and state estimation for the proposed framework. We then validate the proposed model using a significant corpus of user activity on Twitter. Our numerical studies show that with sufficient observations to ensure accurate model learning, the proposed framework explains the observed data better than either a renewal process-based model or a conventional (uncoupled) HMM. We also demonstrate the utility of the proposed approach in predicting the time to the next tweet. Finally, clustering in the model parameter space is shown to result in distinct natural clusters of users characterized by the interaction dynamic between a user and his network. Vasanthan Raghavan, Greg Ver Steeg, Aram Galstyan, Alexander G. Tartakovsky |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2013 | Statistical Tests for Contagion in Observational Social Network StudiesabstractCurrent tests for contagion in social network studies are vulnerable to the confounding effects of latent homophily (i.e., ties form preferentially between individuals with similar hidden traits). We demonstrate a general method to lower bound the strength of causal effects in observational social network studies, even in the presence of arbitrary, unobserved individual traits. Our tests require no parametric assumptions and each test is associated with an algebraic proof. We demonstrate the effectiveness of our approach by correctly deducing the causal effects for examples previously shown to expose defects in existing methodology. Finally, we discuss preliminary results on data taken from the Framingham Heart Study. Greg Ver Steeg, Aram Galstyan |
AISTATS | 2 |
| 2013 | Sentiment Prediction Using Collaborative Filtering
Jihie Kim, Jae-Bong Yoo, Ho Lim, Huida Qiu, Zornitsa Kozareva, Aram Galstyan |
ICWSM | 6 |
| 2013 | Information-theoretic measures of influence based on content dynamicsabstractThe fundamental building block of social influence is for one person to elicit a response in another. Researchers measuring a "response" in social media typically depend either on detailed models of human behavior or on platform-specific cues such as re-tweets, hash tags, URLs, or mentions. Most content on social networks is difficult to model because the modes and motivation of human expression are diverse and incompletely understood. We introduce content transfer, an information-theoretic measure with a predictive interpretation that directly quantifies the strength of the effect of one user's content on another's in a model-free way. Estimating this measure is made possible by combining recent advances in non-parametric entropy estimation with increasingly sophisticated tools for content representation. We demonstrate on Twitter data collected for thousands of users that content transfer is able to capture non-trivial, predictive relationships even for pairs of users not linked in the follower or mention graph. We suggest that this measure makes large quantities of previously under-utilized social media content accessible to rigorous statistical causal analysis. Greg Ver Steeg, Aram Galstyan |
WSDM | 2 |
| 2013 | Continuous strategy replicator dynamics for multi-agent Q-learning
Aram Galstyan |
Auton. Agents Multi Agent Syst. | 1 |
| 2012 | Information transfer in social mediaabstractRecent research has explored the increasingly important role of social media by examining the dynamics of individual and group behavior, characterizing patterns of information diffusion, and identifying influential individuals. In this paper we suggest a measure of causal relationships between nodes based on the information--theoretic notion of transfer entropy, or information transfer. This theoretically grounded measure is based on dynamic information, captures fine--grain notions of influence, and admits a natural, predictive interpretation. Networks inferred by transfer entropy can differ significantly from static friendship networks because most friendship links are not useful for predicting future dynamics. We demonstrate through analysis of synthetic and real-world data that transfer entropy reveals meaningful hidden network structures. In addition to altering our notion of who is influential, transfer entropy allows us to differentiate between weak influence over large groups and strong influence over small groups. Greg Ver Steeg, Aram Galstyan |
WWW | 2 |
| 2011 | Co-Evolution of Selection and Influence in Social NetworksabstractMany networks are complex dynamical systems, where both attributes of nodes and topology of the network (link structure) can change with time. We propose a model of co-evolving networks where both node attributes and network structure evolve under mutual influence. Specifically, we consider a mixed membership stochastic blockmodel, where the probability of observing a link between two nodes depends on their current membership vectors, while those membership vectors themselves evolve in the presence of a link between the nodes. Thus, the network is shaped by the interaction of stochastic processes describing the nodes, while the processes themselves are influenced by the changing network structure. We derive an efficient variational inference procedure for our model, and validate the model on both synthetic and real-world data. Yoon-Sik Cho, Greg Ver Steeg, Aram Galstyan |
AAAI | 3 |
| 2011 | Comparative Analysis of Viterbi Training and Maximum Likelihood Estimation for HMMsabstractWe present an asymptotic analysis of Viterbi Training (VT) and contrast it with a more conventional Maximum Likelihood (ML) approach to parameter estimation in Hidden Markov Models. While ML estimator works by (locally) maximizing the likelihood of the observed data, VT seeks to maximize the probability of the most likely hidden state sequence. We develop an analytical framework based on a generating function formalism and illustrate it on an exactly solvable model of HMM with one unambiguous symbol. For this particular model the ML objective function is continuously degenerate. VT objective, in contrast, is shown to have only finite degeneracy. Furthermore, VT converges faster and results in sparser (simpler) models, thus realizing an automatic Occam's razor for HMM learning. For more general scenario VT can be worse compared to ML but still capable of correctly recovering most of the parameters. Armen E. Allahverdyan, Aram Galstyan |
NIPS | 2 |
| 2011 | A Sequence of Relaxation Constraining Hidden Variable Models
Greg Ver Steeg, Aram Galstyan |
UAI | 2 |
| 2011 | Statistical Mechanics of Semi-Supervised Clustering in Sparse Graphs (Abstract)
Greg Ver Steeg, Aram Galstyan, Armen E. Allahverdyan |
UAI | 2 |
| 2009 | TENTACLES: Self-configuring robotic radio networks in unknown environmentsabstractThis paper presents a bio-inspired, distributed control algorithm calledTENTACLESfor a group of radio robots to move, self-configure and maintain communication between some critical entities (such as humans, command centers, or other systems) in an unknown environment. The basic idea is to direct robots' explorative movements to grow ¿tentacles¿ from entities and establish links when tentacles meet. This approach can self-heal failures of robots and improve communication coverage and quality over time. Experiments in simulations and real robots have shown positive results. Harris Chi Ho Chiu, Bo Ryu, Pedro A. Szekely, Rajiv T. Maheswaran, Craig Milo Rogers, Aram Galstyan, Behnam Salemi, Michael Rubenstein, Wei-Min Shen |
IROS | 7 |
| 2009 | On Maximum a Posteriori Estimation of Hidden Markov Processes
Armen E. Allahverdyan, Aram Galstyan |
UAI | 2 |
| 2007 | Empirical Comparison of "Hard" and "Soft" Label Propagation for Relational Classification
Aram Galstyan, Paul R. Cohen |
ILP | 1 |
| 2006 | Relational Classification Through Three-State Epidemic DynamicsabstractRelational classification in networked data plays an important role in many problems such as text categorization, classification of Web pages, group finding in peer networks, etc. We have previously demonstrated that for a class of label propagating algorithms the underlying dynamics can be modeled as a two-state epidemic process on heterogeneous networks, where infected nodes correspond to classified data instances. We have also suggested a binary classification algorithm that utilizes non-trivial characteristics of epidemic dynamics. In this paper we extend our previous work by considering a three-state epidemic model for label propagation. Specifically, we introduce a new, intermediate state that corresponds to "susceptible" data instances. The utility of the added state is that it allows to control the rates of epidemic spreading, hence making the algorithm more flexible. We show empirically that this extension improves significantly the performance of the algorithm. In particular, we demonstrate that the new algorithm achieves good classification accuracy even for relatively large overlap across the classes Aram Galstyan, Paul R. Cohen |
FUSION | 1 |
| 2006 | Iterative Relational Classification Through Three-State Epidemic Dynamics
Aram Galstyan, Paul R. Cohen |
ISI | 1 |
| 2005 | Inferring Useful Heuristics from the Dynamics of Iterative Relational Classifiers
Aram Galstyan, Paul R. Cohen |
IJCAI | 1 |
| 2005 | Modeling and mathematical analysis of swarms of microscopic robotsabstractThe biologically-inspired swarm paradigm is being used to design self-organizing systems of locally interacting artificial agents. A major difficulty in designing swarms with desired characteristics is understanding the causal relation between individual agent and collective behaviors. Mathematical analysis of swarm dynamics can address this difficulty to gain insight into system design. This paper proposes a framework for mathematical modeling of swarms of microscopic robots that may one day be useful in medical applications. While such devices do not yet exist, the modeling approach can be helpful in identifying various design trade-offs for the robots and be a useful guide for their eventual fabrication. Specifically, we examine microscopic robots that reside in a fluid, for example, a bloodstream, and are able to detect and respond to different chemicals. We present the general mathematical model of a scenario in which robots locate a chemical source. We solve the scenario in one-dimension and show how results can be used to evaluate certain design decisions. Aram Galstyan, Tad Hogg, Kristina Lerman |
SIS | 1 |
| 2005 | Resource Allocation in the Grid with Learning Agents
Aram Galstyan, Karl Czajkowski, Kristina Lerman |
J. Grid Comput. | 1 |
| 2004 | Distributed online localization in sensor networks using a moving targetabstractWe describe a novel method for node localization in a sensor network where there are a fraction of reference nodes with known locations. For application-specific sensor networks, we argue that it makes sense to treat localization through online distributed learning and integrate it with an application task such as target tracking. We propose distributed online algorithm in which sensor nodes use geometric constraints induced by both radio connectivity and sensing to decrease the uncertainty of their position. The sensing constraints, which are caused by a commonly sensed moving target, are usually tighter than connectivity based constraints and lead to a decrease in average localization error over time. Different sensing models, such as radial binary detection and distance-bound estimation, are considered. First, we demonstrate our approach by studying a simple scenario in which a moving beacon broadcasts its own coordinates to the nodes in its vicinity. We then generalize this to the case when instead of a beacon, there is a moving target with a-priori unknown coordinates. The algorithms presented are fully distributed and assume only local information exchange between neighboring nodes. Our results indicate that the proposed method can be used to signicantly enhance the accuracy in position estimation, even when the fraction of reference nodes is small. We compare the efficiency of the distributed algorithms to the case when node positions are estimated using centralized (convex) programming. Finally, simulations using the TinyOS-Nido platform are used to study the performance in more realistic scenarios. Aram Galstyan, Bhaskar Krishnamachari, Kristina Lerman, Sundeep Pattem |
IPSN | 1 |
| 2003 | Macroscopic analysis of adaptive task allocation in robotsabstractWe describe a general mechanism for adaptation in multi-agent systems in which agents modify their behavior in response to changes in the environment or actions of other agents. The agent use memory to estimate the global state of the system from individual observations and adjust their actions accordingly. We present a mathematical model of the dynamics of collective behavior in such systems and apply it to study adaptive task allocation in mobile robots. In this application, the robots task is to forage for red or green pucks. As it travels around the arena, a robot records observations of puck and other robots, and uses these observations to compute the estimated density of each. If it finds there are not enough robots of a specific type, it may switch its foraging state to fill a gap. After a transient, we expect the number of robots in each foraging state to reflect the prevalence of each puck type in the environment. We modelled adaptive task allocation and studied the dynamics of the system for different transition rates between states. We find that for some rates lead to fast convergence times and a steady state solution. Kristina Lerman, Aram Galstyan |
IROS | 2 |
| 2001 | A Macroscopic Analytical Model of Collaboration in Distributed Robotic SystemsabstractIn this article, we present a macroscopic analytical model of collaboration in a group of reactive robots. The model consists of a series of coupled differential equations that describe the dynamics of group behavior. After presenting the general model, we analyze in detail a case study of collaboration, the stick-pulling experiment, studied experimentally and in simulation by Ijspeert et al. [Autonomous Robots, 11, 149-171]. The robots' task is to pull sticks out of their holes, and it can be successfully achieved only through the collaboration of two robots. There is no explicit communication or coordination between the robots. Unlike microscopic simulations (sensor-based or using a probabilistic numerical model), in which computational time scales with the robot group size, the macroscopic model is computationally efficient, because its solutions are independent of robot group size. Analysis reproduces several qualitative conclusions of Ijspeert et al.: namely, the different dynamical regimes for different values of the ratio of robots to sticks, the existence of optimal control parameters that maximize system performance as a function of group size, and the transition from superlinear to sublinear performance as the number of robots is increased. Kristina Lerman, Aram Galstyan, Alcherio Martinoli, Auke Jan Ijspeert |
Artif. Life | 2 |