Robert Sim

dblp:47/1233 · DBLP profile ↗
← Back
48ranked-venue papers
19as first author
17since 2021 · last 2026
0000-0002-0949-791XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 19 first-author · 12 since 2021Systems, architecture and hardware · 14 · 9 first-authorDatabases, data management, data science and information retrieval · 7 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Evaluation Validity in Information Retrieval
Paul Thomas 0001, Nick Craswell, Mark Sanderson, Seth Spielman, Robert Sim, Ryen W. White
SIGIR5
2025 Sweeping Heterogeneity with Smart MoPs: Mixture of Prompts for LLM Task Adaptation
abstract
Prompt instruction tuning is a popular approach to better adjust pretrained LLMs for specific downstream tasks. How to extend this approach to simultaneously handle multiple tasks and data distributions is an interesting question. We propose Mixture of Prompts (MoPs) with smart gating functionality. Our proposed system identifies relevant skills embedded in different groups of prompts and dynamically weighs experts (i.e., collection of prompts) based on the target task. Experiments show that MoPs are resilient to model compression, data source, and task composition, making them highly versatile and applicable in various contexts. In practice, MoPs can simultaneously mitigate prompt training ``interference'' in multi-task, multi-source scenarios (e.g., task and data heterogeneity across sources) and possible implications from model approximations. Empirically, MoPs show particular effectiveness in compressed model scenarios, while maintaining favorable performance in uncompressed settings: MoPs can reduce final perplexity from 9% up to 70% in non-i.i.d. distributed cases and from 3% up to 30% in centralized cases, compared to baselines.
Chen Dun, Mirian Hipolito Garcia, Guoqing Zheng, Ahmed Awadallah 0001, Robert Sim, Anastasios Kyrillidis
AAAI5
2025 Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
abstract
Mixture-of-Experts (MoEs) achieve scalability by dynamically activating subsets of their components. Yet, understanding how expertise emerges through joint training of gating mechanisms and experts remains incomplete, especially in scenarios without clear task partitions. Motivated by inference costs and data heterogeneity, we study how joint training of gating functions and experts can dynamically allocate domain-specific expertise across multiple underlying data distributions. As an outcome of our framework, we develop an instance tailored specifically to decentralized training scenarios, introducing *Dynamically Decentralized Orchestration of MoEs* or *DDOME*. *DDOME* leverages heterogeneity emerging from distributional shifts across decentralized data sources to specialize experts dynamically. By integrating a pretrained common expert to inform a gating function, *DDOME* achieves personalized expert subset selection on-the-fly, facilitating just-in-time personalization. We empirically validate *DDOME* within a Federated Learning (FL) context: *DDOME* attains from 4\% up to an 24\% accuracy improvement over state-of-the-art FL baselines in image and text classification tasks, while maintaining competitive zero-shot generalization capabilities. Furthermore, we provide theoretical insights confirming that the joint gating-experts training is critical for achieving meaningful expert specialization.
Yehya Farhat, Hamza ElMokhtar Shili, Fangshuo Liao, Chen Dun, Mirian Hipolito Garcia, Guoqing Zheng, Ahmed Awadallah 0001, Robert Sim, Dimitrios Dimitriadis, Anastasios Kyrillidis
NeurIPS8
2025 Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
abstract
As the era of autonomous agents making decisions on behalf of users unfolds, ensuring contextual integrity (CI) -- what is the appropriate information to share while carrying out a certain task -- becomes a central question to the field. We posit that CI demands a form of reasoning where the agent needs to reason about the context in which it is operating. To test this, we first prompt LLMs to reason explicitly about CI when deciding what information to disclose. We then extend this approach by developing a reinforcement learning (RL) framework that further instills in models the reasoning necessary to achieve CI. Using a synthetic, automatically created, dataset of only $\sim700$ examples but with diverse contexts and information disclosure norms, we show that our method substantially reduces inappropriate information disclosure while maintaining task performance across multiple model sizes and families. Importantly, improvements transfer from this synthetic dataset to established CI benchmarks such as PrivacyLens that has human annotations and evaluates privacy leakage of AI assistants in actions and tool calls.
Guangchen Lan, Huseyin A. Inan, Sahar Abdelnabi, Janardhan Kulkarni, Lukas Wutschitz, Reza Shokri, Christopher G. Brinton, Robert Sim
NeurIPS8
2024 Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
abstract
Large language models (LLMs) excel in most NLP tasks but also require expensive cloud servers for deployment due to their size, while smaller models that can be deployed on lower cost (e.g., edge) devices, tend to lag behind in terms of response quality. Therefore in this work we propose a hybrid inference approach which combines their respective strengths to save cost and maintain quality. Our approach uses a router that assigns queries to the small or large model based on the predicted query difficulty and the desired quality level. The desired quality level can be tuned dynamically at test time to seamlessly trade quality for cost as per the scenario requirements. In experiments our approach allows us to make up to 40% fewer calls to the large model, with no drop in response quality.
Dujian Ding, Ankur Mallick, Chi Wang 0001, Robert Sim, Subhabrata Mukherjee, Victor Rühle, Laks V. S. Lakshmanan, Ahmed Awadallah 0001
ICLR4
2024 Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation
abstract
We study the problem of in-context learning (ICL) with large language models (LLMs) on private datasets. This scenario poses privacy risks, as LLMs may leak or regurgitate the private examples demonstrated in the prompt. We propose a novel algorithm that generates synthetic few-shot demonstrations from the private dataset with formal differential privacy (DP) guarantees, and show empirically that it can achieve effective ICL. We conduct extensive experiments on standard benchmarks and compare our algorithm with non-private ICL and zero-shot solutions. Our results demonstrate that our algorithm can achieve competitive performance with strong privacy levels. These results open up new possibilities for ICL with privacy protection for a broad range of applications.
Xinyu Tang 0003, Richard Shin, Huseyin A. Inan, Andre Manoel, Niloofar Mireshghallah, Zinan Lin 0001, Sivakanth Gopi, Janardhan Kulkarni, Robert Sim
ICLR9
2024 Privately Aligning Language Models with Reinforcement Learning
abstract
Positioned between pre-training and user deployment, aligning large language models (LLMs) through reinforcement learning (RL) has emerged as a prevailing strategy for training instruction following-models such as ChatGPT. In this work, we initiate the study of privacy-preserving alignment of LLMs through Differential Privacy (DP) in conjunction with RL. Following the influential work of Ziegler et al. (2020), we study two dominant paradigms: (i) alignment via RL without human in the loop (e.g., positive review generation) and (ii) alignment via RL from human feedback (RLHF) (e.g., summarization in a human-preferred way). We give a new DP framework to achieve alignment via RL, and prove its correctness. Our experimental results validate the effectiveness of our approach, offering competitive utility while ensuring strong privacy protections.
Huseyin A. Inan, Arturs Backurs, Varun Chandrasekaran, Janardhan Kulkarni, Robert Sim
ICLR6
2024 TrojanPuzzle: Covertly Poisoning Code-Suggestion Models
abstract
With tools like GitHub Copilot, automatic code suggestion is no longer a dream in software engineering. These tools, based on large language models, are typically trained on massive corpora of code mined from unvetted public sources. As a result, these models are susceptible to data poisoning attacks where an adversary manipulates the model’s training by injecting malicious data. Poisoning attacks could be designed to influence the model’s suggestions at run time for chosen contexts, such as inducing the model into suggesting insecure code payloads. To achieve this, prior attacks explicitly inject the insecure code payload into the training data, making the poison data detectable by static analysis tools that can remove such malicious data from the training set. In this work, we demonstrate two novel attacks, Covert and TrojanPuzzle, that can bypass static analysis by planting malicious poison data in out-of-context regions such as docstrings. Our most novel attack, TrojanPuzzle, goes one step further in generating less suspicious poison data by never explicitly including certain (suspicious) parts of the payload in the poison data, while still inducing a model that suggests the entire payload when completing code (i.e., outside docstrings). This makes TrojanPuzzle robust against signature-based dataset-cleansing methods that can filter out suspicious sequences from the training data. Our evaluation against models of two sizes demonstrates that both Covert and TrojanPuzzle have significant implications for practitioners when selecting code used to train or tune code-suggestion models.
Hojjat Aghakhani, Wei Dai 0007, Andre Manoel, Xavier Fernandes, Anant Kharkar, Christopher Krügel, Giovanni Vigna, David Evans 0001, Benjamin G. Zorn, Robert Sim
SP10
2023 Synthetic Text Generation with Differential Privacy: A Simple and Practical Recipe
abstract
Xiang Yue, Huseyin Inan, Xuechen Li, Girish Kumar, Julia McAnallen, Hoda Shajari, Huan Sun, David Levitan, Robert Sim. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Xiang Yue, Huseyin A. Inan, Julia McAnallen, Hoda Shajari, Huan Sun 0001, David Levitan, Robert Sim
ACL (1)9
2023 Federated Learning Toolkit with Voice-based User Verification Demo
Prathamesh Mandke, Rachel Oberst, Matthias Reisser, Avijit Chakraborty, Christos Louizos, Joseph B. Soriaga, Daniel Madrigal 0001, Andre Manoel, Nalin Singal, Jeff Omhover, Robert Sim
INTERSPEECH11
2023 Analyzing Leakage of Personally Identifiable Information in Language Models
abstract
Language Models (LMs) have been shown to leak information about training data through sentence-level membership inference and reconstruction attacks. Understanding the risk of LMs leaking Personally Identifiable Information (PII) has received less attention, which can be attributed to the false assumption that dataset curation techniques such as scrubbing are sufficient to prevent PII leakage. Scrubbing techniques reduce but do not prevent the risk of PII leakage: in practice scrubbing is imperfect and must balance the trade-off between minimizing disclosure and preserving the utility of the dataset. On the other hand, it is unclear to which extent algorithmic defenses such as differential privacy, designed to guarantee sentence-or user-level privacy, prevent PII disclosure. In this work, we introduce rigorous game-based definitions for three types of PII leakage via black-box extraction, inference, and reconstruction attacks with only API access to an LM. We empirically evaluate the attacks against GPT-2 models fine-tuned with and without defenses in three domains: case law, health care, and e-mails. Our main contributions are (i) novel attacks that can extract up to 10× more PII sequences than existing attacks, (ii) showing that sentence-level differential privacy reduces the risk of PII disclosure but still leaks about 3% of PII sequences, and (iii) a subtle connection between record-level membership inference and PII reconstruction. Code to reproduce all experiments in the paper is available at https://github.com/microsoft/analysing_pii_leakage.
Nils Lukas, Ahmed Salem 0001, Robert Sim, Shruti Tople, Lukas Wutschitz, Santiago Zanella-Béguelin
SP3
2022 Heterogeneous Ensemble Knowledge Transfer for Training Large Models in Federated Learning
abstract
Federated learning (FL) enables edge-devices to collaboratively learn a model without disclosing their private data to a central aggregating server. Most existing FL algorithms require models of identical architecture to be deployed across the clients and server, making it infeasible to train large models due to clients' limited system resources. In this work, we propose a novel ensemble knowledge transfer method named Fed-ET in which small models (different in architecture) are trained on clients, and used to train a larger model at the server. Unlike in conventional ensemble learning, in FL the ensemble can be trained on clients' highly heterogeneous data. Cognizant of this property, Fed-ET uses a weighted consensus distillation scheme with diversity regularization that efficiently extracts reliable consensus from the ensemble while improving generalization by exploiting the diversity within the ensemble. We show the generalization bound for the ensemble of weighted models trained on heterogeneous datasets that supports the intuition of Fed-ET. Our experiments on image and language tasks show that Fed-ET significantly outperforms other state-of-the-art FL algorithms with fewer communicated parameters, and is also robust against high data-heterogeneity.
Yae Jee Cho, Andre Manoel, Gauri Joshi, Robert Sim, Dimitrios Dimitriadis
IJCAI4
2022 UserIdentifier: Implicit User Representations for Simple and Effective Personalized Sentiment Analysis
abstract
Fatemehsadat Mireshghallah, Vaishnavi Shrivastava, Milad Shokouhi, Taylor Berg-Kirkpatrick, Robert Sim, Dimitrios Dimitriadis. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Niloofar Mireshghallah, Vaishnavi Shrivastava, Milad Shokouhi, Taylor Berg-Kirkpatrick, Robert Sim, Dimitrios Dimitriadis
NAACL-HLT5
2022 Understanding Questions that Arise When Working with Business Documents
abstract
While digital assistants are increasingly used to help with various productivity tasks, less attention has been paid to employing them in the domain of business documents. To build an agent that can handle users' information needs in this domain, we must first understand the types of assistance that users desire when working on their documents. In this work, we present results from two user studies that characterize the information needs and queries of authors, reviewers, and readers of business documents. In the first study, we used experience sampling to collect users' questions in-situ as they were working with their documents, and in the second, we built a human-in-the-loop document Q&A system which rendered assistance with a variety of users' questions. Our results have implications for the design of document assistants that complement AI with human intelligence including whether particular skillsets or roles within the document are needed from human respondents, as well as the challenges around such systems.
Farnaz Jahanbakhsh, Elnaz Nouri, Robert Sim, Ryen W. White, Adam Fourney
Proc. ACM Hum. Comput. Interact.3
2021 Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets
abstract
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, Hanna Wallach. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, Hanna M. Wallach
ACL/IJCNLP (1)4
2021 MeetingCoach: An Intelligent Dashboard for Supporting Effective & Inclusive Meetings
abstract
Video-conferencing is essential for many companies, but its limitations in conveying social cues can lead to ineffective meetings. We present MeetingCoach, an intelligent post-meeting feedback dashboard that summarizes contextual and behavioral meeting information. Through an exploratory survey (N=120), we identified important signals (e.g., turn taking, sentiment) and used these insights to create a wireframe dashboard. The design was evaluated with in situ participants (N=16) who helped identify the components they would prefer in a post-meeting dashboard. After recording video-conferencing meetings of eight teams over four weeks, we developed an AI system to quantify the meeting features and created personalized dashboards for each participant. Through interviews and surveys (N=23), we found that reviewing the dashboard helped improve attendees’ awareness of meeting dynamics, with implications for improved effectiveness and inclusivity. Based on our findings, we provide suggestions for future feedback system designs of video-conferencing meetings.
Samiha Samrose, Daniel McDuff, Robert Sim, Jina Suh, Kael Rowan, Javier Hernandez, Sean Rintel, Kevin Moynihan, Mary Czerwinski
CHI3
2021 Privacy Regularization: Joint Privacy-Utility Optimization in LanguageModels
abstract
Fatemehsadat Mireshghallah, Huseyin Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, Robert Sim. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Niloofar Mireshghallah, Huseyin A. Inan, Marcello Hasegawa, Victor Rühle, Taylor Berg-Kirkpatrick, Robert Sim
NAACL-HLT6
2020 Conversations with Documents: An Exploration of Document-Centered Assistance
abstract
The role of conversational assistants has become more prevalent in helping people increase their productivity. Document-centered assistance, for example to help an individual quickly review a document, has seen less significant progress, even though it has the potential to tremendously increase a user's productivity. This type of document-centered assistance is the focus of this paper. Our contributions are three-fold: (1) We first present a survey to understand the space of document-centered assistance and the capabilities people expect in this scenario. (2) We investigate the types of queries that users will pose while seeking assistance with documents, and show that document-centered questions form the majority of these queries. (3) We present a set of initial machine learned models that show that (a) we can accurately detect document-centered questions, and (b) we can build reasonably accurate models for answering such questions. These positive results are encouraging, and suggest that even greater results may be attained with continued study of this interesting and novel problem space. Our findings have implications for the design of intelligent systems to support task completion via natural interactions with documents.
Maartje ter Hoeve, Robert Sim, Elnaz Nouri, Adam Fourney, Maarten de Rijke, Ryen W. White
CHIIR2
2020 Step-wise Recommendation for Complex Task Support
abstract
Digital assistants help people perform simple tasks, including scheduling, home automation, information look up, and question answering. Current assistants offer less support for accomplishing more complex tasks comprising multiple steps. These tasks sometimes require the ability to leverage the capabilities of multiple devices. Ideal smart assistants for such complex scenarios would track task progress and recommend useful and appropriate information at every step of the procedure. To this end, we introduce the novel notion of step-wise recommendation as a means of automatically providing guidance relevant to the current step in a complex task. We employ a common real-life scenario for this purpose: recipe preparation. We demonstrate how a smart assistant can be developed to offer support for complex tasks enabled by multi-modal inputs and outputs through multiple devices (i.e., smart speakers, tablets, or other smart devices available in kitchens). We develop step-wise recommendation models for this scenario and analyze their efficacy for: (1) different prediction tasks (e.g., resources, devices), and (2) different contextual information used to make the prediction (e.g., completed steps, current step, and importantly, future steps). Our recommendation model achieves a prediction accuracy of 83-96%, depending on the prediction task and context used. The findings have implications for the design of intelligent systems to help people accomplish complex tasks.
Elnaz Nouri, Robert Sim, Adam Fourney, Ryen W. White
CHIIR2
2020 Proactive Suggestion Generation: Data and Methods for Stepwise Task Assistance
abstract
Conversational systems such as digital assistants can help users per-form many simple tasks upon request. Looking to the future, these systems will also need to fully support more complex, multi-step tasks (e.g., following cooking instructions), and help users complete those tasks, e.g., via useful and relevant suggestions made during the process. This paper takes the first step towards automatic generation of task-related suggestions. We introduce proactive suggestion generation as a novel task of natural language generation, in which a decision is made to inject a suggestion into an ongoing user dialog and one is then automatically generated. We propose two types of stepwise suggestions: multiple-choice response generation and text generation. We provide several models for each type of suggestion, including binary and multi-class classification, and text generation.
Elnaz Nouri, Robert Sim, Adam Fourney, Ryen W. White
SIGIR2
2020 Toward Activity Discovery in the Personal Web
abstract
Individuals' personal information collections (their emails, files, appointments, web searches, contacts, etc) offer a wealth of insights into the organization and structure of their everyday lives. In this paper we address the task of learning representations of personal information items to capture individuals' ongoing activities, such as projects and tasks: Such representations can be used in activity-centric applications like personal assistants, email clients, and productivity tools to help people better manage their data and time. We propose a graph-based approach that leverages the inherent interconnected structure of personal information collections, and derive efficient, exact techniques to incrementally update representations as new data arrive. We demonstrate the strengths of our graph-based representations against competitive baselines in a novel intrinsic rating task and an extrinsic recommendation task.
Tara Safavi, Adam Fourney, Robert Sim, Marcin Juraszek, Shane Williams, Ned Friend, Danai Koutra, Paul N. Bennett
WSDM3
2019 Task Completion Detection: A Study in the Context of Intelligent Systems
abstract
People can record their pending tasks using to-do lists, digital assistants, and other task management software. In doing so, users of these systems face at least two challenges: (1) they must manually mark their tasks as complete, and (2) when systems proactively remind them about their pending tasks, say, via interruptive notifications, they lack information on task completion status. As a result, people may not realize the full benefits of to-do lists (since these lists can contain both completed and pending tasks) and they may be reminded about tasks they have already done (wasting time and causing frustration). In this paper, we present methods to automatically detect task completion. These inferences can be used to deprecate completed tasks and/or suppress notifications for these tasks (or for other purposes, e.g., task prioritization). Using log data from a popular digital assistant, we analyze temporal dynamics in the completion of tasks and train machine-learned models to detect completion with accuracy exceeding 80% using a variety of features (time elapsed since task creation, task content, email, notifications, user history). The findings have implications for the design of intelligent systems to help people manage their tasks.
Ryen W. White, Ahmed Awadallah 0001, Robert Sim
SIGIR3
2019 Domain Adaptation for Commitment Detection in Email
abstract
People often make commitments to perform future actions. Detecting commitments made in email (e.g., "I'll send the report by end of day'') enables digital assistants to help their users recall promises they have made and assist them in meeting those promises in a timely manner. In this paper, we show that commitments can be reliably extracted from emails when models are trained and evaluated on the same domain (corpus). However, their performance degrades when the evaluation domain differs. This illustrates the domain bias associated with email datasets and a need for more robust and generalizable models for commitment detection. To learn a domain-independent commitment model, we first characterize the differences between domains (email corpora) and then use this characterization to transfer knowledge between them. We investigate the performance of domain adaptation, namely transfer learning, at different granularities: feature-level adaptation and sample-level adaptation. We extend this further using a neural autoencoder trained to learn a domain-independent representation for training samples. We show that transfer learning can help remove domain bias to obtain models with less domain dependence. Overall, our results show that domain differences can have a significant negative impact on the quality of commitment detection models and that transfer learning has enormous potential to address this issue.
Hosein Azarbonyad, Robert Sim, Ryen W. White
WSDM2
2009 Autonomous vision-based robotic exploration and mapping using hybrid maps and particle filters
Robert Sim, James J. Little
Image Vis. Comput.1
2009 Second and Third Canadian Conferences on Computer and Robot Vision
Robert Sim, Greg Mori, Ioannis M. Rekleitis
Image Vis. Comput.1
2007 A Study of the Rao-Blackwellised Particle Filter for Efficient and Accurate Vision-Based SLAM
Robert Sim, Pantelis Elinas, James J. Little
Int. J. Comput. Vis.1
2006 σSLAM: Stereo Vision SLAM using the Rao-Blackwellised Particle Filter and a Novel Mixture Proposal Distribution
abstract
We consider the problem of simultaneous localization and mapping (SLAM) using the Rao-Blackwellised particle filter (RBPF) for the class of indoor mobile robots equipped only with stereo vision. Our goal is to construct dense metric maps of natural 3D point landmarks for large cyclic environments in the absence of accurate landmark position measurements and motion estimates. Our work differs from other approaches because landmark estimates are derived from stereo vision and motion estimates are based on sparse optical flow. We distinguish between landmarks using the scale invariant feature transform (SIFT). This is in contrast to current popular approaches that rely on reliable motion models derived from odometric hardware and accurate landmark measurements obtained with laser sensors. Since our approach depends on a particle filter whose main component is the proposal distribution, we develop and evaluate a novel mixture proposal distribution that allows us to robustly close large loops. We validate our approach experimentally for long camera trajectories processing thousands of images at reasonable frame rates
Pantelis Elinas, Robert Sim, James J. Little
ICRA2
2006 Autonomous vision-based exploration and mapping using hybrid maps and Rao-Blackwellised particle filters
abstract
This paper addresses the problem of exploring and mapping an unknown environment using a robot equipped with a stereo vision sensor. The main contribution of our work is a fully automatic mapping system that operates without the use of active ranger sensors (such as laser or sonic transducers), can operate in real-time and can consistently produce accurate maps of large-scale environments. Our approach implements a Rao-Blackwellised particle filter (RBPF) to solve the simultaneous localization and mapping problem and uses efficient data structures for real-time data association, mapping, and spatial reasoning. We employ a hybrid map representation that infers 3D point landmarks from image features to achieve precise localization, coupled with occupancy grids for safe navigation. This paper describes our framework and implementation, and presents our exploration method, and experimental results illustrating the functionality of the system
Robert Sim, James J. Little
IROS1
2006 Landmark Selection for Vision-Based Navigation
abstract
Recent work in the object recognition community has yielded a class of interest-point-based features that are stable under significant changes in scale, viewpoint, and illumination, making them ideally suited to landmark-based navigation. Although many such features may be visible in a given view of the robot's environment, only a few such features are necessary to estimate the robot's position and orientation. In this paper, we address the problem of automatically selecting, from the entire set of features visible in the robot's environment, the minimum (optimal) set by which the robot can navigate its environment. Specifically, we decompose the world into a small number of maximally sized regions, such that at each position in a given region, the same small set of features is visible. We introduce a novel graph theoretic formulation of the problem, and prove that it is NP-complete. Next, we introduce a number of approximation algorithms and evaluate them on both synthetic and real data. Finally, we use the decompositions from the real image data to measure the localization performance versus the undecomposed map
Pablo Sala, Robert Sim, Ali Shokoufandeh, Sven J. Dickinson
IEEE Trans. Robotics2
2005 Stable Exploration for Bearings-only SLAM
abstract
Recent work on robotic exploration and active sensing has examined a variety of information-theoretic approaches to efficient and convergent map construction. These involve moving an exploring robot to locations in the world where the anticipated information gain is maximized. In this paper we demonstrate that, for map construction using bearings-only information and the Extended Kalman Filter (EKF), driving exploration so as to maximize expected information gain leads to ill-conditioned filter updates and a high probability of divergence between the inferred map and reality. In particular, we present analytical and numerical results demonstrating the effects of blindly applying an information-theoretic approach to bearings-only exploration. Subsequently, we present experimental results demonstrating that an exploration approach that favours the conditioning of the filter update will lead to more accurate maps.
Robert Sim
ICRA1
2005 Global A-Optimal Robot Exploration in SLAM
abstract
It is well-known that the Kalman filter for simultaneous localization and mapping (SLAM) converges to a fully correlated map in the limit of infinite time and data [1]. However, the rate of convergence of the map has a strong dependence on the order of the observations. We show that conventional exploration algorithms for collecting map data are sub-optimal in both the objective function and choice of optimization procedure. We show that optimizing the a-optimal information measure results in a more accurate map than existing approaches, using a greedy, closed-loop strategy. Secondly, we demonstrate that by restricting the planning to an appropriate policy class, we can tractably find non-greedy, global planning trajectories that produce more accurate maps, explicitly planning to close loops even in open-loop scenarios.
Robert Sim, Nicholas Roy
ICRA1
2005 Stabilizing information-driven exploration for bearings-only SLAM using range gating
abstract
This paper examines the problem of information-driven exploration for the purposes of simultaneous localization and mapping (SLAM) with a bearings-only sensor. In another work, we have demonstrated that employing an information-driven approach to exploration with an extended Kalman filter (EKF) can drive the robot to locations in the world where filter updates are ill-conditioned and linearization constraints are violated, potentially destabilizing the filter, and increasing the probability of divergence from the true state estimate. In this paper, we demonstrate an information-driven approach to exploration that preserves the stability of the EKF and produces maps that are significantly more accurate than a conventional information-driven approach. Our method is based on range-gating observations so as to avoid potentially destabilizing updates. We provide simulated experimental results demonstrating the superior performance of our approach over simple outlier gating and over heuristic-driven exploration.
Robert Sim
IROS1
2004 Self-Organizing Visual Maps
Robert Sim, Gregory Dudek
AAAI1
2004 Online Control Policy Optimization for Minimizing Map Uncertainty during Exploration
abstract
Tremendous progress has been made recently in simultaneous localization and mapping of unknown environments. Using sensor and odometry data from an exploring mobile robot, it has become much easier to build high-quality globally consistent maps of many large, real-world environments. To date, however, relatively little attention has been paid to the controllers used to build these maps. Existing exploration strategies usually attempt to cover the largest amount of unknown space as quickly as possible. Few strategies exist for building the most reliable map possible, but the particular control strategy can have a substantial impact on the quality of the resulting map. In this paper, we devise a control algorithm for exploring unknown space that explicitly tries to build as large a map as possible while maintaining as accurate a map as possible. We make use of a parameterized class of spiral trajectory policies, choosing a new parameter setting at every time step to maximize the expected reward of the policy. We do this in the context of building a visual map of an unknown environment, and show that our strategy leads to a higher accuracy map faster than other candidate controllers, including any single choice in our policy class.
Robert Sim, Gregory Dudek, Nicholas Roy
ICRA1
2004 AQUA: an aquatic walking robot
abstract
This paper describes an underwater walking robotic system being developed under the name AQUA, the goals of the AQUA project, the overall hardware and software design, the basic hardware and sensor packages that have been developed, and some initial experiments. The robot is based on the RHex hexapod robot and uses a suite of sensing technologies, primarily based on computer vision and INS, to allow it to navigate and map clear shallow-water environments. The sensor-based navigation and mapping algorithms are based on the use of both artificial floating visual and acoustic landmarks as well as on naturally occurring underwater landmarks and trinocular stereo.
Christina Georgiades, Andrew German, Andrew Hogue, Chris Prahacs, Arlene Ripsman, Robert Sim, Luz Abril Torres-Méndez, Pifu Zhang, Martin Buehler, Gregory Dudek, Michael R. M. Jenkin, Evangelos E. Milios
IROS7
2004 Landmark selection for vision-based navigation
abstract
Recent work in the object recognition community has yielded a class of interest point-based features that are stable under significant changes in scale, viewpoint, and illumination, making them ideally suited to landmark-based navigation. Although many such features may be visible in a given view of the robot's environment, only a few such features are necessary to estimate the robot's position and orientation. In this paper, we address the problem of automatically selecting, from the entire set of features visible in the robot's environment, the minimum (optimal) set by which the robot can navigate its environment. Specifically, we decompose the world into a small number of maximally sized regions such that at each position in a given region, the same small set of features is visible. We introduce a novel graph theoretic formulation of the problem and prove that it is NP-complete. Next, we introduce a number of approximation algorithms and evaluate them on both synthetic and real data.
Pablo Sala, Robert Sim, Ali Shokoufandeh, Sven J. Dickinson
IROS2
2004 Learning generative models of invariant features
abstract
We present a method for learning a set of models of visual features which are invariant to scale and translation in the image domain. The models are constructed by first applying the scale-invariant feature transform (SIFT) to a set of training images, and matching the extracted features across the images, followed by learning the pose-dependent behavior of the features. The modeling process avoids assumptions with respect to scene and imaging geometry, but rather learns the direct mapping from camera pose to feature observation. Such models are useful for applications to robotic tasks, such as localization, as well as visualization tasks. We present the model learning framework, and experimental results illustrating the success of the method for learning models that are useful for robot localization.
Robert Sim, Gregory Dudek
IROS1
2004 Learning Generative Models of Scene Features
Robert Sim, Gregory Dudek
Int. J. Comput. Vis.1
2003 Robodaemon -a device independent, network-oriented, modular mobile robot controller
abstract
We discuss a software environment for multi-robot, multi-platform mobile robot control and simulation. Like others, we have observed that mobile robotics research is greatly facilitated by the availability of a suitable simulator for both vehicle kinematics as well as sensing, and have created an environment that permits this while allowing a large measure of device independence. By using a multiprocessor internet-based architecture, our platform permits multiple users to use a variety of programming interfaces (visual, script-based or various application programming interfaces (API's)) to rapidly prototype methods to control multiple heterogeneous robots both in simulation and in real-world settings. We present an overview of our architecture and discuss its future directions.
Gregory Dudek, Robert Sim
ICRA2
2003 Comparing image-based localization methods
Robert Sim, Gregory Dudek
IJCAI1
2003 Effective exploration strategies for the construction of visual maps
abstract
We consider the effect of exploration policy in the context of the autonomous construction of a visual map of an unknown environment. Like other concurrent mapping and localization (CML) tasks, odometric uncertainty poses the problem of introducing distortions into the map which are difficult to correct without costly on-line or post-processing algorithms. Our problem is further compounded by the implicit nature of the visual map representation, which is designed to accommodate a wide variety of visual phenomena without assuming a particular imaging platform, thereby precluding the inference of scene geometry. Such a representation presents a requirement for a relatively dense sampling of observations of the environment in order to produce reliable models. Our goal is to develop an online policy for exploring an unknown environment which minimizes map distortion while maximizing coverage. We do not depend on costly post-hoc expectation maximization approaches to improve the output, but rather employ extended Kalman filter (EKF) methods to localize each observation once, and rely on the exploration policy to ensure that sufficient information is available to localize the successive observations. We present an experimental analysis of a variety of exploratory policies, in both simulated and real environments, and demonstrate that with an effective policy an accurate map can be constructed.
Robert Sim, Gregory Dudek
IROS1
2001 Learning Generative Models of Scene Features
abstract
We present a method for learning a set of generative models which are suitable for representing variations of selected image-domain features of the scene as a function of changes in the camera viewpoint. Such models are important for robotic tasks, such as probabilistic position estimation (i.e. localization), as well as visualization. Our approach entails the selection of image-domain features, as well as the synthesis of models of their visual behavior. The model we propose is capable of generating maximum likelihood views of automatically selected features, as well as a measure of the likelihood of a particular view from a particular camera position. Training the models involves regularizing observations of the features from known camera locations. The uncertainty of the model is evaluated using cross validation. The features themselves are initially selected automatically as salient points by a measure of visual attention, and are tracked across multiple views. While the motivation for this work is for robot localization, the results have implications for image interpolation, virtual scene reconstruction and object recognition. The paper presents a formulation of the problem and illustrative experimental results.
Robert Sim, Gregory Dudek
CVPR (1)1
2001 Collaborative exploration for the construction of visual maps
abstract
We examine the problem of learning a visual map of the environment while maintaining an accurate pose estimate. Our approach is based on using two robots in a simple collaborative scheme. Without outside information, as a robot collects training images, its position estimate accumulates errors, thus corrupting its knowledge of the positions from which observations are taken. We address this problem by deploying a second robot to observe the first one as it explores, thereby establishing a virtual tether, and enabling an accurate estimate of the robot's position while it constructs the map. We refer to this process as cooperative localization. The images collected during this process are assembled into a representation that allows vision-based position estimation from a single image at a later date. In addition to developing a formalism and concept, we validate our results experimentally and present quantitative results demonstrating the performance of the method in over 90 trials.
Ioannis M. Rekleitis, Robert Sim, Gregory Dudek, Evangelos E. Milios
IROS2
2001 Learning environmental features for pose estimation
Robert Sim, Gregory Dudek
Image Vis. Comput.1
2000 RoboMutts
Robert Sim, Paul Russo, Andrew Grahm, Andrew Blair, Nick Barnes, Alan Blair 0001
RoboCup1
1999 Learning and Evaluating Visual Features for Pose Estimation
abstract
We present a method for learning a set of visual landmarks which are useful for pose estimation. The landmark learning mechanism is designed to be applicable to a wide range of environments, and generalized for different approaches to computing a pose estimate. Initially, each landmark is detected as a focal extremum of a measure of distinctiveness and represented by a principal components encoding which is exploited for matching. Attributes of the observed landmarks can be parameterized using a generic parameterization method and then evaluated in terms of their utility for pose estimation. We present experimental evidence that demonstrates the utility of the method.
Robert Sim, Gregory Dudek
ICCV1
1999 Learning Visual Landmarks for Pose Estimation
abstract
We present an approach to vision-based mobile robot localization, even without an a-priori pose estimate. This is accomplished by learning a set of visual features called image-domain landmarks. The landmark learning mechanism is designed to be applicable to a wide range of environments. Each landmark is detected as a focal extremum of a measure of uniqueness and represented by an appearance-based encoding. Localization is performed using a method that matches observed landmarks to learned prototypes and generates independent position estimates for each match. The independent estimates are then combined to obtain a final position estimate, with an associated uncertainty. Quantitative experimental evidence is presented that demonstrates that accurate pose estimates can be obtained, despite changes to the environment.
Robert Sim, Gregory Dudek
ICRA1
1998 Mobile robot localization from learned landmarks
abstract
Presents an approach to vision-based mobile robot localization. In an attempt to capitalize on the benefits of both image and landmark-based methods, we describe a method that combines their strengths. Images are encoded as a set of visual features called landmarks. Potential landmarks are detected using an attention mechanism implemented as a measure of uniqueness. They are then selected and represented by an appearance-based encoding. Localization is performed using a landmark tracking and interpolation method which obtains an estimate accurate to a fraction of the environment sampling density. Experimental results are shown to confirm the feasibility and accuracy of the method.
Robert Sim, Gregory Dudek
IROS1