VLDB 2026 Research / reviewers in the wild / expert
Mark Roberts
dblp:75/5516
· DBLP profile ↗
24ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-2690-7658ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 first-authorSecurity and privacy · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Probabilistic Hierarchical Goal Network Planning with UCTabstractHierarchical goal networks (HGNs) provide a framework for goal-directed planning by decomposing high-level goals into ordered subgoals. While prior work has examined non-determinism for hierarchical planning (specifically, HTNs), scant work studies how HGNs can help in stochastic settings. We introduce a formalism for probabilistic HGN planning with action-insertion semantics, enabling probabilistic planners to incorporate domain knowledge from goal decomposition methods. We design and evaluate two UCT-based algorithms for solving probabilistic HGN planning problems: an asymptotically optimal approach and a compressed, shared-value approach that optimizes separately for each goal within the goal-subgoal hierarchy. We compare our two UCT-based HGN search algorithms experimentally on modified benchmark domains from the FOND HTN literature. Our results demonstrate that on larger problems, the compressed search converges more quickly and outperforms the asymptotically optimal search. This suggests that HGNs can be effective in probabilistic planning, and compression may yield better performance on large problems in anytime settings with stochastic action outcomes. David H. Chan, Mark Roberts, Dana S. Nau |
AAAI | 2 |
| 2025 | Landmark-Assisted Monte Carlo PlanningabstractLandmarks—conditions that must be satisfied at some point in every solution plan—have contributed to major advancements in classical planning, but they have seldom been used in stochastic domains. We formalize probabilistic landmarks and adapt the UCT algorithm to leverage them as subgoals to decompose MDPs; core to the adaptation is balancing between greedy landmark achievement and final goal achievement. Our results in benchmark domains show that well-chosen landmarks can significantly improve the performance of UCT in online probabilistic planning, while the best balance of greedy versus long-term goal achievement is problem-dependent. The results suggest that landmarks can provide helpful guidance for anytime algorithms solving MDPs. David H. Chan, Mark Roberts, Dana S. Nau |
ECAI | 2 |
| 2025 | An Interaction Specification Language for Robot Application DevelopmentabstractRobot programming languages that represent tasks as graph structures are both popular and accessible among programming novices and experts. However, these languages are largely decoupled from robots' automated task planning capabilities, rendering developers unable to explicitly leverage their robot's ability to plan its own actions. We thereby created the Interaction Specification Language (ISL), which enables de-velopers to import and apply elements from a robot planning do-main in a graph-based programming paradigm. For developers, ISL provides flexibility in the reliance on automated planning. For researchers, the release of our open-source ISL lexer and parser is intended to promote standardization and test-driven development. We additionally provide a metric by which ISL programs can be evaluated. David Porfirio, Mark Roberts, Laura M. Hiatt |
HRI | 2 |
| 2025 | HTN Plan Repair Algorithms Compared: Strengths and Weaknesses of Different MethodsabstractThis paper provides theoretical and empirical comparisons of three recent hierarchical plan repair algorithms: SHOPFIXER, IPYHOPPER, and REWRITE. Our theoretical results show that the three algorithms correspond to three different definitions of the plan repair problem, leading to differences in the algorithms’ search spaces, the repair problems they can solve, and the kinds of repairs they can make. Understanding these distinctions is important when choosing a repair method for any given application. Building on the theoretical results, we evaluate the algorithms empirically in a series of benchmark planning problems. Our empirical results provide more detailed insight into the run- time repair performance of these systems and the coverage of the repair problems solved, based on algorithmic properties such as replanning, chronological backtracking, and back- jumping over plan trees. Paul Zaidins, Robert P. Goldman, Ugur Kuter, Dana S. Nau, Mark Roberts |
ICAPS | 5 |
| 2025 | Automating Curriculum Learning for Reinforcement Learning using a Skill-Based Bayesian Network
Vincent Hsiao, Mark Roberts, Laura M. Hiatt, George Dimitri Konidaris, Dana S. Nau |
AAMAS | 2 |
| 2025 | Uncertainty Expression for Human-Robot Task Communication
David Porfirio, Mark Roberts, Laura M. Hiatt |
AAMAS | 2 |
| 2025 | ToMCAT: Benchmark for Socially Assistive Robots with Theory of Mind of Children Assembling Tangram PuzzlesabstractAssistive robots will be more effective if they can accurately reason about the intentions and beliefs of the user (i.e., have Theory of Mind (ToM)). ToM benchmarks allow us to examine how well an artificial agent (e.g., robot) is able to do ToM reasoning in a given scenario. However, there is a need for ToM benchmarks that are more representative of the challenges faced in assistive robotics. Existing benchmarks from AI and HRI make simplifying assumptions, such as simply defined goals, plans that are indicative of goals, and no user errors. To address the challenges from relaxing these assumptions, we propose the Theory of Mind of Children Assembling Tangrams (ToMCAT) dataset. The data is derived from videos of children building tangram puzzles while being assisted by a social robot. As a baseline benchmark, we evaluated two approaches for how well they can recognize which puzzle this child is building based on a single observation. Analogical reasoning correctly recognized the puzzle more than 75% of the time and had perfect accuracy for puzzle states that were close to complete. However, an out-of-the-box commercial LLM correctly recognized the puzzle only 60% of the time and was accurate on less than 80% of the completed puzzles. Our results suggest that the ToMCAT dataset offers challenges for recognizing the intended puzzle of a child. Furthermore, the dataset provides opportunities to examine additional ToM reasoning capabilities. Overall, the ToMCAT dataset provides a useful benchmark to facilitate the advancement of ToM reasoning for assistive robotics. Jason R. Wilson, Irina Rabkina, Mark Roberts, Laura M. Hiatt |
RO-MAN | 3 |
| 2024 | Goal-Oriented End-User Programming of RobotsabstractEnd-user programming (EUP) tools must balance user control with the robot's ability to plan and act autonomously. Many existing task-oriented EUP tools enforce a specific level of control, e.g., by requiring that users hand-craft detailed sequences of actions, rather than offering users the flexibility to choose the level of task detail they wish to express. We thereby created a novel EUP system, Polaris, that in contrast to most existing EUP tools, uses goal predicates as the fundamental building block of programs. Users can thereby express high-level robot objectives or lower-level checkpoints at their choosing, while an off-the-shelf task planner fills in any remaining program detail. To ensure that goal-specified programs adhere to user expectations of robot behavior, Polaris is equipped with a Plan Visualizer that exposes the planner's output to the user before runtime. In what follows, we describe our design of Polaris and its evaluation with 32 human participants. Our results support the Plan Visualizer's ability to help users craft higher-quality programs. Furthermore, there are strong associations between user perception of the robot and Plan Visualizer usage, and evidence that robot familiarity has a key role in shaping user experience. David Porfirio, Mark Roberts, Laura M. Hiatt |
HRI | 2 |
| 2024 | Satellite Resource Scheduling: Compaction Strategies for Genetic Algorithm Schedulers
L. Darrell Whitley, Ozeas Quevedo de Carvalho, Mark Roberts, Vivint Shetty, Piyabutra Jampathom |
PPSN (4) | 3 |
| 2024 | Understanding CNN fragility when learning with imbalanced dataabstractAbstract Convolutional neural networks (CNNs) have achieved impressive results on imbalanced image data, but they still have difficulty generalizing to minority classes and their decisions are difficult to interpret. These problems are related because the method by which CNNs generalize to minority classes, which requires improvement, is wrapped in a black-box. To demystify CNN decisions on imbalanced data, we focus on their latent features. Although CNNs embed the pattern knowledge learned from a training set in model parameters, the effect of this knowledge is contained in feature and classification embeddings ( FE and CE ). These embeddings can be extracted from a trained model and their global, class properties (e.g., frequency, magnitude and identity) can be analyzed. We find that important information regarding the ability of a neural network to generalize to minority classes resides in the class top-K CE and FE . We show that a CNN learns a limited number of class top-K CE per category, and that their magnitudes vary based on whether the same class is balanced or imbalanced. We hypothesize that latent class diversity is as important as the number of class examples, which has important implications for re-sampling and cost-sensitive methods. These methods generally focus on rebalancing model weights, class numbers and margins; instead of diversifying class latent features. We also demonstrate that a CNN has difficulty generalizing to test data if the magnitude of its top-K latent features do not match the training set. We use three popular image datasets and two cost-sensitive algorithms commonly employed in imbalanced learning for our experiments. Damien Dablain, Kristen N. Jacobson, Colin Bellinger, Mark Roberts, Nitesh V. Chawla |
Mach. Learn. | 4 |
| 2023 | Scheduling Multi-Resource Satellites using Genetic Algorithms and Permutation Based RepresentationsabstractThe U.S. Navy currently deploys Genetic Algorithms to schedule multi-resource satellites. We document this real-world application and also introduce a new synthetic test problem generator. A permutation is used as the representation. A greedy scheduler then converts the permutation into a schedule which can be displayed as a Gantt chart. Surprisingly, there have been few careful comparisons of standard generational Genetic Algorithms and Steady State Genetic Algorithms for these types of problems. In addition, this paper compares different crossover operators for the multi-resource satellite scheduling problem. Finally, we look at two ways of mapping the permutation to a schedule in the form of a Gantt chart. One method gives priority to reducing conflicts, while the other gives priority to reducing overlaps of conflicting tasks. This can produce very different results, even when the evaluation function stays exactly the same. L. Darrell Whitley, Ozeas Quevedo de Carvalho, Mark Roberts, Vivint Shetty, Piyabutra Jampathom |
GECCO | 3 |
| 2023 | Guidelines for a Human-Robot Interaction Specification LanguageabstractDesigning novel application development environments (ADEs) is a growing area of systems research within the human-robot interaction (HRI) community. This research involves the design of a novel system, the ADE, to afford end users and application designers the ability to develop robot applications. Researchers then usually validate their ADEs in the form of user studies or a series of case studies. In this paper, we highlight a problem with the typical approach to conducting ADE research within HRI—there is currently little standardization in how these systems are designed, developed, and validated, leading to difficulty in sharing resources between different research groups and the inability to compare similar ADEs to each other. We argue that a standardized formal representation embedded within an Interaction Specification Language (ISL) can lead to more streamlined development and validation of ADEs for HRI. Furthermore, we discuss several desired characteristics that an ISL should embody. David Porfirio, Mark Roberts, Laura M. Hiatt |
RO-MAN | 2 |
| 2018 | Comparing Reward Shaping, Visual Hints, and Curriculum LearningabstractCommon approaches to learn complex tasks in reinforcement learning include reward shaping, environmental hints, or a curriculum. Yet few studies examine how they compare to each other, when one might prefer one approach, or how they may complement each other. As a first step in this direction, we compare reward shaping, hints, and curricula for a Deep RL agent in the game of Minecraft. We seek to answer whether reward shaping, visual hints, or the curricula have the most impact on performance, which we measure as the time to reach the target, the distance from the target, the cumulative reward, or the number of actions taken. Our analyses show that performance is most impacted by the curriculum used and visual hints; shaping had less impact. For similar navigation tasks, the results suggest that designing an effective curriculum and providing appropriate hints most improve the performance. Common approaches to learn complex tasks in reinforcement learning include reward shaping, environmental hints, or a curriculum, yet few studies examine how they compare to each other. We compare these approaches for a Deep RL agent in the game of Minecraft and show performance is most impacted by the curriculum used and visual hints; shaping had less impact. For similar navigation tasks, this suggests that designing an effective curriculum with hints most improve the performance. Rey Pocius, David Isele, Mark Roberts, David W. Aha |
AAAI | 3 |
| 2016 | Cost-Optimal Algorithms for Planning with Procedural Control KnowledgeabstractThere is an impressive body of work on developing heuristics and other reasoning algorithms to guide search in optimal and anytime planning algorithms for classical planning. However, very little effort has been directed towards developing analogous techniques to guide search towards high-quality solutions in hierarchical planning formalisms like HTN planning, which allows using additional domain-specific procedural control knowledge. In lieu of such techniques, this control knowledge often needs to provide the necessary search guidance to the planning algorithm, which imposes a substantial burden on the domain author and can yield brittle or error-prone domain models. We address this gap by extending recent work on a new hierarchical goal-based planning formalism called Hierarchical Goal Network (HGN) Planning to develop the Hierarchically-Optimal Goal Decomposition Planner (HOpGDP), an HGN planning algorithm that computes hierarchically-optimal plans. HOpGDP is guided by $h_{HL}$, a new HGN planning heuristic that extends existing admissible landmark-based heuristics from classical planning to compute admissible cost estimates for HGN planning problems. Our experimental evaluation across three benchmark planning domains shows that HOpGDP compares favorably to both optimal classical planners due to its ability to use domain-specific procedural knowledge, and a blind-search version of HOpGDP due to the search guidance provided by $h_{HL}$. Vikas Shivashankar, Ron Alford, Mark Roberts, David W. Aha |
ECAI | 3 |
| 2016 | Hierarchical Planning: Relating Task and Goal Decomposition with Task Sharing
Ron Alford, Vikas Shivashankar, Mark Roberts, Jeremy Frank, David W. Aha |
IJCAI | 3 |
| 2015 | Learning to Estimate: A Case-Based Approach to Task Execution Prediction
Bryan Auslander, Michael W. Floyd, Thomas Apker, Benjamin Johnson 0003, Mark Roberts, David W. Aha |
ICCBR | 5 |
| 2013 | Accepting the inevitable: factoring the user into home computer securityabstractHome computer users present unique challenges to computer security. A user's actions frequently affect security without the user understanding how. Moreover, whereas some home users are quite adept at protecting their machines from security threats, a vast majority are not. Current generation security tools, unfortunately, do not tailor security to the home user's needs and actions. In this work, we propose Personalized Attack Graphs (PAG) as a formal technique to model the security risks for the home computer informed by a profile of the user attributes such as preferences, threat perceptions and activities. A PAG also models the interplay between user activities and preferences, attacker strategies, and system activities within the system risk model. We develop a formal model of a user profile to personalize a single, monolithic PAG to different users, and show how to use the user profile to predict user actions. Malgorzata Urbanska, Mark Roberts, Indrajit Ray, Adele E. Howe, Zinta S. Byrne |
CODASPY | 2 |
| 2012 | The Psychology of Security for the Home Computer UserabstractThe home computer user is often said to be the weakest link in computer security. They do not always follow security advice, and they take actions, as in phishing, that compromise themselves. In general, we do not understand why users do not always behave safely, which would seem to be in their best interest. This paper reviews the literature of surveys and studies of factors that influence security decisions for home computer users. We organize the review in four sections: understanding of threats, perceptions of risky behavior, efforts to avoid security breaches and attitudes to security interventions. We find that these studies reveal a lot of reasons why current security measures may not match the needs or abilities of home computer users and suggest future work needed to inform how security is delivered to this user group. Adele E. Howe, Indrajit Ray, Mark Roberts, Malgorzata Urbanska, Zinta S. Byrne |
IEEE Symposium on Security and Privacy | 3 |
| 2009 | Learning from planner performance
Mark Roberts, Adele E. Howe |
Artif. Intell. | 1 |
| 2007 | Harnessing Algorithm Bias in Classical Planning
Mark Roberts |
AAAI | 1 |
| 2006 | Understanding Algorithm Performance on an Oversubscribed Scheduling ApplicationabstractThe best performing algorithms for a particular oversubscribed scheduling application, Air Force Satellite Control Network (AFSCN) scheduling, appear to have little in common. Yet, through careful experimentation and modeling of performance in real problem instances, we can relate characteristics of the best algorithms to characteristics of the application. In particular, we find that plateaus dominate the search spaces (thus favoring algorithms that make larger changes to solutions) and that some randomization in exploration is critical to good performance (due to the lack of gradient information on the plateaus). Based on our explanations of algorithm performance, we develop a new algorithm that combines characteristics of the best performers; the new algorithm's performance is better than the previous best. We show how hypothesis driven experimentation and search modeling can both explain algorithm performance and motivate the design of a new algorithm. Laura Barbulescu, Adele E. Howe, L. Darrell Whitley, Mark Roberts |
J. Artif. Intell. Res. | 4 |
| 2003 | Improving the Scalability of Parallel Jobs by adding Parallel Awareness to the Operating SystemabstractA parallel application benefits from scheduling policies that include a global perspective of the application's process working set. As the interactions among cooperating processes increase, mechanisms to ameliorate waiting within one or more of the processes become more important. In particular, collective operations such as barriers and reductions are extremely sensitive to even usually harmless events such as context switches among members of the process working set. For the last 18 months, we have been researching the impact of random short-lived interruptions such as timer-decrement processing and periodic daemon activity, and developing strategies to minimize their impact on large processor-count SPMD bulk-synchronous programming styles. We present a novel co-scheduling scheme for improving performance of fine-grain collective activities such as barriers and reductions, describe an implementation consisting of operating system kernel modifications and run-time system, and present a set of empirical results comparing the technique with traditional operating system scheduling. Our results indicate a speedup of over 300% on synchronizing collectives. Terry R. Jones, Shawn Dawson, Rob Neely, William G. Tuel Jr., Larry Brenner, Jeffrey Fier, Robert Blackmore, Patrick Caffrey, Brian Maskell, Paul Tomlinson, Mark Roberts |
SC | 11 |
| 1995 | The Aurora RAM CompilerabstractArticle The Aurora RAM compiler Share on Authors: Ajay Chandna University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MI University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MIView Profile , C. David Kibler Hewlett Packard Company, 3404 East Harmony Rd., Ft. Collins, CO Hewlett Packard Company, 3404 East Harmony Rd., Ft. Collins, COView Profile , Richard B. Brown University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MI University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MIView Profile , Mark Roberts University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MI University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MIView Profile , Karem A. Sakallah University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MI University of Michigan, Department of Electrical Engineering & Computer Science, Ann Arbor, MIView Profile Authors Info & Claims DAC '95: Proceedings of the 32nd annual ACM/IEEE Design Automation ConferenceJanuary 1995 Pages 261–266https://doi.org/10.1145/217474.217539Online:01 January 1995Publication History 0citation326DownloadsMetricsTotal Citations0Total Downloads326Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Ajay Chandna, C. David Kibler, Richard B. Brown, Mark Roberts, Karem A. Sakallah |
DAC | 4 |
| 1989 | CEDIF: A Data Driven EDIF ReaderabstractThe Electronic Design Interchange Format (EDIF) allows tools at all levels of the design process to exchange data. This paper describes an efficient method for reading incremental, hierarchical EDIF and the issues involved. The approach taken allows the data to drive the interpretation process by attaching functions to the EDIF keywords. The key to this approach is the representation and interpretation of the EDIF keyword semantics. Mark Roberts |
DAC | 1 |