EDBT 2026 Demo / reviewers in the wild / expert
Josiah Wong
dblp:178/8895
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BEHAVIOR Vision Suite: Customizable Dataset Generation via SimulationabstractThe systematic evaluation and understanding of computer vision models under varying conditions require large amounts of data with comprehensive and customized labels, which real-world vision datasets rarely satisfy. While current synthetic data generators offer a promising alternative, particularly for embodied AI tasks, they often fall short for computer vision tasks due to low asset and rendering quality, limited diversity, and unrealistic physical properties. We introduce the BEHAVIOR Vision Suite (BVS), a set of tools and assets to generate fully customized synthetic data for systematic evaluation of computer vision models, based on the newly developed embodied AI benchmark, BEHAVIOR-1 K. BVS supports a large number of adjustable parameters at the scene level (e.g., lighting, object placement), the object level (e.g., joint configuration, attributes such as “filled” and “folded”), and the camera level (e.g., field of view, focal length). Researchers can arbitrarily vary these parameters during data generation to perform controlled experiments. We showcase three example application scenarios: systematically evaluating the robustness of models across different continuous axes of domain shift, evaluating scene understanding models on the same set of images, and training and evaluating simulation-to-real transfer for a novel vision task: unary and binary state prediction. Project website: https://behavior-vision-suite.github.io/ Yunhao Ge, Yihe Tang, Cem Gökmen, Chengshu Li 0001, Wensi Ai, Benjamin Jose Martinez, Arman Aydin, Mona Anvari, Ayush K. Chakravarthy, Hong-Xing Yu, Josiah Wong, Sanjana Srivastava, Sharon Lee, Shengxin Zha, Laurent Itti, Yunzhu Li, Roberto Martin Martin, Miao Liu 0007, Pengchuan Zhang, Li Fei-Fei 0001, Jiajun Wu 0001 |
CVPR | 12 |
| 2024 | The evolution of the fAIble system to automatically compose and narrate stories for childrenabstractThis article describes our long-term research into automated story generation and our resulting story generation architecture called fAIble that incorporates several innovations. fAIble determines each event that occurs in the tale using a combination of scripted sequences and stochastically chosen events. The probability of an event occurring is based on the skills and personalities of the characters who have agency. Event selection is also influenced by the context of the situation faced by the characters. Each event is associated with a description in grammatically-correct natural language that can be narrated orally via text-to-speech. We describe the evolution of fAIble, its architecture and the results of our independent evaluation of each of the four progressively developed fAIble prototypes (fAIble 0, I, II and III), as tested with human test subjects. On a continuous scale where 0 means unacceptable, 1 means acceptable and 2 means optimal, the composite human test subject rating average from the independent tests of the prototypes was 0.933. The paper also describes a summative assessment where test subjects were asked to review stories from all four prototypes and rank them comparatively. These comparative results indicate an improvement from the original (fAIble 0) to the last one (fAIble III). Avelino J. Gonzalez, Thomas Anchor, Anthony Hevia, Andres Posadas, Josh Wade, Rebeca Amaya Ansag, Kyle A. Benko, Brooke Bottoni, Vera A. Kazakova, Matthew Alvarez, Josiah Wong, Jordan T. Martin, Rainer Knauf, Klaus P. Jantke, Annie S. Wu |
J. Exp. Theor. Artif. Intell. | 11 |
| 2022 | OSCAR: Data-Driven Operational Space Control for Adaptive and Robust Robot ManipulationabstractLearning performant robot manipulation policies can be challenging due to high-dimensional continuous actions and complex physics-based dynamics. This can be alleviated through intelligent choice of action space. Operational Space Control (OSC) has been used as an effective task-space controller for manipulation. Nonetheless, its strength depends on the underlying modeling fidelity, and is prone to failure when there are modeling errors. In this work, we propose OSC for Adaptation and Robustness (OSCAR), a data-driven variant of OSC that compensates for modeling errors by inferring relevant dynamics parameters from online trajectories. OSCAR decomposes dynamics learning into task-agnostic and task-specific phases, decoupling the dynamics dependencies of the robot and the extrinsics due to its environment. This structure enables robust zero-shot performance under out-of-distribution and rapid adaptation to significant domain shifts through additional finetuning. We evaluate our method on a variety of simulated manipulation problems, and find substantial improvements over an array of controller baselines. For more results and information, please visit https://cremebrule.github.io/oscar-web/. Josiah Wong, Viktor Makoviychuk, Anima Anandkumar, Yuke Zhu |
ICRA | 1 |
| 2022 | Detection of driver health condition by monitoring driving behavior through machine learning from observation
Avelino J. Gonzalez, Josiah Wong, Emily M. Thomas, Alec Kerrigan, Lauren Hastings, Andres Posadas, Kevin Negy, Annie S. Wu, Santiago Ontañón, Yi-Ching Lee, Flaura K. Winston |
Expert Syst. Appl. | 2 |
| 2021 | Learning Multi-Arm Manipulation Through Collaborative TeleoperationabstractImitation Learning (IL) is a powerful paradigm to teach robots to perform manipulation tasks by allowing them to learn from human demonstrations collected via teleoperation, but has mostly been limited to single-arm manipulation. However, many real-world tasks require multiple arms, such as lifting a heavy object or assembling a desk. Unfortunately, applying IL to multi-arm manipulation tasks has been challenging –asking a human to control more than one robotic arm can impose significant cognitive burden and is often only possible for a maximum of two robot arms. To address these challenges, we present MULTI-ARM ROBOTURK (MART), a multi-user data collection platform that allows multiple remote users to simultaneously teleoperate a set of robotic arms and collect demonstrations for multi-arm tasks. Using MART, we collected demonstrations for five novel two and three-arm tasks from several geographically separated users. From our data we arrived at a critical insight: most multi-arm tasks do not require global coordination throughout its full duration, but only during specific moments. We show that learning from such data consequently presents challenges for centralized agents that directly attempt to model all robot actions simultaneously, and perform a comprehensive study of different policy architectures with varying levels of centralization on our tasks. Finally, we propose and evaluate a base-residual policy framework that allows trained policies to better adapt to the mixed coordination setting common in multi-arm manipulation, and show that a centralized policy augmented with a decentralized residual model outperforms all other models on our set of benchmark tasks. Additional results and videos at https://roboturk.stanford.edu/multiarm Albert Tung, Josiah Wong, Ajay Mandlekar, Roberto Martin Martin, Yuke Zhu, Li Fei-Fei 0001, Silvio Savarese |
ICRA | 2 |
| 2021 | iGibson 1.0: A Simulation Environment for Interactive Tasks in Large Realistic ScenesabstractWe present iGibson 1.0, a novel simulation environment to develop robotic solutions for interactive tasks in large-scale realistic scenes. Our environment contains 15 fully interactive home-sized scenes with 108 rooms populated with rigid and articulated objects. The scenes are replicas of real-world homes, with distribution and the layout of objects aligned to those of the real world. iGibson 1.0 integrates several key features to facilitate the study of interactive tasks: i) generation of high-quality virtual sensor signals (RGB, depth, segmentation, LiDAR, flow and so on), ii) domain randomization to change the materials of the objects (both visual and physical) and/or their shapes, iii) integrated sampling-based motion planners to generate collision-free trajectories for robot bases and arms, and iv) intuitive human-iGibson interface that enables efficient collection of human demonstrations. Through experiments, we show that the full interactivity of the scenes enables agents to learn useful visual representations that accelerate the training of downstream manipulation tasks. We also show that iGibson features enable the generalization of navigation agents, and that the human-iGibson interface and integrated motion planners facilitate efficient imitation learning of human demonstrated (mobile) manipulation behaviors. iGibson 1.0 is open-source, equipped with comprehensive examples and documentation. For more information, visit our project website: http://svl.stanford.edu/igibson/. Bokui Shen, Fei Xia 0002, Chengshu Li 0002, Roberto Martin Martin, Linxi Fan, Guanzhi Wang, Claudia Pérez-D'Arpino, Shyamal Buch, Sanjana Srivastava, Lyne Tchapmi, Micael Tchapmi, Kent Vainio, Josiah Wong, Li Fei-Fei 0001, Silvio Savarese |
IROS | 13 |
| 2021 | Discovering Tactical Memory From Observed Human Performance in Machine LearningabstractThis article describes an investigation for composing a representation of significant past events that are retained in memory by an observed human actor. These memories influence that actor's performance of a task or making a decision. More specifically, we seek to infer which aspects of the environment and which events have significant effect on an observed actor's future decisions. We introduce a new memory modeling algorithm, memory composition learning, which processes traces of an observed actor's performance, and from these, composes a set of memory features that describe important events in his/her memory that affected these actions. These memory features are subsequently used to produce memory-enhanced traces that can be used by machine learning algorithms to learn memory-influenced behaviors. We implemented a prototype of our approach and evaluated it in two simulated domains, one with synthetic memory-influenced vacuum cleaner agents and one involving human subjects controlling a lawn mower that required memory-influenced behaviors. Results show that our approach is able to discover intuitive representations of tactical memory from observed behavior in both domains, and that these memory representations contributed to improved machine learning of human behavior. Josiah Wong, Avelino J. Gonzalez |
IEEE Trans. Hum. Mach. Syst. | 1 |