Ben Newman

dblp:325/6182 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Question answering and dialogue systems · 50% 3D vision · 50%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › multimodal question answering
embodied question answering
0.812024
OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024
Computer vision › 3D vision
environmental understanding
0.812024
OpenEQA: Embodied Question Answering in the Era of Foundation Models · CVPR 2024

Methods — techniques the papers use, named apart from their topics

large language model evaluation · 0.8foundation model · 0.8
YearPublicationVenuePosition
2024 OpenEQA: Embodied Question Answering in the Era of Foundation Models
abstract
We present a modern formulation of Embodied Question Answering (EQA) as the task of understanding an environment well enough to answer questions about it in natural language. An agent can achieve such an understanding by either drawing upon episodic memory, exemplified by agents on smart glasses, or by actively exploring the environment, as in the case of mobile robots. We accompany our formulation with OpenEQA - the first open-vocabulary benchmark dataset for EQA supporting both episodic memory and active exploration use cases. OpenEQA contains over 1600 high-quality human generated questions drawn from over 180 real-world environments. In addition to the dataset, we also provide an automatic LLM-powered evaluation protocol that has excellent correlation with human judgement. Using this dataset and evaluation protocol, we evaluate several state-of-the-art foundation models including GPT-4V, and find that they significantly lag behind human-level performance. Consequently, OpenEQA stands out as a straightforward, measurable, and practically rele-vant benchmark that poses a considerable challenge to current generation offoundation models. We hope this inspires and stimulates future research at the intersection of Embod-ied AI, conversational agents, and world models.
Arjun Majumdar, Anurag Ajay, Xiaohan Zhang 0002, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Silwal, Paul McVay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma 0001, Vincent-Pierre Berges, Shiqi Zhang 0001, Pulkit Agrawal 0001, Yonatan Bisk, Dhruv Batra, Mrinal Kalakrishnan, Franziska Meier, Chris Paxton 0001, Alexander Sax, Aravind Rajeswaran
CVPR13
2022 Work-in-Progress - Using Virtual Reality in Museums to Assist Historical Learning
abstract
Museums are always looking for new ways to attract and engage their visitors. Virtual reality (VR) is becoming an increasingly viable option for doing this and enables the recreation of historical sites/structures for visitors to explore. This paper describes a VR application being developed in partnership with the Bawdsey Radar Museum to recreate the radar transmitter towers that stood on the site and were pivotal in winning the Battle of Britain. This paper describes the application’s development and discusses user trial results to determine whether VR can be effectively used within a museum setting to assist in historical learning.
Adam Morsman, Ben Newman, Stephanie Bally, Jack Johns, Hannah Cook, Daniel McHugh, Grace Williams
iLRN2