Luis Pesqueira

dblp:243/3021 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0008-9625-9186ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Video understanding and tracking · 65% 3D vision · 35%
Human-computer interaction and pervasive computing
2 papers
Wearable and physiological sensing · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Wearable and physiological sensing › wearable camera › egocentric vision
egocentric sensing
1.622025
Reading Recognition in the Wild · NeurIPS 2025
Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild · ECCV (24) 2024
Computer vision › Video understanding and tracking
activity recognition
0.912025
Reading Recognition in the Wild · NeurIPS 2025
Computer vision › 3D vision
egocentric vision
0.912025
Reading Recognition in the Wild · NeurIPS 2025
Computer vision › Video understanding and tracking › motion analysis
human motion analysis
0.812024
Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild · ECCV (24) 2024
Wearable and physiological sensing › wearable display
smart glasses
0.312025
Reading Recognition in the Wild · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

transformer · 1.7head pose · 1.7eye gaze · 1.7
YearPublicationVenuePosition
2025 Reading Recognition in the Wild
abstract
To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism. Code, model, and data will be public.
Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx, Lambert Mathias, Kiran K. Somasundaram, Luis Pesqueira, James Fort, Sheroze Sheriffdeen, Omkar M. Parkhi, Carl Yuheng Ren, Mi Zhang 0002, Yuning Chai, Richard A. Newcombe, Hyo Jin Kim 0004
NeurIPS7
2024 Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild
Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang 0002, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim 0004, Kevin Bailey, David Soriano Fosas, C. Karen Liu, Ziwei Liu 0002, Jakob J. Engel, Renzo De Nardi, Richard A. Newcombe
ECCV (24)7