VLDB 2026 Research / reviewers in the wild / expert
Victor Gomes
dblp:284/4484
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Generative modeling · 44% Video understanding and tracking · 44% Language models and text generation · 13% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 50% Immersive interaction · 50% | |
| Computer graphics and multimedia
1 paper |
Virtual and augmented reality · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding |
0.8 | 1 | 2024 | Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual Modalities · CVPR 2024 |
Machine learning › Generative modeling › autoregressive model
multimodal autoregressive model |
0.8 | 1 | 2024 | Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual Modalities · CVPR 2024 |
Immersive interaction
head-mounted display |
0.8 | 1 | 2024 | Googly Eyes: Exploring Effects of Displaying User's Eyes Outward on a Virtual Reality Head-Mounted Display on User Experience · VR 2024 |
Human-robot interaction
nonverbal communication |
0.8 | 1 | 2024 | Googly Eyes: Exploring Effects of Displaying User's Eyes Outward on a Virtual Reality Head-Mounted Display on User Experience · VR 2024 |
Natural language and speech › Language models and text generation › language modeling
multimodal language modeling |
0.2 | 1 | 2024 | Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual Modalities · CVPR 2024 |
Virtual and augmented reality › immersive interaction
co-located collaboration |
0.2 | 1 | 2024 | Googly Eyes: Exploring Effects of Displaying User's Eyes Outward on a Virtual Reality Head-Mounted Display on User Experience · VR 2024 |
Methods — techniques the papers use, named apart from their topics
prototype · 1.5between-subjects user study · 1.5sequence partitioning · 0.8autoregressive modeling · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The development of polyseme learning under uncertainty
Alexander S. LaTourrette, Victor Gomes, Kantinka Tangen, John C. Trueswell |
CogSci | 2 |
| 2024 | Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual ModalitiesabstractOne of the main challenges of multimodal learning is combining multiple heterogeneous modalities, e.g., video, audio, and text. Video and audio are obtained at much higher rates than text and are roughly aligned in time. They are often not synchronized with text, which comes as a global context, e.g. a title, or a description. Furthermore, video and audio inputs are of much larger volumes, and grow as the video length increases, which naturally requires more compute dedicated to these modalities, and makes modeling of long-range dependencies harder. We here decouple the multimodal modeling, dividing it into separate autoregressive models, processing the inputs according to the characteristics of the modalities. We propose a multimodal model, consisting of an autoregressive component for the time-synchronized modalities (audio and video), and an autoregressive component for the context modalities which are not necessarily aligned in time but are still sequential. To address the long-sequences of the video-audio inputs, we further partition the video and audio sequences in consecutive snippets and autoregressively process their representations. To that end, we propose a Combiner mechanism, which models the audio-video information jointly, producing compact but expressive representations. This allows us to scale to 512 input video frames without increase in model parameters. Our approach achieves the state-of-the-art on multiple well established multimodal benchmarks. It effectively addresses the high computational demand of media inputs by learning compact representations, controlling the sequence length of the audio-video feature representations, and modeling their dependencies in time. A. J. Piergiovanni, Isaac Noble, Dahun Kim, Michael S. Ryoo, Victor Gomes, Anelia Angelova |
CVPR | 5 |
| 2024 | Googly Eyes: Exploring Effects of Displaying User's Eyes Outward on a Virtual Reality Head-Mounted Display on User ExperienceabstractHead-mounted displays (HMDs) in virtual reality (VR) occlude the upper face of the wearing users, leading to decreased nonverbal communication cues toward outside users. In this paper, we discuss Googly Eyes, a high-fidelity prototype that displays an illustration of the HMD-wearing user’s eyes in real-time in front of an HMD. We explored the effects of the Googly Eyes on the user experience of non-HMD users. For this, we designed and developed a collaborative asymmetrical co-located task performed by an HMD-wearing user and a non-HMD (tablet) user, and we conducted a between-subjects user study where we compared the Googly Eyes with a baseline HMD experience without any external eye depiction. In this paper, we discuss the system, task, user study details, and results along with implications for future studies. Evren Bozgeyikli, Lila Bozgeyikli, Victor Gomes |
VR | 3 |
| 2020 | Not what you expect: The relationship between violation of expectation and negation
Victor Gomes, Yubin Huh, John C. Trueswell |
CogSci | 1 |