Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Philip Schroeder

dblp:330/4377 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Video understanding and tracking · 44% Reinforcement learning · 44% Robot manipulation · 13%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › multi-agent reinforcement learning
recursive reasoning
0.912025
ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks · NeurIPS 2025
Computer vision › Video understanding and tracking › deep video understanding
video reasoning
0.912025
ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

vision-language model · 0.9in-context learning · 0.9
YearPublicationVenuePosition
2025 THREAD: Thinking Deeper with Recursive Spawning
abstract
Philip Schroeder, Nathaniel W. Morgan, Hongyin Luo, James R. Glass. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Philip Schroeder, Nathaniel Morgan, Hongyin Luo, James R. Glass
NAACL (Long Papers)1
2025 ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
abstract
Vision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their utility in embodied settings, which require reasoning over long frame sequences from a continuous stream of visual input at each moment of a task attempt. To address this limitation, we propose ROVER (Reasoning Over VidEo Recursively), a framework that enables the model to recursively decompose long-horizon video trajectories into segments corresponding to shorter subtasks within the trajectory. In doing so, ROVER facilitates more focused and accurate reasoning over temporally localized frame sequences without losing global context. We evaluate ROVER, implemented using an in-context learning approach, on diverse OpenX Embodiment videos and on a new dataset derived from RoboCasa that consists of 543 videos showing both expert and perturbed non-expert trajectories across 27 manipulation tasks. ROVER outperforms strong baselines across three video reasoning tasks: task progress estimation, frame-level natural language reasoning, and video question answering. We observe that, by reducing the number of frames the model reasons over at each timestep, ROVER mitigates model hallucinations, especially during unexpected or non-optimal moments of a trajectory. In addition, by enabling the implementation of a subtask-specific sliding context window, ROVER's time complexity scales linearly with video length, an asymptotic improvement over baselines.
Philip Schroeder, Ondrej Biza, Thomas Weng, Hongyin Luo, James R. Glass
NeurIPS1