VLDB 2026 Research / reviewers in the wild / expert
Reza Pourreza 0002
dblp:02/7252-2
· DBLP profile ↗
12ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-9667-1209ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?abstractMulti-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering instructions but also detecting their successful execution, as well as identifying and alerting users to mistakes, all of which has to happen in real-time. This requires models that are not turn-based, but that can react asynchronously to a video stream, as well as video data showing users performing tasks including mistakes and their corrections. To this end, we introduce Qualcomm Interactive Cooking, a new benchmark and dataset built upon CaptainCook4D, which contains user mistakes during task execution. Our dataset and benchmark features densely annotated, timed instructions and feedback messages, specifically including mistake alerts precisely timestamped to their visual occurrence in the video. We evaluate state-of-the-art multi-modal LLMs on the Qualcomm Interactive Cooking benchmark and introduce LiveMamba, a streaming multi-modal LLM designed for interactive instructional guidance. This work provides the first dedicated benchmark and a strong baseline for developing and evaluating on live, situated coaching. Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh, Reza Pourreza 0002, Litian Liu, Sunny Panchal, Leonid Sigal, Roland Memisevic |
NeurIPS | 4 |
| 2024 | Look, Remember and Reason: Grounded Reasoning in Videos with Language ModelsabstractMulti-modal language models (LM) have recently shown promising performance
in high-level reasoning tasks on videos. However, existing methods still fall short in tasks like causal or compositional spatiotemporal reasoning over actions, in which model predictions need to be grounded in fine-grained low-level details, such as object motions and object interactions.
In this work, we propose training an LM end-to-end on low-level surrogate tasks, including object detection, re-identification, and tracking, to endow the model with the required low-level visual capabilities. We show that a two-stream video encoder with spatiotemporal attention is effective at capturing the required static and motion-based cues in the video. By leveraging the LM's ability to perform the low-level surrogate tasks, we can cast reasoning in videos as the three-step process of *Look, Remember, Reason*, wherein visual information is extracted using low-level visual skills step-by-step and then integrated to arrive at a final answer. We demonstrate the effectiveness of our framework on diverse visual reasoning tasks from the ACRE, CATER, Something-Else and STAR datasets. Our approach is trainable end-to-end and surpasses state-of-the-art task-specific methods across these tasks by a large margin. Apratim Bhattacharyya, Sunny Panchal, Reza Pourreza 0002, Mingu Lee, Pulkit Madan, Roland Memisevic |
ICLR | 3 |
| 2024 | What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated InteractionabstractVision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, where each turn must be stepped (i.e., prompted) by the user. Open-ended, asynchronous interactions, where an AI model may proactively deliver timely responses or feedback based on the unfolding situation in real-time, are an open challenge. In this work, we present the QEVD benchmark and dataset, which explores human-AI interaction in the challenging, yet controlled, real-world domain of fitness coaching – a task which intrinsically requires monitoring live user activity and providing immediate feedback. The benchmark requires vision-language models to recognize complex human actions, identify possible mistakes, and provide appropriate feedback in real-time. Our experiments reveal the limitations of existing state-of-the-art vision-language models for such asynchronous situated interactions. Motivated by this, we propose a simple end-to-end streaming baseline that can respond asynchronously to human actions with appropriate feedback at the appropriate time. Sunny Panchal, Apratim Bhattacharyya, Guillaume Berger, Antoine Mercier 0005, Cornelius Böhm, Florian Dietrichkeit, Reza Pourreza 0002, Pulkit Madan, Mingu Lee, Mark Todorovich, Ingo Bax, Roland Memisevic |
NeurIPS | 7 |
| 2023 | A neural video codec with spatial rate-distortion controlabstractNeural video compression algorithms are nearly competitive with hand-crafted codecs in terms of rate-distortion performance and subjective quality. However, many neural codecs are inflexible black boxes, and give users little to no control over the reconstruction quality and bitrate. In this work, we present a flexible neural video codec that combines ideas from variable-bitrate codecs and region-of-interest-based coding. By conditioning our model on a global rate-distortion tradeoff parameter and a region-of-interest (ROI) mask, we obtain dynamic control over the per-frame bitrate and the reconstruction quality in the ROI at test time. The resulting codec enables practical use cases such as coding under bitrate constraints with fixed ROI quality, while taking a negligible hit in performance compared to a fixed-rate model. We find that our codec performs best on sequences with complex motion, where we substantially outperform non-ROI codecs in the region of interest with Bjøntegaard-Delta rate savings exceeding 60%. Noor Fathima Ghouse, Jens Petersen, Guillaume Sautière, Auke J. Wiggers, Reza Pourreza 0002 |
WACV | 5 |
| 2023 | Boosting neural video codecs by exploiting hierarchical redundancyabstractIn video compression, coding efficiency is improved by reusing pixels from previously decoded frames via motion and residual compensation. We define two levels of hierarchical redundancy in video frames: 1) first-order: redundancy in pixel space, i.e., similarities in pixel values across neighboring frames, which is effectively captured using motion and residual compensation, 2) second-order: redundancy in motion and residual maps due to smooth motion in natural videos. While most of the existing neural video coding literature addresses first-order redundancy, we tackle the problem of capturing second-order redundancy in neural video codecs via predictors. We introduce generic motion and residual predictors that learn to extrapolate from previously decoded data. These predictors are lightweight, and can be employed with most neural video codecs in order to improve their rate-distortion performance. Moreover, while RGB is the dominant colorspace in neural video coding literature, we introduce general modifications for neural video codecs to embrace the YUV420 colorspace and report YUV420 results. Our experiments show that using our predictors with a well-known neural video codec leads to 38% and 34% bitrate savings in RGB and YUV420 colorspaces measured on the UVG dataset. Reza Pourreza 0002, Hoang Le, Amir Said, Guillaume Sautière, Auke J. Wiggers |
WACV | 1 |
| 2022 | GameCodec: Neural Cloud Gaming Video Codec
Hoang Le, Reza Pourreza 0002, Amir Said, Guillaume Sautière, Auke J. Wiggers |
BMVC | 2 |
| 2022 | Optimized Learned Entropy Coding Parameters for Practical Neural-Based Image and Video CompressionabstractNeural-based image and video codecs are significantly more power-efficient when weights and activations are quantized to low-precision integers. While there are general-purpose techniques for reducing quantization effects, large losses can occur when specific entropy coding properties are not considered. This work analyzes how entropy coding is affected by parameter quantizations, and provides a method to minimize losses. It is shown that, by using a certain type of coding parameters to be learned, uniform quantization becomes practically optimal, also simplifying the minimization of code memory requirements. The mathematical properties of the new representation are presented, and its effectiveness is demonstrated by coding experiments, showing that good results can be obtained with precision as low as 4 bits per network output, and practically no loss with 8 bits. Amir Said, Reza Pourreza 0002, Hoang Le |
ICIP | 2 |
| 2022 | MobileCodec: neural inter-frame video compression on mobile devicesabstractRealizing the potential of neural codecs on real-world mobile devices is a big technological challenge due to the inherent conflict between the computational complexity of deep networks and the power-constrained mobile hardware performance. We demonstrate practical feasibility by leveraging Qualcomm's innovation and technology, bridging the gap from neural network-based model simulations to operation on a mobile device powered by Snapdragon® technology. We show the first-ever inter-frame neural video decoder running on a commercial mobile phone, decompressing high-definition videos in real-time while maintaining a low bitrate and high visual quality, comparable to conventional codecs. Hoang Le, Amir Said, Guillaume Sautière, Yang Yang 0010, Pranav Shrestha, Reza Pourreza 0002, Auke J. Wiggers |
MMSys | 8 |
| 2022 | Differentiable bit-rate estimation for neural-based video codec enhancementabstractNeural networks (NN) can improve standard video compression by pre-and post-processing the encoded video. For optimal NN training, the standard codec needs to be replaced with a codec proxy that can provide derivatives of estimated bit-rate and distortion, which are used for gradient back-propagation. Since entropy coding of standard codecs is designed to take into account non-linear dependencies between transform coefficients, bit-rates cannot be well approximated with simple per-coefficient estimators. This paper presents a new approach for bit-rate estimation that is similar to the type employed in training end-to-end neural codecs, and able to efficiently take into account those statistical dependencies. It is defined from a mathematical model that provides closed-form formulas for the estimates and their gradients, reducing the computational complexity. Experimental results demonstrate the method’s accuracy in estimating HEVC/H.265 codec bit-rates. Amir Said, Manish Kumar Singh 0002, Reza Pourreza 0002 |
PCS | 3 |
| 2021 | Extending Neural P-frame Codecs for B-frame CodingabstractWhile most neural video codecs address P-frame coding (predicting each frame from past ones), in this paper we address B-frame compression (predicting frames using both past and future reference frames). Our B-frame solution is based on the existing P-frame methods. As a result, B-frame coding capability can easily be added to an existing neural codec. The basic idea of our B-frame coding method is to interpolate the two reference frames to generate a single reference frame and then use it together with an existing P-frame codec to encode the input B-frame. Our studies show that the interpolated frame is a much better reference for the P-frame codec compared to using the previous frame as is usually done. Our results show that using the proposed method with an existing P-frame codec can lead to 28.5% saving in bit-rate on the UVG dataset compared to the P-frame codec while generating the same video quality. Reza Pourreza 0002, Taco Cohen |
ICCV | 1 |
| 2020 | Feedback Recurrent Autoencoder for Video Compression
Adam Golinski, Reza Pourreza 0002, Yang Yang 0010, Guillaume Sautière, Taco Cohen |
ACCV (4) | 2 |
| 2020 | Parallelized Rate-Distortion Optimized Quantization Using Deep LearningabstractRate-Distortion Optimized Quantization (RDOQ) has played an important role in the coding performance of recent video compression standards such as H.264/AVC, H.265/HEVC, VP9 and AV1. This scheme yields significant reductions in bit-rate at the expense of relatively small increases in distortion. Typically, RDOQ algorithms are prohibitively expensive to implement on real-time hardware encoders due to their sequential nature and their need to frequently obtain entropy coding costs. This work addresses this limitation using a neural network-based approach, which learns to trade-off rate and distortion during offline supervised training. As these networks are based solely on standard arithmetic operations that can be executed on existing neural network hardware, no additional area-on-chip needs to be reserved for dedicated RDOQ circuitry. We train two classes of neural networks, a fully-convolutional network and an auto-regressive network, and evaluate each as a post-quantization step designed to refine cheap quantization schemes such as scalar quantization (SQ). Both network architectures are designed to have a low computational overhead. After training they are integrated into the HM 16.20 implementation of HEVC, and their video coding performance is evaluated on a subset of the H.266/VVC SDR common test sequences. Comparisons are made to RDOQ and SQ implementations in HM16.20. Our method achieves 1.64% BD-rate savings on luminosity compared to the HM SQ anchor, and on average reaches 45% of the performance of the iterative HM RDOQ algorithm. Dana Kianfar, Auke J. Wiggers, Amir Said, Reza Pourreza 0002, Taco Cohen |
MMSP | 4 |