VLDB 2026 Research / reviewers in the wild / expert
Reno Kriz
dblp:220/2001
· DBLP profile ↗
12ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-0239-9989ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 100% | |
| Artificial intelligence
8 papers |
Machine translation · 29% Language models and text generation · 25% Vision and language · 21% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 100% |
Topics — the 26 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval models |
1.9 | 2 | 2026 | Multi-Vector Index Compression in Any Modality · SIGIR 2026 MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion · SIGIR 2025 |
Information retrieval › retrieval models › neural retrieval
multi-vector retrieval |
1.0 | 1 | 2026 | Multi-Vector Index Compression in Any Modality · SIGIR 2026 |
Natural language and speech › Machine translation
speech translation |
0.9 | 1 | 2025 | Whisper-UT: A Unified Translation Framework for Speech and Text · EMNLP 2025 |
Computer vision › Vision and language › video-text retrieval
text-to-video retrieval |
0.9 | 1 | 2025 | Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval · CVPR 2025 |
Information retrieval
cross-modal retrieval |
0.9 | 1 | 2025 | Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval · CVPR 2025 |
Information retrieval › retrieval models › neural retrieval
late interaction |
0.9 | 1 | 2025 | Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval · CVPR 2025 |
Information retrieval
multimedia analysis and retrieval |
0.9 | 1 | 2025 | MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion · SIGIR 2025 |
Information retrieval › ranking
rank aggregation |
0.9 | 1 | 2025 | MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion · SIGIR 2025 |
Multimedia analysis and retrieval
cross-modal retrieval |
0.9 | 1 | 2025 | MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval · CVPR 2025 |
Multimedia analysis and retrieval
video retrieval |
0.9 | 1 | 2025 | MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval · CVPR 2025 |
Computer vision › Vision and language
cross-modal retrieval |
0.7 | 1 | 2023 | MultiVENT: Multilingual Videos of Events and Aligned Natural Text · NeurIPS 2023 |
Information retrieval
cross-language information retrieval |
0.7 | 1 | 2023 | MultiVENT: Multilingual Videos of Events and Aligned Natural Text · NeurIPS 2023 |
Information retrieval › multimedia analysis and retrieval
video retrieval |
0.7 | 1 | 2023 | MultiVENT: Multilingual Videos of Events and Aligned Natural Text · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › calibration
model calibration |
0.6 | 1 | 2022 | Ambiguous Images With Human Judgments for Robust Visual Event Classification · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.6 | 1 | 2022 | Ambiguous Images With Human Judgments for Robust Visual Event Classification · NeurIPS 2022 |
Computer vision › Video understanding and tracking › event recognition
video event recognition |
0.6 | 1 | 2022 | Ambiguous Images With Human Judgments for Robust Visual Event Classification · NeurIPS 2022 |
Natural language and speech › Machine translation
parallel corpora |
0.5 | 1 | 2021 | BiSECT: Learning to Split and Rephrase Sentences with Bitexts · EMNLP (1) 2021 |
Natural language and speech › Machine translation › parallel corpora
sentence alignment |
0.5 | 1 | 2021 | BiSECT: Learning to Split and Rephrase Sentences with Bitexts · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › text generation › text simplification
sentence simplification |
0.5 | 1 | 2021 | BiSECT: Learning to Split and Rephrase Sentences with Bitexts · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › text generation › text simplification › sentence simplification
split and rephrase |
0.5 | 1 | 2021 | BiSECT: Learning to Split and Rephrase Sentences with Bitexts · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › decoding
decoding strategy |
0.4 | 1 | 2019 | Comparison of Diverse Decoding Methods from Conditional Language Models · ACL (1) 2019 |
Natural language and speech › Language models and text generation › decoding
diverse decoding |
0.4 | 1 | 2019 | Comparison of Diverse Decoding Methods from Conditional Language Models · ACL (1) 2019 |
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
0.3 | 1 | 2018 | Learning Translations via Images with a Massively Multilingual Image Dataset · ACL (1) 2018 |
Natural language and speech › Machine translation
multimodal machine translation |
0.3 | 1 | 2018 | Learning Translations via Images with a Massively Multilingual Image Dataset · ACL (1) 2018 |
Computer vision › Vision and language
multimodal understanding |
0.3 | 1 | 2025 | MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval · CVPR 2025 |
Information retrieval
evaluation |
0.3 | 1 | 2025 | MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 1.7token-wise interaction · 1.7query and visual expansion · 1.7dual sigmoid loss · 1.7video retrieval baselines · 1.3human uncertainty judgments · 1.1index compression · 1.0reciprocal rank fusion · 0.9modality-aware weighting · 0.9multimodal models · 0.7multimodal model · 0.7sequence-to-sequence model · 0.5machine translation · 0.5beam search · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Vector Index Compression in Any Modality
Hanxiang Qin, Alexander Martin 0006, Rohan Jha, Chunsheng Zuo, Reno Kriz, Benjamin Van Durme |
SIGIR | 5 |
| 2025 | MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video RetrievalabstractEfficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on matching descriptive but vague queries with small collections of professionally edited, English-centric videos. To address this gap, we introduce MultiVENT 2.0, a large-scale, multilingual event-centric video retrieval benchmark featuring a collection of more than 218,000 news videos and over 3,900 queries targeting specific world events. These queries specifically target information found in the visual content, audio, embedded text, and text metadata of the videos, requiring systems leverage all these sources to succeed at the task. Preliminary results show that state-of-the-art vision-language models struggle significantly with this task, and while alternative approaches show promise, they are still insufficient to adequately address this problem. These findings underscore the need for more robust multimodal retrieval systems, as effective video retrieval is a crucial step towards multimodal content understanding and generation. Reno Kriz, Kate Sanders 0002, David Etter, Kenton Murray, Cameron Carpenter, Hannah Recknor, Jimena Guallar-Blasco, Alexander Martin 0006, Eugene Yang 0001, Benjamin Van Durme |
CVPR | 1 |
| 2025 | Video-ColBERT: Contextualized Late Interaction for Text-to-Video RetrievalabstractIn this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and efficient mechanism for fine-grained similarity assessment between queries and videos. Video-ColBERT is built upon three main components: a fine-grained spatial and temporal token-wise interaction, query and visual expansions, and a dual sigmoid loss during training. We find that this interaction and training paradigm leads to strong individual, yet compatible, representations for encoding video content. These representations lead to increases in performance on common text-to-video retrieval benchmarks compared to other bi-encoder methods. Arun V. Reddy, Alexander Martin 0006, Eugene Yang 0001, Andrew Yates, Kate Sanders 0002, Kenton Murray, Reno Kriz, Celso de Melo, Benjamin Van Durme, Rama Chellappa |
CVPR | 7 |
| 2025 | Whisper-UT: A Unified Translation Framework for Speech and TextabstractCihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz, Keith Cunningham, Kenton Murray, Kevin Duh, Luis Tavarez-Arce, Paul McNamee, Sanjeev Khudanpur. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Cihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz, Keith Cunningham, Kenton Murray, Kevin Duh, Luis Tavarez-Arce, Paul McNamee, Sanjeev Khudanpur |
EMNLP | 4 |
| 2025 | MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank FusionabstractVideos inherently contain multiple modalities, including visual events, text overlays, sounds, and speech, all of which are important for retrieval. However, state-of-the-art multimodal language models like VAST and LanguageBind are built on vision-language models (VLMs), and thus overly prioritize visual signals. Retrieval benchmarks further reinforce this bias by focusing on visual queries and neglecting other modalities. We create a search system MMMORRF that extracts text and features from both visual and audio modalities and integrates them with a novel modality-aware weighted reciprocal rank fusion. MMMORRF is both effective and efficient, demonstrating practicality in searching videos based on users' information needs instead of visual descriptive queries. We evaluate MMMORRF on MultiVENT 2.0 and TVR, two multimodal benchmarks designed for more targeted information needs, and find that it improves nDCG@20 by 81% over leading multimodal encoders and 37% over single-modality retrieval. Saron Samuel, Dan DeGenaro, Jimena Guallar-Blasco, Kate Sanders 0002, Seun Eisape, Arun V. Reddy, Alexander Martin 0006, Andrew Yates, Eugene Yang 0001, Cameron Carpenter, David Etter, Efsun Selin Kayi, Matthew Wiesner, Kenton Murray, Reno Kriz |
SIGIR | 15 |
| 2023 | MultiVENT: Multilingual Videos of Events and Aligned Natural TextabstractEveryday news coverage has shifted from traditional broadcasts towards a wide range of presentation formats such as first-hand, unedited video footage. Datasets that reflect the diverse array of multimodal, multilingual news sources available online could be used to teach models to benefit from this shift, but existing news video datasets focus on traditional news broadcasts produced for English-speaking audiences. We address this limitation by constructing MultiVENT, a dataset of multilingual, event-centric videos grounded in text documents across five target languages. MultiVENT includes both news broadcast videos and non-professional event footage, which we use to analyze the state of online news videos and how they can be leveraged to build robust, factually accurate models. Finally, we provide a model for complex, multilingual video retrieval to serve as a baseline for information retrieval using MultiVENT. Kate Sanders 0002, David Etter, Reno Kriz, Benjamin Van Durme |
NeurIPS | 3 |
| 2022 | Did that happen? Predicting Social Media Posts that are Indicative of what happened in a scene: A case study of a TV showabstractWhile popular Television (TV) shows are airing, some users interested in these shows publish social media posts about the show. Analyzing social media posts related to a TV show can be beneficial for gaining insights about what happened during scenes of the show. This is a challenging task partly because a significant number of social media posts associated with a TV show or event may not clearly describe what happened during the event. In this work, we propose a method to predict social media posts (associated with scenes of a TV show) that are indicative of what transpired during the scenes of the show. We evaluate our method on social media (Twitter) posts associated with an episode of a popular TV show, Game of Thrones. We show that for each of the identified scenes, with high AUC’s, our method can predict posts that are indicative of what happened in a scene from those that are not-indicative. Based on Twitters policy, we will make the Tweeter ID’s of the Twitter posts used for this work publicly available. Anietie Andy, Reno Kriz, Sharath Chandra Guntuku, Derry Wijaya, Chris Callison-Burch |
LREC | 2 |
| 2022 | Ambiguous Images With Human Judgments for Robust Visual Event ClassificationabstractContemporary vision benchmarks predominantly consider tasks on which humans can achieve near-perfect performance. However, humans are frequently presented with visual data that they cannot classify with 100% certainty, and models trained on standard vision benchmarks achieve low performance when evaluated on this data. To address this issue, we introduce a procedure for creating datasets of ambiguous images and use it to produce SQUID-E ("Squidy"), a collection of noisy images extracted from videos. All images are annotated with ground truth values and a test set is annotated with human uncertainty judgments. We use this dataset to characterize human uncertainty in vision tasks and evaluate existing visual event classification models. Experimental results suggest that existing vision models are not sufficiently equipped to provide meaningful outputs for ambiguous images and that datasets of this nature can be used to assess and improve such models through model training and direct evaluation of model calibration. These findings motivate large-scale ambiguous dataset creation and further research focusing on noisy visual data. Kate Sanders 0002, Reno Kriz, Anqi Liu 0001, Benjamin Van Durme |
NeurIPS | 2 |
| 2021 | BiSECT: Learning to Split and Rephrase Sentences with BitextsabstractAn important task in NLP applications such as sentence simplification is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary.We introduce a novel dataset and a new model for this 'split and rephrase' task.Our BISECT training data consists of 1 million long English sentences paired with shorter, meaning-equivalent English sentences.We obtain these by extracting 1-2 sentence alignments in bilingual parallel corpora and then using machine translation to convert both sides of the corpus into the same language.BISECT contains higher quality training examples than previous Split and Rephrase corpora, with sentence splits that require more significant modifications.We categorize examples in our corpus, and use these categories in a novel model that allows us to target specific regions of the input sentence to be split and edited.Moreover, we show that models trained on BISECT can perform a wider variety of split operations and improve upon previous state-of-the-art approaches in automatic and human evaluations.1 Joongwon Kim, Mounica Maddela, Reno Kriz, Wei Xu 0004, Chris Callison-Burch |
EMNLP (1) | 3 |
| 2019 | Comparison of Diverse Decoding Methods from Conditional Language ModelsabstractWhile conditional language models have greatly improved in their ability to output high-quality natural language, many NLP applications benefit from being able to generate a diverse set of candidate sequences.Diverse decoding strategies aim to, within a givensized candidate list, cover as much of the space of high-quality outputs as possible, leading to improvements for tasks that re-rank and combine candidate outputs.Standard decoding methods, such as beam search, optimize for generating high likelihood sequences rather than diverse ones, though recent work has focused on increasing diversity in these methods.In this work, we perform an extensive survey of decoding-time strategies for generating diverse outputs from conditional language models.We also show how diversity can be improved without sacrificing quality by oversampling additional candidates, then filtering to the desired number. Daphne Ippolito, Reno Kriz, João Sedoc, Maria Kustikova, Chris Callison-Burch |
ACL (1) | 2 |
| 2018 | Learning Translations via Images with a Massively Multilingual Image DatasetabstractJohn Hewitt, Daphne Ippolito, Brendan Callahan, Reno Kriz, Derry Tanti Wijaya, Chris Callison-Burch. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. John Hewitt, Daphne Ippolito, Brendan Callahan, Reno Kriz, Derry Wijaya, Chris Callison-Burch |
ACL (1) | 4 |
| 2018 | Simplification Using Paraphrases and Context-Based Lexical SubstitutionabstractReno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Reno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch |
NAACL-HLT | 1 |