Reno Kriz

dblp:220/2001 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-0239-9989ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 100%
Artificial intelligence
8 papers
Machine translation · 29% Language models and text generation · 25% Vision and language · 21%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 100%

Topics — the 26 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models
1.922026
Multi-Vector Index Compression in Any Modality · SIGIR 2026
MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion · SIGIR 2025
Information retrieval › retrieval models › neural retrieval
multi-vector retrieval
1.012026
Multi-Vector Index Compression in Any Modality · SIGIR 2026
Natural language and speech › Machine translation
speech translation
0.912025
Whisper-UT: A Unified Translation Framework for Speech and Text · EMNLP 2025
Computer vision › Vision and language › video-text retrieval
text-to-video retrieval
0.912025
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval · CVPR 2025
Information retrieval
cross-modal retrieval
0.912025
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval · CVPR 2025
Information retrieval › retrieval models › neural retrieval
late interaction
0.912025
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval · CVPR 2025
Information retrieval
multimedia analysis and retrieval
0.912025
MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion · SIGIR 2025
Information retrieval › ranking
rank aggregation
0.912025
MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion · SIGIR 2025
Multimedia analysis and retrieval
cross-modal retrieval
0.912025
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval · CVPR 2025
Multimedia analysis and retrieval
video retrieval
0.912025
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval · CVPR 2025
Computer vision › Vision and language
cross-modal retrieval
0.712023
MultiVENT: Multilingual Videos of Events and Aligned Natural Text · NeurIPS 2023
Information retrieval
cross-language information retrieval
0.712023
MultiVENT: Multilingual Videos of Events and Aligned Natural Text · NeurIPS 2023
Information retrieval › multimedia analysis and retrieval
video retrieval
0.712023
MultiVENT: Multilingual Videos of Events and Aligned Natural Text · NeurIPS 2023
Machine learning › Trustworthy machine learning › calibration
model calibration
0.612022
Ambiguous Images With Human Judgments for Robust Visual Event Classification · NeurIPS 2022
Machine learning › Trustworthy machine learning
uncertainty estimation
0.612022
Ambiguous Images With Human Judgments for Robust Visual Event Classification · NeurIPS 2022
Computer vision › Video understanding and tracking › event recognition
video event recognition
0.612022
Ambiguous Images With Human Judgments for Robust Visual Event Classification · NeurIPS 2022
Natural language and speech › Machine translation
parallel corpora
0.512021
BiSECT: Learning to Split and Rephrase Sentences with Bitexts · EMNLP (1) 2021
Natural language and speech › Machine translation › parallel corpora
sentence alignment
0.512021
BiSECT: Learning to Split and Rephrase Sentences with Bitexts · EMNLP (1) 2021
Natural language and speech › Language models and text generation › text generation › text simplification
sentence simplification
0.512021
BiSECT: Learning to Split and Rephrase Sentences with Bitexts · EMNLP (1) 2021
Natural language and speech › Language models and text generation › text generation › text simplification › sentence simplification
split and rephrase
0.512021
BiSECT: Learning to Split and Rephrase Sentences with Bitexts · EMNLP (1) 2021
Natural language and speech › Language models and text generation › decoding
decoding strategy
0.412019
Comparison of Diverse Decoding Methods from Conditional Language Models · ACL (1) 2019
Natural language and speech › Language models and text generation › decoding
diverse decoding
0.412019
Comparison of Diverse Decoding Methods from Conditional Language Models · ACL (1) 2019
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.312018
Learning Translations via Images with a Massively Multilingual Image Dataset · ACL (1) 2018
Natural language and speech › Machine translation
multimodal machine translation
0.312018
Learning Translations via Images with a Massively Multilingual Image Dataset · ACL (1) 2018
Computer vision › Vision and language
multimodal understanding
0.312025
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval · CVPR 2025
Information retrieval
evaluation
0.312025
MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.7token-wise interaction · 1.7query and visual expansion · 1.7dual sigmoid loss · 1.7video retrieval baselines · 1.3human uncertainty judgments · 1.1index compression · 1.0reciprocal rank fusion · 0.9modality-aware weighting · 0.9multimodal models · 0.7multimodal model · 0.7sequence-to-sequence model · 0.5machine translation · 0.5beam search · 0.4
YearPublicationVenuePosition
2026 Multi-Vector Index Compression in Any Modality
Hanxiang Qin, Alexander Martin 0006, Rohan Jha, Chunsheng Zuo, Reno Kriz, Benjamin Van Durme
SIGIR5
2025 MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
abstract
Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on matching descriptive but vague queries with small collections of professionally edited, English-centric videos. To address this gap, we introduce MultiVENT 2.0, a large-scale, multilingual event-centric video retrieval benchmark featuring a collection of more than 218,000 news videos and over 3,900 queries targeting specific world events. These queries specifically target information found in the visual content, audio, embedded text, and text metadata of the videos, requiring systems leverage all these sources to succeed at the task. Preliminary results show that state-of-the-art vision-language models struggle significantly with this task, and while alternative approaches show promise, they are still insufficient to adequately address this problem. These findings underscore the need for more robust multimodal retrieval systems, as effective video retrieval is a crucial step towards multimodal content understanding and generation.
Reno Kriz, Kate Sanders 0002, David Etter, Kenton Murray, Cameron Carpenter, Hannah Recknor, Jimena Guallar-Blasco, Alexander Martin 0006, Eugene Yang 0001, Benjamin Van Durme
CVPR1
2025 Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
abstract
In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and efficient mechanism for fine-grained similarity assessment between queries and videos. Video-ColBERT is built upon three main components: a fine-grained spatial and temporal token-wise interaction, query and visual expansions, and a dual sigmoid loss during training. We find that this interaction and training paradigm leads to strong individual, yet compatible, representations for encoding video content. These representations lead to increases in performance on common text-to-video retrieval benchmarks compared to other bi-encoder methods.
Arun V. Reddy, Alexander Martin 0006, Eugene Yang 0001, Andrew Yates, Kate Sanders 0002, Kenton Murray, Reno Kriz, Celso de Melo, Benjamin Van Durme, Rama Chellappa
CVPR7
2025 Whisper-UT: A Unified Translation Framework for Speech and Text
abstract
Cihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz, Keith Cunningham, Kenton Murray, Kevin Duh, Luis Tavarez-Arce, Paul McNamee, Sanjeev Khudanpur. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Cihan Xiao, Matthew Wiesner, Debashish Chakraborty, Reno Kriz, Keith Cunningham, Kenton Murray, Kevin Duh, Luis Tavarez-Arce, Paul McNamee, Sanjeev Khudanpur
EMNLP4
2025 MMMORRF: Multimodal Multilingual MOdularized Reciprocal Rank Fusion
abstract
Videos inherently contain multiple modalities, including visual events, text overlays, sounds, and speech, all of which are important for retrieval. However, state-of-the-art multimodal language models like VAST and LanguageBind are built on vision-language models (VLMs), and thus overly prioritize visual signals. Retrieval benchmarks further reinforce this bias by focusing on visual queries and neglecting other modalities. We create a search system MMMORRF that extracts text and features from both visual and audio modalities and integrates them with a novel modality-aware weighted reciprocal rank fusion. MMMORRF is both effective and efficient, demonstrating practicality in searching videos based on users' information needs instead of visual descriptive queries. We evaluate MMMORRF on MultiVENT 2.0 and TVR, two multimodal benchmarks designed for more targeted information needs, and find that it improves nDCG@20 by 81% over leading multimodal encoders and 37% over single-modality retrieval.
Saron Samuel, Dan DeGenaro, Jimena Guallar-Blasco, Kate Sanders 0002, Seun Eisape, Arun V. Reddy, Alexander Martin 0006, Andrew Yates, Eugene Yang 0001, Cameron Carpenter, David Etter, Efsun Selin Kayi, Matthew Wiesner, Kenton Murray, Reno Kriz
SIGIR15
2023 MultiVENT: Multilingual Videos of Events and Aligned Natural Text
abstract
Everyday news coverage has shifted from traditional broadcasts towards a wide range of presentation formats such as first-hand, unedited video footage. Datasets that reflect the diverse array of multimodal, multilingual news sources available online could be used to teach models to benefit from this shift, but existing news video datasets focus on traditional news broadcasts produced for English-speaking audiences. We address this limitation by constructing MultiVENT, a dataset of multilingual, event-centric videos grounded in text documents across five target languages. MultiVENT includes both news broadcast videos and non-professional event footage, which we use to analyze the state of online news videos and how they can be leveraged to build robust, factually accurate models. Finally, we provide a model for complex, multilingual video retrieval to serve as a baseline for information retrieval using MultiVENT.
Kate Sanders 0002, David Etter, Reno Kriz, Benjamin Van Durme
NeurIPS3
2022 Did that happen? Predicting Social Media Posts that are Indicative of what happened in a scene: A case study of a TV show
abstract
While popular Television (TV) shows are airing, some users interested in these shows publish social media posts about the show. Analyzing social media posts related to a TV show can be beneficial for gaining insights about what happened during scenes of the show. This is a challenging task partly because a significant number of social media posts associated with a TV show or event may not clearly describe what happened during the event. In this work, we propose a method to predict social media posts (associated with scenes of a TV show) that are indicative of what transpired during the scenes of the show. We evaluate our method on social media (Twitter) posts associated with an episode of a popular TV show, Game of Thrones. We show that for each of the identified scenes, with high AUC’s, our method can predict posts that are indicative of what happened in a scene from those that are not-indicative. Based on Twitters policy, we will make the Tweeter ID’s of the Twitter posts used for this work publicly available.
Anietie Andy, Reno Kriz, Sharath Chandra Guntuku, Derry Wijaya, Chris Callison-Burch
LREC2
2022 Ambiguous Images With Human Judgments for Robust Visual Event Classification
abstract
Contemporary vision benchmarks predominantly consider tasks on which humans can achieve near-perfect performance. However, humans are frequently presented with visual data that they cannot classify with 100% certainty, and models trained on standard vision benchmarks achieve low performance when evaluated on this data. To address this issue, we introduce a procedure for creating datasets of ambiguous images and use it to produce SQUID-E ("Squidy"), a collection of noisy images extracted from videos. All images are annotated with ground truth values and a test set is annotated with human uncertainty judgments. We use this dataset to characterize human uncertainty in vision tasks and evaluate existing visual event classification models. Experimental results suggest that existing vision models are not sufficiently equipped to provide meaningful outputs for ambiguous images and that datasets of this nature can be used to assess and improve such models through model training and direct evaluation of model calibration. These findings motivate large-scale ambiguous dataset creation and further research focusing on noisy visual data.
Kate Sanders 0002, Reno Kriz, Anqi Liu 0001, Benjamin Van Durme
NeurIPS2
2021 BiSECT: Learning to Split and Rephrase Sentences with Bitexts
abstract
An important task in NLP applications such as sentence simplification is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary.We introduce a novel dataset and a new model for this 'split and rephrase' task.Our BISECT training data consists of 1 million long English sentences paired with shorter, meaning-equivalent English sentences.We obtain these by extracting 1-2 sentence alignments in bilingual parallel corpora and then using machine translation to convert both sides of the corpus into the same language.BISECT contains higher quality training examples than previous Split and Rephrase corpora, with sentence splits that require more significant modifications.We categorize examples in our corpus, and use these categories in a novel model that allows us to target specific regions of the input sentence to be split and edited.Moreover, we show that models trained on BISECT can perform a wider variety of split operations and improve upon previous state-of-the-art approaches in automatic and human evaluations.1
Joongwon Kim, Mounica Maddela, Reno Kriz, Wei Xu 0004, Chris Callison-Burch
EMNLP (1)3
2019 Comparison of Diverse Decoding Methods from Conditional Language Models
abstract
While conditional language models have greatly improved in their ability to output high-quality natural language, many NLP applications benefit from being able to generate a diverse set of candidate sequences.Diverse decoding strategies aim to, within a givensized candidate list, cover as much of the space of high-quality outputs as possible, leading to improvements for tasks that re-rank and combine candidate outputs.Standard decoding methods, such as beam search, optimize for generating high likelihood sequences rather than diverse ones, though recent work has focused on increasing diversity in these methods.In this work, we perform an extensive survey of decoding-time strategies for generating diverse outputs from conditional language models.We also show how diversity can be improved without sacrificing quality by oversampling additional candidates, then filtering to the desired number.
Daphne Ippolito, Reno Kriz, João Sedoc, Maria Kustikova, Chris Callison-Burch
ACL (1)2
2018 Learning Translations via Images with a Massively Multilingual Image Dataset
abstract
John Hewitt, Daphne Ippolito, Brendan Callahan, Reno Kriz, Derry Tanti Wijaya, Chris Callison-Burch. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
John Hewitt, Daphne Ippolito, Brendan Callahan, Reno Kriz, Derry Wijaya, Chris Callison-Burch
ACL (1)4
2018 Simplification Using Paraphrases and Context-Based Lexical Substitution
abstract
Reno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Reno Kriz, Eleni Miltsakaki, Marianna Apidianaki, Chris Callison-Burch
NAACL-HLT1