VLDB 2026 Research / reviewers in the wild / expert
Edison Marrese-Taylor
dblp:137/6899
· DBLP profile ↗
17ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MED-COREASONER: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-ReasoningabstractWhile reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local languages, limiting equitable global medical deployment. To bridge this gap, we introduce Med-CoReasoner, a language-informed co-reasoning framework that elicits parallel English and local-language reasoning, abstracts them into structured concepts, and integrates local clinical knowledge into an English logical scaffold via concept-level alignment and retrieval. This design combines the structural robustness of English reasoning with the practice-grounded expertise encoded in local languages. To evaluate multilingual medical reasoning beyond multiple-choice settings, we construct MultiMed-X, a benchmark covering seven languages with expert-annotated long-form question answering and natural language inference tasks, comprising 350 instances per language. Experiments across three benchmarks show that Med-CoReasoner improves multilingual reasoning performance by an average of 5%, with particularly substantial gains in low-resource languages. Moreover, model distillation and expert evaluation analysis further confirm that Med-CoReasoner produces clinically sound and culturally grounded reasoning traces. Sherry T. Tong, Jiwoong Sohn, Ding Xia, Piyalitt Ittichaiwong, Kanyakorn Veerakanjana, Hyunjae Kim, Qingyu Chen 0001, Edison Marrese-Taylor, Kazuma Kobayashi, Akiko Aizawa, Irene Li |
ACL (1) | 11 |
| 2025 | MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model EvaluationabstractWeihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li |
EMNLP | 28 |
| 2025 | Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone EnsemblingabstractContrastive Language-Image Pretraining (CLIP) stands out as a prominent method for image representation learning. Various architectures, from vision transformers~(ViTs) to convolutional networks (ResNets) have been trained with CLIP to serve as general solutions to diverse vision tasks.
This paper explores the differences across various CLIP-trained vision backbones.
Despite using the same data and training objective, we find that these architectures have notably different representations,
different classification performance across datasets, and different robustness properties to certain types of image perturbations.
Our findings indicate a remarkable possible synergy across backbones
by leveraging their respective strengths.
In principle, classification accuracy could be improved by over 40 percentage with an informed selection of the optimal backbone per test example.
Using this insight, we develop a straightforward yet powerful approach to adaptively ensemble multiple backbones.
The approach uses as few as one labeled example per class
to tune the adaptive combination of backbones.
On a large collection of datasets, the method achieves a remarkable increase in accuracy of up to 39.1\% over the best single backbone, well beyond traditional ensembles. Cristian Rodriguez Opazo, Ehsan Abbasnejad, Damien Teney, Hamed Damirchi, Edison Marrese-Taylor, Anton van den Hengel |
ICLR | 5 |
| 2025 | Language Models can Categorize System Inputs for Performance AnalysisabstractDominic Sobhani, Ruiqi Zhong, Edison Marrese-Taylor, Keisuke Sakaguchi, Yutaka Matsuo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dominic Sobhani, Ruiqi Zhong, Edison Marrese-Taylor, Keisuke Sakaguchi, Yutaka Matsuo |
NAACL (Long Papers) | 3 |
| 2024 | Annotations for Exploring Food Tweets from Multiple AspectsabstractThis research builds upon the Latvian Twitter Eater Corpus (LTEC), which is focused on the narrow domain of tweets related to food, drinks, eating and drinking. LTEC has been collected for more than 12 years and reaching almost 3 million tweets with the basic information as well as extended automatically and manually annotated metadata. In this paper we supplement the LTEC with manually annotated subsets of evaluation data for machine translation, named entity recognition, timeline-balanced sentiment analysis, and text-image relation classification. We experiment with each of the data sets using baseline models and highlight future challenges for various modelling approaches. Matiss Rikters, Rinalds Viksna, Edison Marrese-Taylor |
LREC/COLING | 3 |
| 2024 | Media Bias Detection Across Families of Language ModelsabstractIffat Maab, Edison Marrese-Taylor, Sebastian Padó, Yutaka Matsuo. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Iffat Maab, Edison Marrese-Taylor, Sebastian Padó, Yutaka Matsuo |
NAACL-HLT | 2 |
| 2023 | Memory-efficient Temporal Moment Localization in Long VideosabstractCristian Rodriguez-Opazo, Edison Marrese-Taylor, Basura Fernando, Hiroya Takamura, Qi Wu. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Cristian Rodriguez Opazo, Edison Marrese-Taylor, Basura Fernando, Hiroya Takamura, Qi Wu 0001 |
EACL | 2 |
| 2023 | Target-Aware Contextual Political Bias Detection in NewsabstractIffat Maab, Edison Marrese-Taylor, Yutaka Matsuo. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Iffat Maab, Edison Marrese-Taylor, Yutaka Matsuo |
IJCNLP (1) | 2 |
| 2022 | Open-domain Video Commentary GenerationabstractEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topić, Yusuke Miyao, Ichiro Kobayashi, Hiroya Takamura. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Edison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic, Yusuke Miyao, Ichiro Kobayashi 0001, Hiroya Takamura |
EMNLP | 1 |
| 2021 | Variational Inference for Learning Representations of Natural Language EditsabstractDocument editing has become a pervasive component of production of information, with version control systems enabling edits to be efficiently stored and applied. In light of this, the task of learning distributed representations of edits has been recently proposed. With this in mind, we propose a novel approach that employs variational inference to learn a continuous latent space of vector representations to capture the underlying semantic information with regard to the document editing process. We achieve this by introducing a latent variable to explicitly model the aforementioned features. This latent variable is then combined with a document representation to guide the generation of an edited-version of this document. Additionally, to facilitate standardized automatic evaluation of edit representations, which has heavily relied on direct human input thus far, we also propose a suite of downstream tasks, PEER, specifically designed to measure the quality of edit representations in the context of natural language processing. Edison Marrese-Taylor, Machel Reid, Yutaka Matsuo |
AAAI | 1 |
| 2021 | DORi: Discovering Object Relationships for Moment Localization of a Natural Language Query in a VideoabstractThis paper studies the task of temporal moment localization in long untrimmed videos using natural language queries. Given a query sentence, the goal is to determine the start and end of the relevant segment within the video. Our key innovation is to learn a video feature embedding through a language-conditioned message-passing algorithm suitable for temporal moment localization which captures the relationships between humans, objects and activities in the video. These relationships are obtained by a spatial sub-graph that contextualizes the scene representation using detected objects and human features conditioned in the language query. Moreover, a temporal sub-graph captures the activities within the video through time. Our method is evaluated on three standard benchmark datasets, and we also introduce YouCookII as a new benchmark for this task. Experiments show our method outperforms state-of-the-art methods on these datasets, confirming the effectiveness of our approach. Cristian Rodriguez Opazo, Edison Marrese-Taylor, Basura Fernando, Hongdong Li, Stephen Gould |
WACV | 2 |
| 2020 | VCDM: Leveraging Variational Bi-encoding and Deep Contextualized Word Representations for Improved Definition ModelingabstractIn this paper, we tackle the task of definition modeling, where the goal is to learn to generate definitions of words and phrases. Existing approaches for this task are discriminative, combining distributional and lexical semantics in an implicit rather than direct way. To tackle this issue we propose a generative model for the task, introducing a continuous latent variable to explicitly model the underlying relationship between a phrase used within a context and its definition. We rely on variational inference for estimation and leverage contextualized word embeddings for improved performance. Our approach is evaluated on four existing challenging benchmarks with the addition of two new datasets, “Cambridge” and the first non-English corpus “Robert”, which we release to complement our empirical study. Our Variational Contextual Definition Modeler (VCDM) achieves state-of-the-art performance in terms of automatic and human evaluation metrics, demonstrating the effectiveness of our approach. Machel Reid, Edison Marrese-Taylor, Yutaka Matsuo |
EMNLP (1) | 2 |
| 2020 | Learning to Describe Editing Activities in Collaborative Environments: A Case Study on GitHub and Wikipedia
Edison Marrese-Taylor, Pablo Loyola, Jorge A. Balazs, Yutaka Matsuo |
PACLIC | 1 |
| 2020 | Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided AttentionabstractThis paper studies the problem of temporal moment localization in a long untrimmed video using natural language as the query. Given an untrimmed video and a query sentence, the goal is to determine the start and end of the relevant visual moment in the video that corresponds to the query sentence. While most previous works have tackled this by a propose-and-rank approach, we introduce a more efficient, end-to-end trainable, and proposal-free approach that is built upon three key components: a dynamic filter which adaptively transfers language information to visual domain attention map, a new loss function to guide the model to attend the most relevant part of the video, and soft labels to cope with annotation uncertainties. Our method is evaluated on three standard benchmark datasets, Charades-STA, TACoS and ActivityNet-Captions. Experimental results show our method outperforms state-of-the-art methods on these datasets, confirming the effectiveness of the method. We believe the proposed dynamic filter-based guided attention mechanism will prove valuable for other vision and language tasks as well. Cristian Rodriguez Opazo, Edison Marrese-Taylor, Fatemehsadat Saleh, Hongdong Li, Stephen Gould |
WACV | 2 |
| 2018 | Content Aware Source Code Change Description GenerationabstractWe propose to study the generation of descriptions from source code changes by integrating the messages included on code commits and the intra-code documentation inside the source in the form of docstrings.Our hypothesis is that although both types of descriptions are not directly aligned in semantic terms -one explaining a change and the other the actual functionality of the code being modified-there could be certain common ground that is useful for the generation.To this end, we propose an architecture that uses the source codedocstring relationship to guide the description generation.We discuss the results of the approach comparing against a baseline based on a sequence-to-sequence model, using standard automatic natural language generation metrics as well as with a human study, thus offering a comprehensive view of the feasibility of the approach. Pablo Loyola, Edison Marrese-Taylor, Jorge A. Balazs, Yutaka Matsuo, Fumiko Satoh |
INLG | 2 |
| 2014 | A novel deterministic approach for aspect-based opinion mining in tourism products reviews
Edison Marrese-Taylor, Juan D. Velásquez 0001, Felipe Bravo-Marquez |
Expert Syst. Appl. | 1 |
| 2013 | Identifying Customer Preferences about Tourism Products Using an Aspect-based Opinion Mining ApproachabstractIn this study we extend Bing Liu's aspect-based opinion mining technique to apply it to the tourism domain. Using this extension, we also offer an approach for considering a new alternative to discover consumer preferences about tourism products, particularly hotels and restaurants, using opinions available on the Web as reviews. An experiment is also conducted, using hotel and restaurant reviews obtained from TripAdvisor, to evaluate our proposals. Results showed that tourism product reviews available on web sites contain valuable information about customer preferences that can be extracted using an aspect-based opinion mining approach. The proposed approach proved to be very effective in determining the sentiment orientation of opinions, achieving a precision and recall of 90%. However, on average, the algorithms were only capable of extracting 35% of the explicit aspect expressions. Edison Marrese-Taylor, Juan D. Velásquez 0001, Felipe Bravo-Marquez, Yutaka Matsuo |
KES | 1 |