VLDB 2026 Research / reviewers in the wild / expert
Ji Ma 0004
dblp:253/2346-4
· DBLP profile ↗
13ranked-venue papers
2as first author
10since 2021 · last 2025
0009-0009-2102-8209ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Massive Sound Embedding Benchmark (MSEB)abstractAudio is a critical component of multimodal perception, and any truly intelligent system must demonstrate a wide range of auditory capabilities. These capabilities include transcription, classification, retrieval, reasoning, segmentation, clustering, reranking, and reconstruction. Fundamentally, each task involves transforming a raw audio signal into a meaningful 'embedding'—be it a single vector, a sequence of continuous or discrete representations, or another structured form—which then serves as the basis for generating the task's final response. To accelerate progress towards robust machine auditory intelligence, we present the Massive Sound Embedding Benchmark (MSEB): an extensible framework designed to evaluate the auditory components of any multimodal system. In its first release, MSEB offers a comprehensive suite of eight core tasks, with more planned for the future, supported by diverse datasets, including the new, large-scale Simple Voice Questions (SVQ) dataset. Our initial experiments establish clear performance headrooms, highlighting the significant opportunity to improve real-world multimodal experiences where audio is a core signal. We encourage the research community to use MSEB to assess their algorithms and contribute to its growth. The library is publicly hosted at https://github.com/google-research/mseb. Georg Heigold, Ehsan Variani, Tom Bagby, Cyril Allauzen, Ji Ma 0004, Shankar Kumar, Michael Riley 0001 |
NeurIPS | 5 |
| 2024 | OpenMSD: Towards Multilingual Scientific Documents Similarity MeasurementabstractWe develop and evaluate multilingual scientific documents similarity measurement models in this work. Such models can be used to find related papers in different languages, which can help multilingual researchers find and explore papers more efficiently. We propose the first multilingual scientific documents dataset, Open-access Multilingual Scientific Documents (OpenMSD), which has 74M papers in 103 languages and 778M citation pairs. With OpenMSD, we develop multilingual SDSM models by adjusting and extending the state-of-the-art methods designed for English SDSM tasks. We find that: (i)Some highly successful methods in English SDSM yield significantly worse performance in multilingual SDSM. (ii)Our best model, which enriches the non-English papers with English summaries, outperforms strong baselines by 7% (in mean average precision) on multilingual SDSM tasks, without compromising the performance on English SDSM tasks. Ji Ma 0004, Ivan Korotkov, Keith B. Hall, Dana Alon, Donald Metzler |
LREC/COLING | 2 |
| 2024 | HYRR: Hybrid Infused Reranking for Passage RetrievalabstractExisting passage retrieval systems typically adopt a two-stage retrieve-then-rerank pipeline. To obtain an effective reranking model, many prior works have focused on improving the model architectures, such as leveraging powerful pretrained large language models (LLM) and designing better objective functions. However, less attention has been paid to the issue of collecting high-quality training data. In this paper, we propose HYRR, a framework for training robust reranking models. Specifically, we propose a simple but effective approach to select training data using hybrid retrievers. Our experiments show that the rerankers trained with HYRR are robust to different first-stage retrievers. Moreover, evaluations using MS MARCO and BEIR data sets demonstrate our proposed framework effectively generalizes to both supervised and zero-shot retrieval settings. Jing Lu 0014, Keith B. Hall, Ji Ma 0004, Jianmo Ni |
LREC/COLING | 3 |
| 2023 | Promptagator: Few-shot Dense Retrieval From 8 Examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma 0004, Yi Luan, Jianmo Ni, Jing Lu 0014, Anton Bakalov, Kelvin Guu, Keith B. Hall, Ming-Wei Chang |
ICLR | 3 |
| 2023 | Learning List-Level Domain-Invariant Representations for RankingabstractDomain adaptation aims to transfer the knowledge learned on (data-rich) source domains to (low-resource) target domains, and a popular method is invariant representation learning, which matches and aligns the data distributions on the feature space. Although this method is studied extensively and applied on classification and regression problems, its adoption on ranking problems is sporadic, and the few existing implementations lack theoretical justifications. This paper revisits invariant representation learning for ranking. Upon reviewing prior work, we found that they implement what we call item-level alignment, which aligns the distributions of the items being ranked from all lists in aggregate but ignores their list structure. However, the list structure should be leveraged, because it is intrinsic to ranking problems where the data and the metrics are defined and computed on lists, not the items by themselves. To close this discrepancy, we propose list-level alignment—learning domain-invariant representations at the higher level of lists. The benefits are twofold: it leads to the first domain adaptation generalization bound for ranking, in turn providing theoretical support for the proposed method, and it achieves better empirical transfer performance for unsupervised domain adaptation on ranking tasks, including passage reranking. Ruicheng Xian, Honglei Zhuang, Zhen Qin 0001, Hamed Zamani, Jing Lu 0014, Ji Ma 0004, Kai Hui 0001, Han Zhao 0002, Xuanhui Wang, Michael Bendersky |
NeurIPS | 6 |
| 2023 | RankT5: Fine-Tuning T5 for Text Ranking with Ranking LossesabstractPretrained language models such as BERT have been shown to be exceptionally effective for text ranking. However, there are limited studies on how to leverage more powerful sequence-to-sequence models such as T5. Existing attempts usually formulate text ranking as a classification problem and rely on postprocessing to obtain a ranked list. In this paper, we propose RankT5 and study two T5-based ranking model structures, an encoder-decoder and an encoder-only one, so that they not only can directly output ranking scores for each query-document pair, but also can be fine-tuned with pairwise or listwise ranking losses to optimize ranking performance. Our experiments show that the proposed models with ranking losses can achieve substantial ranking performance gains on different public text ranking data sets. Moreover, ranking models fine-tuned with listwise ranking losses have better zero-shot ranking performance on out-of-domain data than models fine-tuned with classification losses. Honglei Zhuang, Zhen Qin 0001, Rolf Jagerman, Kai Hui 0001, Ji Ma 0004, Jing Lu 0014, Jianmo Ni, Xuanhui Wang, Michael Bendersky |
SIGIR | 5 |
| 2023 | QAmeleon: Multilingual QA with Only 5 ExamplesabstractAbstract The availability of large, high-quality datasets has been a major driver of recent progress in question answering (QA). Such annotated datasets, however, are difficult and costly to collect, and rarely exist in languages other than English, rendering QA technology inaccessible to underrepresented languages. An alternative to building large monolingual training datasets is to leverage pre-trained language models (PLMs) under a few-shot learning setting. Our approach, QAmeleon, uses a PLM to automatically generate multilingual data upon which QA models are fine-tuned, thus avoiding costly annotation. Prompt tuning the PLM with only five examples per language delivers accuracy superior to translation-based baselines; it bridges nearly 60% of the gap between an English-only baseline and a fully-supervised upper bound fine-tuned on almost 50,000 hand-labeled examples; and consistently leads to improvements compared to directly fine-tuning a QA model on labeled examples in low resource settings. Experiments on the TyDiqa-GoldP and MLQA benchmarks show that few-shot prompt tuning for data synthesis scales across languages and is a viable alternative to large-scale annotation.1 Priyanka Agrawal, Christopher Alberti, Fantine Huot, Joshua Maynez, Ji Ma 0004, Sebastian Ruder, Kuzman Ganchev, Dipanjan Das 0001, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | Large Dual Encoders Are Generalizable RetrieversabstractJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, Yinfei Yang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jianmo Ni, Chen Qu 0001, Jing Lu 0014, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma 0004, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, Yinfei Yang |
EMNLP | 6 |
| 2021 | Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question GenerationabstractA major obstacle to the wide-spread adoption of neural retrieval models is that they require large supervised training sets to surpass traditional term-based techniques, which are constructed from raw corpora.In this paper, we propose an approach to zero-shot learning for passage retrieval that uses synthetic question generation to close this gap.The question generation system is trained on general domain data, but is applied to documents in the targeted domain.This allows us to create arbitrarily large, yet noisy, question-passage relevance pairs that are domain specific.Furthermore, when this is coupled with a simple hybrid termneural model, first-stage retrieval performance can be improved further.Empirically, we show that this is an effective strategy for building neural passage retrieval models in the absence of large training corpora.Depending on the domain, this technique can even approach the accuracy of supervised models. Ji Ma 0004, Ivan Korotkov, Yinfei Yang, Keith B. Hall, Ryan T. McDonald |
EACL | 1 |
| 2021 | Multi-stage Training with Improved Negative Contrast for Neural Passage RetrievalabstractIn the context of neural passage retrieval, we study three promising techniques: synthetic data generation, negative sampling, and fusion. We systematically investigate how these techniques contribute to the performance of the retrieval system and how they complement each other. We propose a multi-stage framework comprising of pre-training with synthetic data, fine-tuning with labeled data, and negative sampling at both stages. We study six negative sampling strategies and apply them to the fine-tuning stage and, as a noteworthy novelty, to the synthetic data that we use for pre-training. Also, we explore fusion methods that combine negatives from different strategies. We evaluate our system using two passage retrieval tasks for open-domain QA and using MS MARCO. Our experiments show that augmenting the negative contrast in both stages is effective to improve passage retrieval accuracy and, importantly, they also show that synthetic data generation and negative sampling have additive benefits. Moreover, using the fusion of different kinds allows us to reach performance that establishes a new state-of-the-art level in two of the tasks we evaluated. Jing Lu 0014, Gustavo Hernández Ábrego, Ji Ma 0004, Jianmo Ni, Yinfei Yang |
EMNLP (1) | 3 |
| 2018 | State-of-the-art Chinese Word Segmentation with Bi-LSTMsabstractA wide variety of neural-network architectures have been proposed for the task of Chinese word segmentation.Surprisingly, we find that a bidirectional LSTM model, when combined with standard deep learning techniques and best practices, can achieve better accuracy on many of the popular datasets as compared to models based on more complex neuralnetwork architectures.Furthermore, our error analysis shows that out-of-vocabulary words remain challenging for neural-network models, and many of the remaining errors are unlikely to be fixed through architecture changes.Instead, more effort should be made on exploring resources for further improvement. Ji Ma 0004, Kuzman Ganchev, David Weiss 0001 |
EMNLP | 1 |
| 2017 | Natural Language Processing with Small Feed-Forward NetworksabstractWe show that small and shallow feedforward neural networks can achieve near state-of-the-art results on a range of unstructured and structured language processing tasks while being considerably cheaper in memory and computational requirements than deep recurrent models.Motivated by resource-constrained environments like mobile phones, we showcase simple techniques for obtaining such small neural network models, and investigate different tradeoffs when deciding how to allocate a small memory budget. Jan A. Botha, Emily Pitler, Ji Ma 0004, Anton Bakalov, Alex Salcianu, David Weiss 0001, Ryan T. McDonald, Slav Petrov |
EMNLP | 3 |
| 2016 | Generalized Transition-based Dependency Parsing via Control ParametersabstractIn this paper, we present a generalized transition-based parsing framework where parsers are instantiated in terms of a set of control parameters that constrain transitions between parser states.This generalization provides a unified framework to describe and compare various transitionbased parsing approaches from both a theoretical and empirical perspective.This includes well-known transition systems, but also previously unstudied systems. Bernd Bohnet, Ryan T. McDonald, Emily Pitler, Ji Ma 0004 |
ACL (1) | 4 |