VLDB 2026 Research / reviewers in the wild / expert
Jianmo Ni
dblp:161/2449
· DBLP profile ↗
23ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0002-6863-8073ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ActionPiece: Contextually Tokenizing Action Sequences for Generative RecommendationabstractGenerative recommendation (GR) is an emerging paradigm where user actions are tokenized into discrete token patterns and autoregressively generated as predictions. However, existing GR models tokenize each action independently, assigning the same fixed tokens to identical actions across all sequences without considering contextual relationships. This lack of context-awareness can lead to suboptimal performance, as the same action may hold different meanings depending on its surrounding context. To address this issue, we propose ActionPiece to explicitly incorporate context when tokenizing action sequences. In ActionPiece, each action is represented as a set of item features. Given the action sequence corpora, we construct the vocabulary by merging feature patterns as new tokens, based on their co-occurrence frequency both within individual sets and across adjacent sets. Considering the unordered nature of feature sets, we further introduce set permutation regularization, which produces multiple segmentations of action sequences with the same semantics. Our code is available at: https://github.com/google-deepmind/action_piece. Yupeng Hou, Jianmo Ni, Zhankui He, Noveen Sachdeva, Wang-Cheng Kang, Ed H. Chi, Julian J. McAuley, Zhiyuan Cheng 0002 |
ICML | 2 |
| 2024 | HYRR: Hybrid Infused Reranking for Passage RetrievalabstractExisting passage retrieval systems typically adopt a two-stage retrieve-then-rerank pipeline. To obtain an effective reranking model, many prior works have focused on improving the model architectures, such as leveraging powerful pretrained large language models (LLM) and designing better objective functions. However, less attention has been paid to the issue of collecting high-quality training data. In this paper, we propose HYRR, a framework for training robust reranking models. Specifically, we propose a simple but effective approach to select training data using hybrid retrievers. Our experiments show that the rerankers trained with HYRR are robust to different first-stage retrievers. Moreover, evaluations using MS MARCO and BEIR data sets demonstrate our proposed framework effectively generalizes to both supervised and zero-shot retrieval settings. Jing Lu 0014, Keith B. Hall, Ji Ma 0004, Jianmo Ni |
LREC/COLING | 4 |
| 2024 | Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense RetrievalabstractNandan Thakur, Jianmo Ni, Gustavo Hernandez Abrego, John Wieting, Jimmy Lin, Daniel Cer. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nandan Thakur, Jianmo Ni, Gustavo Hernández Ábrego, John Wieting, Jimmy Lin, Daniel M. Cer |
NAACL-HLT | 2 |
| 2024 | Improving Data Efficiency for Recommenders and LLMsabstractIn recent years, massive transformer-based architectures have driven breakthrough performance in practical applications like autoregressive text-generation (LLMs) and click-prediction (recommenders). A common recipe for success is to train large models on massive web-scale datasets [3, 15], e.g., modern recommenders are trained on billions of user-item click events, and LLMs are trained on trillions of tokens extracted from the public internet. We are close to hitting the computational and economical limits of scaling up the size of these models, and we expect the next frontier of gains to come from improving the: (i) data quality of the training dataset, and (ii) data efficiency of the extremely expensive training procedure. Inspired by this shift, we present a set of “data-centric” techniques for recommendation and language models that summarizes a dataset into a terse data summary, which is both (i) high-quality, i.e., trains better quality models, and (ii) improves the data-efficiency of the overall training procedure. We propose techniques from two disparate data frameworks: (i) data selection (a.k.a., coreset construction) methods that sample portions of the dataset using grounded heuristics, and (ii) data distillation techniques that generate synthetic examples which are optimized to retain the signals needed for training high-quality models. Overall, this work sheds light on the challenges and opportunities offered by data optimization in web-scale systems, a particularly relevant focus as the recommendation community grapples with the grand challenge of leveraging LLMs. Noveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang, Jianmo Ni, James Caverlee, Lichan Hong, Ed H. Chi, Zhiyuan Cheng 0002 |
RecSys | 4 |
| 2023 | A Suite of Generative Tasks for Multi-Level Multimodal Webpage UnderstandingabstractAndrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown, Bryan Plummer, Kate Saenko, Jianmo Ni, Mandy Guo. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Andrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown, Bryan A. Plummer, Kate Saenko, Jianmo Ni, Mandy Guo |
EMNLP | 7 |
| 2023 | Promptagator: Few-shot Dense Retrieval From 8 Examples
Zhuyun Dai, Vincent Y. Zhao, Ji Ma 0004, Yi Luan, Jianmo Ni, Jing Lu 0014, Anton Bakalov, Kelvin Guu, Keith B. Hall, Ming-Wei Chang |
ICLR | 5 |
| 2023 | Efficient Data Representation Learning in Google-scale Systemsabstract"Garbage in, Garbage out" is a familiar maxim to ML practitioners and researchers, because the quality of a learned data representation is highly crucial to the quality of any ML model that consumes it as an input. To handle systems that serve billions of users at millions of queries per second (QPS), we need representation learning algorithms with significantly improved efficiency. At Google, we have dedicated thousands of iterations to develop a set of powerful techniques that efficiently learn high quality data representations. We have thoroughly validated these methods through offline evaluation, online A/B testing, and deployed these in over 50 models across major Google products. In this paper, we consider a generalized data representation learning problem that allows us to identify feature embeddings and crosses as common challenges. We propose two solutions, including: 1. Multi-size Unified Embedding to learn high-quality embeddings; and 2. Deep Cross Network V2 for learning effective feature crosses. We discuss the practical challenges we encountered and solutions we developed during deployment to production systems, compare with SOTA methods, and report offline and online experimental results. This work sheds light on the challenges and opportunities for developing next-gen algorithms for web-scale systems. Zhiyuan Cheng 0002, Wang-Cheng Kang, Benjamin Coleman, Yin Zhang 0011, Jianmo Ni, Jonathan Valverde, Lichan Hong, Ed H. Chi |
RecSys | 6 |
| 2023 | RankT5: Fine-Tuning T5 for Text Ranking with Ranking LossesabstractPretrained language models such as BERT have been shown to be exceptionally effective for text ranking. However, there are limited studies on how to leverage more powerful sequence-to-sequence models such as T5. Existing attempts usually formulate text ranking as a classification problem and rely on postprocessing to obtain a ranked list. In this paper, we propose RankT5 and study two T5-based ranking model structures, an encoder-decoder and an encoder-only one, so that they not only can directly output ranking scores for each query-document pair, but also can be fine-tuned with pairwise or listwise ranking losses to optimize ranking performance. Our experiments show that the proposed models with ranking losses can achieve substantial ranking performance gains on different public text ranking data sets. Moreover, ranking models fine-tuned with listwise ranking losses have better zero-shot ranking performance on out-of-domain data than models fine-tuned with classification losses. Honglei Zhuang, Zhen Qin 0001, Rolf Jagerman, Kai Hui 0001, Ji Ma 0004, Jing Lu 0014, Jianmo Ni, Xuanhui Wang, Michael Bendersky |
SIGIR | 7 |
| 2023 | Scaling Up Models and Data with t5x and seqioabstractScaling up training datasets and model parameters have benefited neural network-based language models, but also present challenges like distributed compute, input data bottlenecks and reproducibility of results. We introduce two simple and scalable software libraries that simplify these issues: t5x enables training large language models at scale, while seqio enables reproducible input and evaluation pipelines. These open-source libraries have been used to train models with hundreds of billions of parameters on multi-terabyte datasets. Configurations and instructions for T5-like and GPT-like models are also provided. The libraries can be found at https://github.com/google-research/t5x and https://github.com/google/seqio. Adam Roberts, Hyung Won Chung, Anselm Levskaya, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio B. Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Kathleen Kenealy, Kehang Han, Michelle Casbon, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Tachard Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, Andrea Gesmundo |
J. Mach. Learn. Res. | 24 |
| 2022 | Exploring Dual Encoder Architectures for Question AnsweringabstractDual encoders have been used for questionanswering (QA) and information retrieval (IR) tasks with good results.Previous research focuses on two major types of dual encoders, Siamese Dual Encoder (SDE), with parameters shared across two encoders, and Asymmetric Dual Encoder (ADE), with two distinctly parameterized encoders.In this work, we explore different ways in which the dual encoder can be structured, and evaluate how these differences can affect their efficacy in terms of QA retrieval tasks.By evaluating on MS MARCO, open domain NQ and the Mul-tiReQA benchmarks, we show that SDE performs significantly better than ADE.We further propose three different improved versions of ADEs by sharing or freezing parts of the architectures between two encoder towers.We find that sharing parameters in projection layers would enable ADEs to perform competitively with or outperform SDEs.We further explore and explain why parameter sharing in projection layer significantly improves the efficacy of the dual encoders, by directly probing the embedding spaces of the two encoder towers with t-SNE algorithm (van der Maaten and Hinton, 2008). Jianmo Ni, Dan Bikel, Enrique Alfonseca, Chen Qu 0001, Imed Zitouni |
EMNLP | 2 |
| 2022 | SHARE: a System for Hierarchical Assistive Recipe EditingabstractThe large population of home cooks with dietary restrictions is under-served by existing cooking resources and recipe generation models.To help them, we propose the task of controllable recipe editing: adapt a base recipe to satisfy a user-specified dietary constraint.This task is challenging, and cannot be adequately solved with human-written ingredient substitution rules or existing end-to-end recipe generation models.We tackle this problem with SHARE: a System for Hierarchical Assistive Recipe Editing, which performs simultaneous ingredient substitution before generating natural-language steps using the edited ingredients.By decoupling ingredient and step editing, our step generator can explicitly integrate the available ingredients.Experiments on the novel RecipePairs dataset-83K pairs of similar recipes where each recipe satisfies one of seven dietary constraints-demonstrate that SHARE produces convincing, coherent recipes that are appropriate for a target dietary constraint.We further show through human evaluations and real-world cooking trials that recipes edited by SHARE can be easily followed by home cooks to create appealing dishes. Yufei Li 0001, Jianmo Ni, Julian J. McAuley |
EMNLP | 3 |
| 2022 | Large Dual Encoders Are Generalizable RetrieversabstractJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, Yinfei Yang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jianmo Ni, Chen Qu 0001, Jing Lu 0014, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma 0004, Vincent Y. Zhao, Yi Luan, Keith B. Hall, Ming-Wei Chang, Yinfei Yang |
EMNLP | 1 |
| 2022 | ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran 0002, Dara Bahri, Jianmo Ni, Jai Gupta 0001, Kai Hui 0001, Sebastian Ruder, Donald Metzler |
ICLR | 10 |
| 2022 | Transformer Memory as a Differentiable Search IndexabstractIn this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of the model. To this end, we introduce the Differentiable Search Index (DSI), a new paradigm that learns a text-to-text model that maps string queries directly to relevant docids; in other words, a DSI model answers queries directly using only its parameters, dramatically simplifying the whole retrieval process. We study variations in how documents and their identifiers are represented, variations in training procedures, and the interplay between models and corpus sizes. Experiments demonstrate that given appropriate design choices, DSI significantly outperforms strong baselines such as dual encoder models. Moreover, DSI demonstrates strong generalization capabilities, outperforming a BM25 baseline in a zero-shot setup. Yi Tay, Vinh Q. Tran 0002, Mostafa Dehghani 0001, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin 0001, Kai Hui 0001, Zhe Zhao 0001, Jai Gupta 0001, Tal Schuster, William W. Cohen, Donald Metzler |
NeurIPS | 4 |
| 2021 | Multi-stage Training with Improved Negative Contrast for Neural Passage RetrievalabstractIn the context of neural passage retrieval, we study three promising techniques: synthetic data generation, negative sampling, and fusion. We systematically investigate how these techniques contribute to the performance of the retrieval system and how they complement each other. We propose a multi-stage framework comprising of pre-training with synthetic data, fine-tuning with labeled data, and negative sampling at both stages. We study six negative sampling strategies and apply them to the fine-tuning stage and, as a noteworthy novelty, to the synthetic data that we use for pre-training. Also, we explore fusion methods that combine negatives from different strategies. We evaluate our system using two passage retrieval tasks for open-domain QA and using MS MARCO. Our experiments show that augmenting the negative contrast in both stages is effective to improve passage retrieval accuracy and, importantly, they also show that synthetic data generation and negative sampling have additive benefits. Moreover, using the fusion of different kinds allows us to reach performance that establishes a new state-of-the-art level in two of the tasks we evaluated. Jing Lu 0014, Gustavo Hernández Ábrego, Ji Ma 0004, Jianmo Ni, Yinfei Yang |
EMNLP (1) | 4 |
| 2020 | Interview: Large-scale Modeling of Media Dialog with Discourse Patterns and Knowledge GroundingabstractIn this work, we perform the first large-scale analysis of discourse in media dialog and its impact on generative modeling of dialog turns, with a focus on interrogative patterns and use of external knowledge.Discourse analysis can help us understand modes of persuasion, entertainment, and information elicitation in such settings, but has been limited to manual review of small corpora.We introduce Interview-a large-scale (105K conversations) media dialog dataset collected from news interview transcripts-which allows us to investigate such patterns at scale.We present a dialog model that leverages external knowledge as well as dialog acts via auxiliary losses and demonstrate that our model quantitatively and qualitatively outperforms strong discourse-agnostic baselines for dialog modeling-generating more specific and topical responses in interview-style conversations. Bodhisattwa Prasad Majumder, Jianmo Ni, Julian J. McAuley |
EMNLP (1) | 3 |
| 2020 | Addressing Marketing Bias in Product RecommendationsabstractModern collaborative filtering algorithms seek to provide personalized product recommendations by uncovering patterns in consumer-product interactions. However, these interactions can be biased by how the product is marketed, for example due to the selection of a particular human model in a product image. These correlations may result in the underrepresentation of particular niche markets in the interaction data; for example, a female user who would potentially like motorcycle products may be less likely to interact with them if they are promoted using stereotypically 'male' images. Mengting Wan, Jianmo Ni, Rishabh Misra, Julian J. McAuley |
WSDM | 2 |
| 2019 | Generating Personalized Recipes from Historical User PreferencesabstractBodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni, Julian McAuley. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bodhisattwa Prasad Majumder, Jianmo Ni, Julian J. McAuley |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Justifying Recommendations using Distantly-Labeled Reviews and Fine-Grained AspectsabstractJianmo Ni, Jiacheng Li, Julian McAuley. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jianmo Ni, Jiacheng Li 0003, Julian J. McAuley |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Scalable and Accurate Dialogue State Tracking via Hierarchical Sequence GenerationabstractLiliang Ren, Jianmo Ni, Julian McAuley. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Liliang Ren, Jianmo Ni, Julian J. McAuley |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Modeling Heart Rate and Activity Data for Personalized Fitness RecommendationabstractActivity logs collected from wearable devices (e.g. Apple Watch, Fitbit, etc.) are a promising source of data to facilitate a wide range of applications such as personalized exercise scheduling, workout recommendation, and heart rate anomaly detection. However, such data are heterogeneous, noisy, diverse in scale and resolution, and have complex interdependencies, making them challenging to model. In this paper, we develop context-aware sequential models to capture the personalized and temporal patterns of fitness data. Specifically, we propose FitRec - an LSTM-based model that captures two levels of context information: context within a specific activity, and context across a user's activity history. We are specifically interested in (a) estimating a user's heart rate profile for a candidate activity; and (b) predicting and recommending suitable activities on this basis. We evaluate our model on a novel dataset containing over 250 thousand workout records coupled with hundreds of millions of parallel sensor measurements (e.g. heart rate, GPS) and metadata. We demonstrate that the model is able to learn contextual, personalized, and activity-specific dynamics of users' heart rate profiles during exercise. We evaluate the proposed model against baselines on several personalized recommendation tasks, showing the promise of using wearable data for activity modeling and recommendation. Jianmo Ni, Larry Muhlstein, Julian J. McAuley |
WWW | 1 |
| 2017 | Estimating Reactions and Recommending Products with Generative Models of ReviewsabstractTraditional approaches to recommendation focus on learning from large volumes of historical feedback to estimate simple numerical quantities (Will a user click on a product? Make a purchase? etc.). Natural language approaches that model information like product reviews have proved to be incredibly useful in improving the performance of such methods, as reviews provide valuable auxiliary information that can be used to better estimate latent user preferences and item properties. In this paper, rather than using reviews as an inputs to a recommender system, we focus on generating reviews as the model’s output. This requires us to efficiently model text (at the character level) to capture the preferences of the user, the properties of the item being consumed, and the interaction between them (i.e., the user’s preference). We show that this can model can be used to (a) generate plausible reviews and estimate nuanced reactions; (b) provide personalized rankings of existing reviews; and (c) recommend existing products more effectively. Jianmo Ni, Zachary C. Lipton, Sharad Vikram, Julian J. McAuley |
IJCNLP(1) | 1 |
| 2017 | Interconnection Allocation Between Functional Units and Registers in High-Level SynthesisabstractData path interconnection on VLSI chips usually consumes a significant amount of both power and area. In this paper, we focus on the port assignment problem for binary commutative operators for interconnection complexity reduction. First, the port assignment problem is formulated on a constraint graph, and a practical method is proposed to find a valid and initial solution. For solution optimization, an elementary spanning-tree-transformation-based local search algorithm is proposed. To improve the efficiency of optimization, a matrix formulation, which meets the simplex tabuleau format, is proposed and thus the simplex method is adopted for optimization. Moreover, operation pivoting and successive pivoting are discussed for algorithm speedup. The experimental results show that on the randomly generated test cases, the matrix-based algorithm shows the highest solution optimality and is five times faster than the elementary transformation method. On the real high-level synthesis benchmarks, the matrix-based method reduced 14% interconnections, while the previous greedy algorithm reduced 8% on average. Cong Hao, Jianmo Ni, Nan Wang 0003, Takeshi Yoshimura |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |