Liam Collins

dblp:170/1157 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
abstract
Xueying Ding, Xingyue Huang, Mingxuan Ju, Liam Collins, Yozen Liu, Leman Akoglu, Neil Shah, Tong Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xueying Ding, Xingyue Huang, Mingxuan Ju, Liam Collins, Yozen Liu, Leman Akoglu, Neil Shah
ACL (1)4
2026 Semantic IDs for Recommender Systems at Snapchat: Use Cases, Technical Challenges, and Design Choices
abstract
Effective item identifiers (IDs) are an important component for recommender systems (RecSys) in practice, and are commonly adopted in many use cases such as retrieval and ranking. IDs can encode collaborative filtering signals within training data, such that RecSys models can extrapolate during the inference and personalize the prediction based on users' behavioral histories. Recently, Semantic IDs (SIDs) have become a trending paradigm for RecSys. In comparison to the conventional atomic ID, an SID is an ordered list of codes, derived from tokenizers such as residual quantization, applied to semantic representations commonly extracted from foundation models or collaborative signals. SIDs have drastically smaller cardinality than the atomic counterpart, and induce semantic clustering in the ID space. At Snapchat, we apply SIDs as auxiliary features for ranking models, and also explore SIDs as additional retrieval sources in different ML applications. In this paper, we discuss practical technical challenges we encountered while applying SIDs, experiments we have conducted, and design choices we have iterated to mitigate these challenges. Backed by promising offline results on both internal data and academic benchmarks as well as online A/B studies, SID variants have been launched in multiple production models with positive metrics impact.
Clark Mingxuan Ju, Tong Zhao 0003, Leonardo Neves, Liam Collins, Bhuvesh Kumar, Jiwen Ren, Wenfeng Zhuo, Jinchao Li, Karthik Iyer, Peicheng Yu, Manish Malik, Neil Shah
SIGIR4
2026 Sequential Data Augmentation for Generative Recommendation
abstract
Generative recommendation plays a crucial role in personalized systems, predicting users' future interactions from their historical behavior sequences. A critical yet underexplored factor in training these models is data augmentation, the process of constructing training data from user interaction histories. By shaping the training distribution, data augmentation directly and often substantially affects model generalization and performance. Nevertheless, in much of the existing work, this process is simplified, applied inconsistently, or treated as a minor design choice, without a systematic and principled understanding of its effects.
Bhuvesh Kumar, Mingxuan Ju, Tong Zhao 0003, Kijung Shin, Neil Shah, Liam Collins
WSDM7
2025 Generative Recommendation with Semantic IDs: A Practitioner's Handbook
abstract
Generative recommendation (GR) has gained increasing attention for its promising performance compared to traditional models. A key factor contributing to the success of GR is the semantic ID (SID), which converts continuous semantic representations (e.g., from large language models) into discrete ID sequences. However, varied modeling techniques, hyper-parameters, and experimental setups in existing literature make direct comparisons between GR proposals challenging. Furthermore, the absence of an open-source, unified framework hinders systematic benchmarking and extension, slowing model iteration. To address this challenge, our work introduces and open-sources a framework for Generative Recommendation with semantic ID, namely GRID, specifically designed for modularity to facilitate easy component swapping and accelerate idea iteration. Using GRID, we systematically experiment with and ablate different components of GR models with SIDs on public benchmarks. Our comprehensive experiments with GRID reveal that many overlooked architectural components in GR models with SIDs substantially impact performance. This offers both novel insights and validates the utility of an open-source platform for robust benchmarking and GR research advancement. GRID is open-sourced at https://github.com/snap-research/GRID.
Clark Mingxuan Ju, Liam Collins, Leonardo Neves, Bhuvesh Kumar, Louis Yufeng Wang, Tong Zhao 0003, Neil Shah
CIKM2
2025 Revisiting Self-attention for Cross-domain Sequential Recommendation
abstract
Sequential recommendation is a popular paradigm in modern recommender systems. In particular, one challenging problem in this space is cross-domain sequential recommendation (CDSR), which aims to predict future behaviors given user interactions across multiple domains. Existing CDSR frameworks are mostly built on the self-attention transformer and seek to improve by explicitly injecting additional domain-specific components (e.g. domain-aware module blocks). While these additional components help, we argue they overlook the core self-attention module already present in the transformer, a naturally powerful tool to learn correlations among behaviors. In this work, we aim to improve the CDSR performance for simple models from a novel perspective of enhancing the self-attention. Specifically, we introduce a Pareto-optimal self-attention and formulate the cross-domain learning as a multi-objective problem, where we optimize the recommendation task while dynamically minimizing the cross-domain attention scores. Our approach automates knowledge transfer in CDSR (dubbed as AutoCDSR) - it not only mitigates negative transfer but also encourages complementary knowledge exchange among auxiliary domains. Based on the idea, we further introduce AutoCDSR+, a more performant variant with slight additional cost. Our proposal is easy to implement and works as a plug-and-play module that can be incorporated into existing transformer-based recommenders. Besides flexibility, it is practical to deploy because it brings little extra computational overheads without heavy hyper-parameter tuning. We conduct experiments over both large-scale production recommender data as well as academic benchmarks, where AutoCDSR consistently enhances the performance of base transformers, enabling simple models to perform on par with state-of-the-art with less overhead (e.g., 4x faster than state-of-the-art CDSR models). AutoCDSR on average improves Recall@10 for SASRec and Bert4Rec by 9.8% and 16.0% and NDCG@10 by 12.0% and 16.7%, respectively. Code is available at https://github.com/snap-research/AutoCDSR.
Clark Mingxuan Ju, Leonardo Neves, Bhuvesh Kumar, Liam Collins, Tong Zhao 0003, Yuwei Qiu, Qing Dou, Sohail Nizam, Sen Yang 0026, Neil Shah
KDD (2)4
2025 Provable Meta-Learning with Low-Rank Adaptations
abstract
The power of foundation models (FMs) lies in their capacity to learn highly expressive representations that can be adapted to a broad spectrum of tasks. However, these pretrained models require additional training stages to become effective for downstream applications. In the multi-task setting, prior works have shown empirically that specific meta-learning approaches for preparing a model for future adaptation through parameter-efficient fine-tuning (PEFT) can outperform standard retraining methods, but the mechanism of the benefits of meta-learning has been largely unexplored. We introduce a framework for generic PEFT-based meta-learning to learn a model that can easily adapt to unseen tasks. For linear models using LoRA, we show that standard retraining is provably suboptimal for finding an adaptable set of parameters and provide strict performance guarantees for our proposed method. We verify these theoretical insights through experiments on synthetic data as well as real-data vision and language tasks. We observe significant performance benefits using a simple implementation of our proposed meta-learning scheme during retraining relative to the conventional approach.
Jacob L. Block, Sundararajan Srinivasan, Liam Collins, Aryan Mokhtari, Sanjay Shakkottai
NeurIPS3
2025 Learning Universal User Representations Leveraging Cross-domain User Intent at Snapchat
abstract
The development of powerful user representations is a key factor in the success of recommender systems (RecSys). Online platforms employ a range of RecSys techniques to personalize user experience across diverse in-app surfaces. User representations are often learned individually through user's historical interactions within each surface and user representations across different surfaces can be shared post-hoc as auxiliary features or additional retrieval sources. While effective, such schemes cannot directly encode collaborative filtering signals across different surfaces, hindering its capacity to discover complex relationships between user behaviors and preferences across the whole platform. To bridge this gap at Snapchat, we seek to conduct universal user modeling (UUM) across different in-app surfaces, learning general-purpose user representations which encode behaviors across surfaces. Instead of replacing domain-specific representations, UUM representations capture cross-domain trends, enriching existing representations with complementary information. This work discusses our efforts in developing initial UUM versions, practical challenges, technical choices and modeling and research directions with promising offline performance. Following successful A/B testing, UUM representations have been launched in production, powering multiple use cases and demonstrating their value. UUM embedding has been incorporated into (i) Long-form Video embedding-based retrieval, leading to 2.78% increase in Long-form Video Open Rate, (ii) Long-form Video L2 ranking, with 19.2% increase in Long-form Video View Time sum, (iii) Lens L2 ranking, leading to 1.76% increase in Lens play time, and (iv) Notification L2 ranking, with 0.87% increase in Notification Open Rate.
Clark Mingxuan Ju, Leonardo Neves, Bhuvesh Kumar, Liam Collins, Tong Zhao 0003, Yuwei Qiu, Qing Dou, Yang Zhou 0063, Sohail Nizam, Rengim Ozturk, Yvette Liu, Sen Yang 0021, Manish Malik, Neil Shah
SIGIR4
2024 Image Segmentation of Bacterial Biofilms to Study Pathogen-Surface Interactions
abstract
Biofilms are structured communities of microbial cells encased in a self-produced extracellular polymeric substance (EPS) matrix, which adhere to biotic or abiotic surfaces and are ubiquitous in natural, industrial, and clinical settings. Biofilms play critical roles in various ecosystems and pose significant challenges in healthcare due to their resistance to antibiotics. Understanding biofilms is complex due to their heterogeneous nature and dynamic behavior. Biofilms exhibit spatial and temporal variations in their structure, composition, and function. These variations are heavily influenced by environmental conditions, microbial species involved, and the physical and chemical properties of the surfaces they colonize. With the biofilm formation occurring in multiple stages a deeper understanding of the fundamental mechanisms for attachment and the subsequent cascade of physical and chemical events that promote biofilm growth and propagation is essential for developing materials that resist biofilm formation and enhance antimicrobial therapies across various fields. However, the properties that can be collected change dynamically from small to large area biofilm imaging. While small area biofilm studies reveal details of individual cell structures, large area biofilm studies can help in studying dynamic changes in biofilms which have a potential for commercial application.Atomic Force Microscopy (AFM) is a valuable tool in biofilm research, offering detailed structural and mechanical insights. It enables imaging of cell structures, interactions, and mechanical properties of biofilms with nanoscale resolution. By scanning a sharp probe over the biofilm surface and measuring the forces between the probe and the sample, AFM can achieve nanometer-scale resolution, revealing detailed structural features of biofilms. Higher-resolution imaging techniques like AFM are crucial for studying individual cell structures, fine features, and appendages such as flagella and pili. These advanced imaging methods provide detailed insights into cellular morphology and interactions that are often not visible with lower-resolution techniques. However, AFM has its own limitations. AFM has a slow scanning process and typically has small scanning ranges of up to 100 μm × 100 μm limiting its ability to capture dynamic changes in biofilms across macroscopic scales. Currently, we have used an automated scanning platform to image large macro-scale (~mm) biofilms with nanoscale resolution, which equates to a large amount of data to be processed manually. Therefore, automated methods of image processing and analysis are necessary, to capture the growth patterns of biofilms.In this study, we integrate AFM with multiple cutting edge AI techniques and developing methods for AFM imaging will enhance the ability to study biofilms comprehensively, leading to significant advancements in this field. Here we present the role of computer vision-based image segmentation techniques in capturing topological properties of bacteria in biofilms which provide valuable insights about pathogen surface interactions. To this end we used You Only Look Once (YOLO v8) image segmentation pretrained models to perform segmentation of AFM bacterial images of dried films of Pantoea sp. YR343. This is a gram-negative bacterium known for promoting plant growth. Previous research has shown that Pantoea sp. YR343 forms biofilms with a honeycomb morphology on hydrophobic surfaces, but it does not readily attach to hydrophilic surfaces. We analyzed high resolution Panteoa bacterial AFM images to quantify morphological characteristics of the bacteria that can provide insights of the attachment patterns of the bacteria on abiotic surfaces. Using the image segmentation models we captures several morphological properties such as total cell count, cell orientation, cell eccentricity, cell perimeter and cell area for every bacterium in the image. The images collected from AFM are in the Gwydion (.gwy) format which are initially converted into regular portable network graph (png) format. These files were then flattened using a combination of thresholding methods to even out the tilt in the images which is a pre-requisite for analyzing AFM images. The flattened images are then normalized and fed into the image analysis pipeline performed using the Roboflow platform. The pipeline includes steps like annotation, data split, image resizing and data augmentation. The obtained dataset is fed into the pre-trained Yolov8 image segmentation model which is widely applied to perform image segmentation tasks. The model gave a validation Precision and Recall and of 0.64 and 0.78 respectively and a Mean Average Precision (mAP50) of 0.73. The obtained modes were then deployed to obtain the dimensions of the masks for every bacterium in the image. The masks are then used to calculate cell properties like cell area, count, perimeter, orientation, centroids and eccentricity using Scikit-Image’s Regionprops method. With varying degrees of mask thresholding, every bacterium in the image was identified. The methodology is integrated into the automated analysis pipelines of AFM image analysis we developed in-house. The study demonstrates a methodology to integrate highly versatile image segmentation models to specific AFM image analysis tasks to enable rapid analysis of large- area biofilms at nanoscale resolution. Although promising, we still believe that for YOLO to gain traction in AFM contexts, further customization may be required to fine-tune the model outputs for detecting nanoscale objects and differentiating subtle textures within AFM data.
Sita Sirisha Madugula, Blythe Dumerer, Ruben Millan-Solsona, Checa Nualart Marti, Liam Collins, Rama K. Vasudevan, Retterer Scott, Lance Zhang, Spencer Cox, Jennifer Lynn Morrell-Falvey
IEEE Big Data5
2024 Provable Multi-Task Representation Learning by Two-Layer ReLU Neural Networks
abstract
An increasingly popular machine learning paradigm is to pretrain a neural network (NN) on many tasks offline, then adapt it to downstream tasks, often by re-training only the last linear layer of the network. This approach yields strong downstream performance in a variety of contexts, demonstrating that multitask pretraining leads to effective feature learning. Although several recent theoretical studies have shown that shallow NNs learn meaningful features when either (i) they are trained on a *single* task or (ii) they are *linear*, very little is known about the closer-to-practice case of *nonlinear* NNs trained on *multiple* tasks. In this work, we present the first results proving that feature learning occurs during training with a nonlinear model on multiple tasks. Our key insight is that multi-task pretraining induces a pseudo-contrastive loss that favors representations that align points that typically have the same label across tasks. Using this observation, we show that when the tasks are binary classification tasks with labels depending on the projection of the data onto an $r$-dimensional subspace within the $d\gg r$-dimensional input space, a simple gradient-based multitask learning algorithm on a two-layer ReLU NN recovers this projection, allowing for generalization to downstream tasks with sample and neuron complexity independent of $d$. In contrast, we show that with high probability over the draw of a single task, training on this single task cannot guarantee to learn all $r$ ground-truth features.
Liam Collins, Seyed Hamed Hassani, Mahdi Soltanolkotabi, Aryan Mokhtari, Sanjay Shakkottai
ICML1
2024 In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
abstract
A striking property of transformers is their ability to perform in-context learning (ICL), a machine learning framework in which the learner is presented with a novel context during inference implicitly through some data, and tasked with making a prediction in that context. As such, that learner must adapt to the context without additional training. We explore the role of *softmax* attention in an ICL setting where each context encodes a regression task. We show that an attention unit learns a window that it uses to implement a nearest-neighbors predictor adapted to the landscape of the pretraining tasks. Specifically, we show that this window widens with decreasing Lipschitzness and increasing label noise in the pretraining tasks. We also show that on low-rank, linear problems, the attention unit learns to project onto the appropriate subspace before inference. Further, we show that this adaptivity relies crucially on the softmax activation and thus cannot be replicated by the linear activation often studied in prior theoretical analyses.
Liam Collins, Advait Parulekar, Aryan Mokhtari, Sujay Sanghavi, Sanjay Shakkottai
NeurIPS1
2023 InfoNCE Loss Provably Learns Cluster-Preserving Representations
abstract
The goal of contrasting learning is to learn a representation that preserves underlying clusters by keeping samples with similar content, e.g. the “dogness” of a dog, close to each other in the space generated by the representation. A common and successful approach for tackling this unsupervised learning problem is minimizing the InfoNCE loss associated with the training samples, where each sample is associated with their augmentations (positive samples such as rotation, crop) and a batch of negative samples (unrelated samples). To the best of our knowledge, it was unanswered if the representation learned by minimizing the InfoNCE loss preserves the underlying data clusters, as it only promotes learning a representation that is faithful to augmentations, i.e., an image and its augmentations have the same representation. Our main result is to show that the representation learned by InfoNCE with a finite number of negative samples is also consistent with respect to {\em clusters} in the data, under the condition that the augmentation sets within clusters may be non-overlapping but are close and intertwined, relative to the complexity of the learning function class.
Advait Parulekar, Liam Collins, Karthikeyan Shanmugam 0001, Aryan Mokhtari, Sanjay Shakkottai
COLT2
2023 Meta-Learning for Image-Guided Millimeter-Wave Beam Selection in Unseen Environments
abstract
The use of alternate modalities, like images, for fast beamforming in the millimeter wave (mmWave)-band is being proposed to ensure high bandwidth connectivity in vehicular scenarios typically seen in the context of autonomous cars. Considering the dynamic deployment conditions, a car may encounter new environments which were not explicitly included in an apriori training dataset. In this paper, we propose to use the Model-Agnostic Meta-Learning (MAML) framework on the image data of the mmWave vehicle-to-infrastructure beam selection FLASH dataset, to overcome the generalization issues of a pre-trained model in unseen non-line-of-sight (NLOS) connectivity environments. MAML has additional advantages over traditional deep-learning techniques: (i) it uses a fraction of the data which, in turn, simplifies data collection and storage, and (ii) it results in equal or higher accuracy in optimal beam selection compared to the case when the new environment dataset is fully available during initial training. We show that our MAML implementation improves test accuracy of beam selection by up to 86% with fine-tuning when encountering an unseen NLOS environment compared to conventional supervised learning.
Jerry Gu, Liam Collins, Debashri Roy, Aryan Mokhtari, Sanjay Shakkottai, Kaushik R. Chowdhury
ICASSP2
2022 MAML and ANIL Provably Learn Representations
abstract
Recent empirical evidence has driven conventional wisdom to believe that gradient-based meta-learning (GBML) methods perform well at few-shot learning because they learn an expressive data representation that is shared across tasks. However, the mechanics of GBML have remained largely mysterious from a theoretical perspective. In this paper, we prove that two well-known GBML methods, MAML and ANIL, as well as their first-order approximations, are capable of learning common representation among a set of given tasks. Specifically, in the well-known multi-task linear representation learning setting, they are able to recover the ground-truth representation at an exponentially fast rate. Moreover, our analysis illuminates that the driving force causing MAML and ANIL to recover the underlying representation is that they adapt the final layer of their model, which harnesses the underlying task diversity to improve the representation in all directions of interest. To the best of our knowledge, these are the first results to show that MAML and/or ANIL learn expressive representations and to rigorously explain why they do so.
Liam Collins, Aryan Mokhtari, Sewoong Oh, Sanjay Shakkottai
ICML1
2022 FedAvg with Fine Tuning: Local Updates Lead to Representation Learning
abstract
The Federated Averaging (FedAvg) algorithm, which consists of alternating between a few local stochastic gradient updates at client nodes, followed by a model averaging update at the server, is perhaps the most commonly used method in Federated Learning. Notwithstanding its simplicity, several empirical studies have illustrated that the model output by FedAvg leads to a model that generalizes well to new unseen tasks after a few fine-tuning steps. This surprising performance of such a simple method, however, is not fully understood from a theoretical point of view. In this paper, we formally investigate this phenomenon in the multi-task linear regression setting. We show that the reason behind the generalizability of the FedAvg output is FedAvg’s power in learning the common data representation among the clients’ tasks, by leveraging the diversity among client data distributions via multiple local updates between communication rounds. We formally establish the iteration complexity required by the clients for proving such result in the setting where the underlying shared representation is a linear map. To the best of our knowledge, this is the first result showing that FedAvg learns an expressive representation in any setting. Moreover, we show that multiple local updates between communication rounds are necessary for representation learning, as distributed gradient methods that make only one local update between rounds provably cannot recover the ground-truth representation in the linear setting, and empirically yield neural network representations that generalize drastically worse to new clients than those learned by FedAvg trained on heterogeneous image classification datasets.
Liam Collins, Seyed Hamed Hassani, Aryan Mokhtari, Sanjay Shakkottai
NeurIPS1
2021 Exploiting Shared Representations for Personalized Federated Learning
abstract
Deep neural networks have shown the ability to extract universal feature representations from data such as images and text that have been useful for a variety of learning tasks. However, the fruits of representation learning have yet to be fully-realized in federated settings. Although data in federated settings is often non-i.i.d. across clients, the success of centralized deep learning suggests that data often shares a global {\em feature representation}, while the statistical heterogeneity across clients or tasks is concentrated in the {\em labels}. Based on this intuition, we propose a novel federated learning framework and algorithm for learning a shared data representation across clients and unique local heads for each client. Our algorithm harnesses the distributed computational power across clients to perform many local-updates with respect to the low-dimensional local parameters for every update of the representation. We prove that this method obtains linear convergence to the ground-truth representation with near-optimal sample complexity in a linear setting, demonstrating that it can efficiently reduce the problem dimension for each client. Further, we provide extensive experimental results demonstrating the improvement of our method over alternative personalized federated learning approaches in heterogeneous settings.
Liam Collins, Seyed Hamed Hassani, Aryan Mokhtari, Sanjay Shakkottai
ICML1
2020 Task-Robust Model-Agnostic Meta-Learning
abstract
Meta-learning methods have shown an impressive ability to train models that rapidly learn new tasks. However, these methods only aim to perform well in expectation over tasks coming from some particular distribution that is typically equivalent across meta-training and meta-testing, rather than considering worst-case task performance. In this work we introduce the notion of ``task-robustness'' by reformulating the popular Model-Agnostic Meta-Learning (MAML) objective \citep{finn2017model} such that the goal is to minimize the maximum loss over the observed meta-training tasks. The solution to this novel formulation is task-robust in the sense that it places equal importance on even the most difficult and/or rare tasks. This also means that it performs well over all distributions of the observed tasks, making it robust to shifts in the task distribution between meta-training and meta-testing. We present an algorithm to solve the proposed min-max problem, and show that it converges to an $\epsilon$-accurate point at the optimal rate of $\mathcal{O}(1/\epsilon^2)$ in the convex setting and to an $(\epsilon, \delta)$-stationary point at the rate of $\mathcal{O}(\max\{1/\epsilon^5, 1/\delta^5\})$ in nonconvex settings. We also provide an upper bound on the new task generalization error that captures the advantage of minimizing the worst-case task loss, and demonstrate this advantage in sinusoid regression and image classification experiments.
Liam Collins, Aryan Mokhtari, Sanjay Shakkottai
NeurIPS1