VLDB 2026 Research / reviewers in the wild / expert
Manish Malik
dblp:169/0346
· DBLP profile ↗
6ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic IDs for Recommender Systems at Snapchat: Use Cases, Technical Challenges, and Design ChoicesabstractEffective item identifiers (IDs) are an important component for recommender systems (RecSys) in practice, and are commonly adopted in many use cases such as retrieval and ranking. IDs can encode collaborative filtering signals within training data, such that RecSys models can extrapolate during the inference and personalize the prediction based on users' behavioral histories. Recently, Semantic IDs (SIDs) have become a trending paradigm for RecSys. In comparison to the conventional atomic ID, an SID is an ordered list of codes, derived from tokenizers such as residual quantization, applied to semantic representations commonly extracted from foundation models or collaborative signals. SIDs have drastically smaller cardinality than the atomic counterpart, and induce semantic clustering in the ID space. At Snapchat, we apply SIDs as auxiliary features for ranking models, and also explore SIDs as additional retrieval sources in different ML applications. In this paper, we discuss practical technical challenges we encountered while applying SIDs, experiments we have conducted, and design choices we have iterated to mitigate these challenges. Backed by promising offline results on both internal data and academic benchmarks as well as online A/B studies, SID variants have been launched in multiple production models with positive metrics impact. Clark Mingxuan Ju, Tong Zhao 0003, Leonardo Neves, Liam Collins, Bhuvesh Kumar, Jiwen Ren, Wenfeng Zhuo, Jinchao Li, Karthik Iyer, Peicheng Yu, Manish Malik, Neil Shah |
SIGIR | 17 |
| 2026 | LLM-Enhanced Topical Trend Detection at SnapchatabstractAutomatic detection of topical trends at scale is both challenging and essential for maintaining a dynamic content ecosystem on social media platforms. In this work, we present a large-scale system for identifying emerging topical trends on Snapchat, one of the world's largest short-video social platforms. Our system integrates multimodal topic extraction, time-series burst detection, and LLM-based consolidation and enrichment to enable accurate and timely trend discovery. To the best of our knowledge, this is the first published end-to-end system for topical trend detection on short-video platforms at production scale. Continuous offline human evaluation over six months demonstrates high precision in identifying meaningful trends. The system has been deployed in production at global scale and applied to downstream surfaces including content ranking and search, driving measurable improvements in content freshness and user experience. Hangqi Zhao, Jay Li, Abhiruchi Bhattacharya, Cong Ni, Jason Yeung, Jinchao Ye, Akshat Malu, Manish Malik |
SIGIR | 9 |
| 2025 | Learning Universal User Representations Leveraging Cross-domain User Intent at SnapchatabstractThe development of powerful user representations is a key factor in the success of recommender systems (RecSys). Online platforms employ a range of RecSys techniques to personalize user experience across diverse in-app surfaces. User representations are often learned individually through user's historical interactions within each surface and user representations across different surfaces can be shared post-hoc as auxiliary features or additional retrieval sources. While effective, such schemes cannot directly encode collaborative filtering signals across different surfaces, hindering its capacity to discover complex relationships between user behaviors and preferences across the whole platform. To bridge this gap at Snapchat, we seek to conduct universal user modeling (UUM) across different in-app surfaces, learning general-purpose user representations which encode behaviors across surfaces. Instead of replacing domain-specific representations, UUM representations capture cross-domain trends, enriching existing representations with complementary information. This work discusses our efforts in developing initial UUM versions, practical challenges, technical choices and modeling and research directions with promising offline performance. Following successful A/B testing, UUM representations have been launched in production, powering multiple use cases and demonstrating their value. UUM embedding has been incorporated into (i) Long-form Video embedding-based retrieval, leading to 2.78% increase in Long-form Video Open Rate, (ii) Long-form Video L2 ranking, with 19.2% increase in Long-form Video View Time sum, (iii) Lens L2 ranking, leading to 1.76% increase in Lens play time, and (iv) Notification L2 ranking, with 0.87% increase in Notification Open Rate. Clark Mingxuan Ju, Leonardo Neves, Bhuvesh Kumar, Liam Collins, Tong Zhao 0003, Yuwei Qiu, Qing Dou, Yang Zhou 0063, Sohail Nizam, Rengim Ozturk, Yvette Liu, Sen Yang 0021, Manish Malik, Neil Shah |
SIGIR | 13 |
| 2020 | Investigating teams of neuro-typical and neuro-atypical students learning together using COGLE: A multi case studyabstractThis Work in Progress Research paper aims to contribute to theories relevant to trust, self-efficacy and team effectiveness in engineering student teams. Self-efficacy and trust in teammates are both crucial for team effectiveness. Borrego et al. in a review on team effectiveness within engineering education have highlighted the scarcity of research on psychological constructs, such as trust. This work was inspired by their call for more research that connects engineering education research with the industrial and organisational psychology literature to improve engineering education practice and the outcomes relating to team working. Team working depends on social and communication skills of individual teammates. However, collaborative teams can experience socio-communication challenges. These can be even more pronounced in neuro-atypical (NT) students. With an increasing number of students, hidden or diagnosed, who are neurologically atypical (NAT) within engineering courses investigating ways to support development of trust and self-efficacy has become even more important. Using two real-world case studies, the efficacy of the Computer Orchestrated Group Learning Environment (COGLE), a novel software intervention that supports the development of trust and self-efficacy of individuals in teams of neuro-typical and neuro-atypical students, is investigated using qualitative and quantitative methods. In particular to answer the two research questions: 1. How does the use of COGLE affect the self-efficacy of NT and NAT engineering students learning together? 2. How does the use of COGLE affect the development of trust between a team of NT and NAT engineering students learning together? The case studies show how COGLE can be used within two pedagogical approaches: Flipped Classroom and Project Based Learning, which are commonly used in engineering education. The learning gain data and related effect sizes from both cases show that COGLE was successful in enhancing self-efficacy in all students. Furthermore, both cases show three very interesting results relating to trust: firstly, the teammates developed trust in each other in just 4 two-hour sessions; secondly, the students, including the neuro-atypical students, were able to correct their trust due to varied interactions enabled by COGLE; and finally, as trust and self-efficacy was enhanced before students were asked to work together on a collaborative activity, it helped both neuro-typical and neuro-atypical students to be fully involved in team work, thereby improving the team's effectiveness. The implication for practice is that COGLE can be used to effectively prepare all students for as shown by learning gain and increased levels of trust and enhance team effectiveness. Manish Malik, Julie-Ann Sime |
FIE | 1 |
| 2018 | Conversational Query Understanding Using Sequence to Sequence ModelingabstractUnderstanding conversations is crucial to enabling conversational search in technologies such as chatbots, digital assistants, and smart home devices that are becoming increasingly popular. Conventional search engines are powerful at answering open domain queries but are mostly capable of stateless search. In this paper, we define a conversational query as a query that depends on the context of the current conversation, and we formulate the conversational query understanding problem as context-aware query reformulation, where the goal is to reformulate the conversational query into a search engine friendly query in order to satisfy users» information needs in conversational settings. Such context-aware query reformulation problem lends itself to sequence to sequence modeling. We present a large scale open domain dataset of conversational queries and various sequence to sequence models that are learned from this dataset. The best model correctly reformulates over half of all conversational queries, showing the potential of sequence to sequence modeling for this task. Gary Ren, Xiaochuan Ni 0001, Manish Malik, Qifa Ke |
WWW | 3 |
| 2014 | Understanding the use of paper and online logbooks for final year undergraduate engineering projectsabstractIn industry an engineer is often required to keep a logbook for recording developments within projects. In higher education, logbooks are a commonly used tool thought to be one that encourages active independent learning and reflective thinking. In School of Engineering, at University of Portsmouth, paper and more recently online logbooks have been in use for recording work for final year projects and project based learning tasks. The work presented here benefits from a unique opportunity within the School of Engineering, where online logbooks alongside traditional paper based logbooks are being used within final year projects. A recent cohort of students (N=127) on ENG600 project module was given the option, through their Supervisors, to use paper logbooks and or online logbooks for recording their work. This work aims to investigate the use of both paper and online logbooks. A mix of Qualitative Research methods and quantitative techniques will be used in this project. The use of content analysis will provide an insight into student reflections and their motivations for using their logbook. Furthermore focus groups, involving live editing of documents in an individual and collaborative fashion, will be used to gather more data for analysis. Quantitative methods (questionnaire, analytics and quantitative content analysis) will also be used in this study. When this work is completed, it will provide guidance and comparison on using the two types of logbooks, backed by knowledge of student motivations and approaches. Manish Malik |
FIE | 1 |