Chris Newell

dblp:68/11509 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
abstract
Today’s most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new datasets called PixMo, including a dataset of highly detailed image captions for pre-training, a free-form image Q&A dataset for fine-tuning, and an innovative 2D pointing dataset, all collected without the use of external VLMs. The success of our approach relies on careful modeling choices, a well- tuned training pipeline, and, most critically, the quality of our newly collected datasets. Our best-in-class 72B model not only outperforms others in the class of open weight and data models, but also outperforms larger proprietary models including Claude 3.5 Sonnet, and Gemini 1.5 Pro and Flash, second only to GPT-4o based on both academic benchmarks and on a large human evaluation. Our model weights, new datasets, and source code are available at https://molmo.allenai.org/blog.
Matt Deitke, Sangho Lee 0008, Rohun Tripathi, Yue Yang 0006, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, Yen-Sung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch, Andrew Head, Rose Hendrix, Favyen Bastani, Eli VanderBilt, Nathan Lambert 0001, Yvonne Chou, Arnavi Chheda, Jenna Sparks, Sam Skjonsberg, Michael Schmitz 0002, Aaron Sarnat, Byron Bischoff, Pete Walsh 0001, Chris Newell, Piper Wolters, Tanmay Gupta, Kuo-Hao Zeng, Jon Borchardt, Dirk Groeneveld, Crystal Nam, Sophie Lebrecht, Caitlin Wittlif, Carissa Schoenick, Oscar Michel, Ranjay Krishna, Luca Weihs, Noah A. Smith, Hannaneh Hajishirzi, Ross B. Girshick, Ali Farhadi, Aniruddha Kembhavi
CVPR33
2024 When Cohesion Lies in the Embedding Space: Embedding-Based Reference-Free Metrics for Topic Segmentation
abstract
In this paper we propose a new framework and new methods for the reference-free evaluation of topic segmentation systems directly in the embedding space. Specifically, we define a common framework for reference-free, embedding-based topic segmentation metrics, and show how this applies to an existing metric. We then define new metrics, based on a previously defined cohesion score, Average Relative Proximity. Using this approach, we show that Large Language Models (LLMs) yield features that, if used correctly, can strongly correlate with traditional topic segmentation metrics based on costly and rare human annotations, while outperforming existing reference-free metrics borrowed from clustering evaluation in most domains. We then show that smaller language models specifically fine-tuned for different sentence-level tasks can outperform LLMs several orders of magnitude larger. Via a thorough comparison of our metric’s performance across different datasets, we see that conversational data present the biggest challenge in this framework. Finally, we analyse the behaviour of our metrics in specific error cases, such as those of under-generation and moving of ground truth topic boundaries, and show that our metrics behave more consistently than other reference-free methods.
Iacopo Ghinassi, Lin Wang 0009, Chris Newell, Matthew Purver
LREC/COLING3
2023 Exploring Pre-Trained Neural Audio Representations for Audio Topic Segmentation
abstract
Recent works have shown that audio embeddings can improve automatic topic segmentation of formats such as radio shows. In this work we expand the work in that direction by showing how and which publicly available, pre-trained neural audio embeddings can perform the task, without the need of any further fine-tuning of the audio encoders. The ranking of the encoders suggest that neural encoders pre-trained for speaker diarization and general purpose audio classification are the best suited to be used as features, beating non-neural baselines. We show that we can obtain perfect results on a newly created random dataset similar to the one used in previous work. We also show for the first time results on real-world data, proving that our method can be applied to actual radio shows with good results, but the choice of audio encoders is extremely important in order to achieve those. Finally, by releasing the datasets we used we make the contribution of providing the first (to our knowledge) publicly available, free of charge datasets for audio topic segmentation of media products.
Iacopo Ghinassi, Matthew Purver, Huy Phan, Chris Newell
ICME4
2023 Semantic and Lexical Token Based Vectors Improve Precision of Recommendations for TV Programmes
abstract
Advances in the digitalisation of data have led to large archives of content in media companies. These archives include multimodal data and metadata associated with each media programme. Relating content across different mediums of data and metadata has thus become an emergent challenge, with applications to popular domains such as programme recommendation. In this paper, we worked with combinations of content similarity measures computed from the distances between different forms of textual data obtained from subtitle files and metadata obtained from the genres of programmes. The different forms of textual representations we considered were neural semantic and topic vectors, and a weighted Jaccard distance encoding lexical token rareness. The late fusion combination of these four distances provided the best recommendation results. For a weekly dataset of 145 TV programmes, it increased the precision of the genre-based recommendations by 5.76%. In a monthly dataset of 906 programmes, it achieved an increase of 1.5%. This combination was more efficient than one with audio and video files.
Taner Cagali, Hadi Wazni, Saba Nazir, Mehrnoosh Sadrzadeh, Chris Newell
ISM5
2023 Multimodal Topic Segmentation of Podcast Shows with Pre-trained Neural Encoders
abstract
We present two multimodal models for topic segmentation of podcasts built on pre-trained neural text and audio embeddings. We show that results can be improved by combining different modalities; but also by combining different encoders from the same modality, especially general-purpose sentence embeddings with specifically fine-tuned ones. We also show that audio embeddings can be substituted with two simple features related to sentence duration and inter-sentential pauses with comparable results. Finally, we publicly release our two datasets, the first in our knowledge publicly and freely available multimodal datasets for topic segmentation.
Iacopo Ghinassi, Lin Wang 0009, Chris Newell, Matthew Purver
ICMR3
2021 Enhancing Personalised Recommendations with the Use of Multimodal Information
abstract
Whenever we watch a TV show or movie, we process a substantial amount of information that is conveyed to us via various multimedia mediums, in particular: visual, textual, and audio. These data signify distinctive properties that aid in creating a unique motion picture experience. In effort to not only produce a more personalised recommender system, but also tackle the problem of popularity bias, we develop a system that incorporates the use of multimodal information. Specifically, we investigate the correlation between features that are extracted using state of the art techniques and deep learning models from visual characteristics, audio patterns and subtitles. The framework is evaluated on a dataset comprising of 145 BBC TV programmes against genre and user baselines. We demonstrate that personalised recommendations can not only be improved with the use of multimodal information, but also outperform genre and user-based models in terms of diversity, whilst maintaining matching levels of accuracy.
Taner Cagali, Mehrnoosh Sadrzadeh, Chris Newell
ISM3
2020 Audiovisual, Genre, Neural and Topical Textual Embeddings for TV Programme Content Representation
abstract
TV programmes have their contents described by multiple means: textual subtitles, audiovisual files, and metadata such as genres. In order to represent these contents, we develop vectorial representations for their low-level multimodal features, group them with simple clustering techniques, and combine them using middle and late fusion. For textual features, we use LSI and Doc2Vec neural embeddings; for audio, MFCC's and Bags of Audio Words; for visual, SIFT, and Bags of Visual Words. We apply our model to a dataset of BBC TV programmes and use a standard recommender and pairwise similarity matrices of content vectors to estimate viewers' behaviours. The late fusion of genre, audio and video vectors with both of the textual embeddings significantly increase the precision and diversity of the results.
Saba Nazir, Taner Cagali, Mehrnoosh Sadrzadeh, Chris Newell
ISM4
2017 Updatable, Accurate, Diverse, and Scalable Recommendations for Interactive Applications
abstract
Recommender systems form the backbone of many interactive systems. They incorporate user feedback to personalize the user experience typically via personalized recommendation lists. As users interact with a system, an increasing amount of data about a user’s preferences becomes available, which can be leveraged for improving the systems’ performance. Incorporating these new data into the underlying recommendation model is, however, not always straightforward. Many models used by recommender systems are computationally expensive and, therefore, have to perform offline computations to compile the recommendation lists. For interactive applications, it is desirable to be able to update the computed values as soon as new user interaction data is available: updating recommendations in interactive time using new feedback data leads to better accuracy and increases the attraction of the system to the users. Additionally, there is a growing consensus that accuracy alone is not enough and user satisfaction is also dependent on diverse recommendations. In this work, we tackle this problem of updating personalized recommendation lists for interactive applications in order to provide both accurate and diverse recommendations. To that end, we explore algorithms that exploit random walks as a sampling technique to obtain diverse recommendations without compromising on efficiency and accuracy. Specifically, we present a novel graph vertex ranking recommendation algorithm called RP 3 β that reranks items based on three-hop random walk transition probabilities. We show empirically that RP 3 β provides accurate recommendations with high long-tail item frequency at the top of the recommendation list. We also present approximate versions of RP 3 β and the two most accurate previously published vertex ranking algorithms based on random walk transition probabilities and show that these approximations converge with an increasing number of samples. To obtain interactively updatable recommendations, we additionally show how our algorithm can be extended for online updates at interactive speeds. The underlying random walk sampling technique makes it possible to perform the updates without having to recompute the values for the entire dataset. In an empirical evaluation with three real-world datasets, we show that RP 3 β provides highly accurate and diverse recommendations that can easily be updated with newly gathered information at interactive speeds (≪ 100 ms ).
Bibek Paudel, Fabian Christoffel, Chris Newell, Abraham Bernstein
ACM Trans. Interact. Intell. Syst.3
2015 Blockbusters and Wallflowers: Accurate, Diverse, and Scalable Recommendations with Random Walks
abstract
User satisfaction is often dependent on providing accurate and diverse recommendations. In this paper, we explore scalable algorithms that exploit random walks as a sampling technique to obtain diverse recommendations without compromising on accuracy. Specifically, we present a novel graph vertex ranking recommendation algorithm called RP^3_beta that re-ranks items based on 3-hop random walk transition probabilities. We show empirically, that RP^3_beta provides accurate recommendations with high long-tail item frequency at the top of the recommendation list. We also present scalable approximate versions of RP^3_beta and the two most accurate previously published vertex ranking algorithms based on random walk transition probabilities and show that these approximations converge with increasing number of samples.
Fabian Christoffel, Bibek Paudel, Chris Newell, Abraham Bernstein
RecSys3
2013 Design and evaluation of a client-side recommender system
abstract
Most recommender systems found on the web are server-based and centralised. However, it can be difficult to maintain the responsiveness with this approach when there are large numbers of concurrent users. In this demonstration we present an alternative approach where major parts of the recommender system are implemented in scripts run by the user's client system.
Chris Newell, Libby Miller
RecSys1
2012 Explaining the user experience of recommender systems
abstract
Research on recommender systems typically focuses on the accuracy of prediction algorithms. Because accuracy only partially constitutes the user experience of a recommender system, this paper proposes a framework that takes a user-centric approach to recommender system evaluation. The framework links objective system aspects to objective user behavior through a series of perceptual and evaluative constructs (called subjective system aspects and experience, respectively). Furthermore, it incorporates the influence of personal and situational characteristics on the user experience. This paper reviews how current literature maps to the framework and identifies several gaps in existing work. Consequently, the framework is validated with four field trials and two controlled experiments and analyzed using Structural Equation Modeling. The results of these studies show that subjective system aspects and experience variables are invaluable in explaining why and how the user experience of recommender systems comes about . In all studies we observe that perceptions of recommendation quality and/or variety are important mediators in predicting the effects of objective system aspects on the three components of user experience: process (e.g. perceived effort, difficulty), system (e.g. perceived system effectiveness) and outcome (e.g. choice satisfaction). Furthermore, we find that these subjective aspects have strong and sometimes interesting behavioral correlates (e.g. reduced browsing indicates higher system effectiveness). They also show several tradeoffs between system aspects and personal and situational characteristics (e.g. the amount of preference feedback users provide is a tradeoff between perceived system usefulness and privacy concerns). These results, as well as the validated framework itself, provide a platform for future research on the user-centric evaluation of recommender systems.
Bart P. Knijnenburg, Martijn C. Willemsen, Zeno Gantner, Hakan Soncu, Chris Newell
User Model. User Adapt. Interact.5