Sarah Marzen

dblp:143/7392 · also Sarah E. Marzen · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 67% Knowledge representation and reasoning · 33%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
belief representation
0.812024
Transformers Represent Belief State Geometry in their Residual Stream · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › explainable AI › interpretable deep learning
residual stream analysis
0.812024
Transformers Represent Belief State Geometry in their Residual Stream · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability
0.812024
Transformers Represent Belief State Geometry in their Residual Stream · NeurIPS 2024
Information theory › algorithmic information theory
optimal predictor
0.212024
Transformers Represent Belief State Geometry in their Residual Stream · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

next-token prediction · 1.5linear representation analysis · 1.5
YearPublicationVenuePosition
2024 Transformers Represent Belief State Geometry in their Residual Stream
abstract
What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the meta-dynamics of belief updating over hidden states of the data- generating process. Leveraging the theory of optimal prediction, we anticipate and then find that belief states are linearly represented in the residual stream of transformers, even in cases where the predicted belief state geometry has highly nontrivial fractal structure. We investigate cases where the belief state geometry is represented in the final residual stream or distributed across the residual streams of multiple layers, providing a framework to explain these observations. Furthermore we demonstrate that the inferred belief states contain information about the entire future, beyond the local next-token prediction that the transformers are explicitly trained on. Our work provides a general framework connecting the structure of training data to the geometric structure of activations inside transformers.
Adam S. Shai, Lucas Teixeira, Alexander Gietelink Oldenziel, Sarah Marzen, Paul M. Riechers
NeurIPS4
2017 Revisiting Perceptual Distortion for Natural Images: Mean Discrete Structural Similarity Index
abstract
A challenge in image processing is quantifying the perceptual quality of distorted images. Solutions to this problem allow lossy compression algorithms to be more easily and accurately evaluated. Motivated by failings of mean-squared error (MSE/PSNR), Wang, Bovik, and others proposed a perceptual image measure called mean structural similarity (MSSIM), which decomposes the distortion of image patches into three components: a difference in mean luminance, a difference in luminance variance, and a difference in structure. We present a new measure, mean discrete structural similarity (MDSSIM), that replaces the structural comparison of MSSIM with the Hamming distance between suitably discretized original and distorted image patches. To assess its performance, we apply this new image measure to a standard human psychophysics dataset, the LIVE Image Quality Assessment Database (Release 2). The high correlation of MDSSIM with human scores suggests, consistent with experiment and well-known results about lossy compression, that the human visual system may be fundamentally concerned with discrete structure in natural images.
Christopher Hillar, Sarah Marzen
DCC2