Estefany Kelly Buchanan

dblp:280/1198 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0003-1448-5662ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 31% Trustworthy machine learning · 21% Efficient and distributed learning · 14%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model inference
inference-time techniques
0.912025
An Architecture Search Framework for Inference-Time Techniques · ICML 2025
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.912025
An Architecture Search Framework for Inference-Time Techniques · ICML 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
An Architecture Search Framework for Inference-Time Techniques · ICML 2025
Machine learning › Learning paradigms
weakly supervised learning
0.912025
Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for Verification · NeurIPS 2025
Bioinformatics and computational biology › neuroscience › neuroinformatics
neural data analysis
0.912025
Extracting task-relevant preserved dynamics from contrastive aligned neural recordings · NeurIPS 2025
Bioinformatics and computational biology › computational neuroscience
neural decoding
0.912025
Extracting task-relevant preserved dynamics from contrastive aligned neural recordings · NeurIPS 2025
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.612022
Deep Ensembles Work, But Are They Necessary? · NeurIPS 2022
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.612022
Deep Ensembles Work, But Are They Necessary? · NeurIPS 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
Deep Ensembles Work, But Are They Necessary? · NeurIPS 2022
Machine learning › Trustworthy machine learning
uncertainty estimation
0.612022
Deep Ensembles Work, But Are They Necessary? · NeurIPS 2022
Computer vision › 3D vision › motion capture
animal pose tracking
0.412020
Deep Graph Pose: a semi-supervised deep graphical model for improved animal pose tracking · NeurIPS 2020
Computer vision › 3D vision
pose estimation
0.412020
Deep Graph Pose: a semi-supervised deep graphical model for improved animal pose tracking · NeurIPS 2020
Machine learning › Representation and self-supervised learning
contrastive learning
0.312025
Extracting task-relevant preserved dynamics from contrastive aligned neural recordings · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
model distillation
0.312025
Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for Verification · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.112020
Deep Graph Pose: a semi-supervised deep graphical model for improved animal pose tracking · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.7weak supervision · 0.9repeated sampling · 0.9linear dynamical systems · 0.9linear dynamical system · 0.9iterative revision · 0.9ensemble · 0.9cross-encoder distillation · 0.9architecture search · 0.9neural network scaling · 0.6deep ensembles · 0.6
YearPublicationVenuePosition
2025 An Architecture Search Framework for Inference-Time Techniques
abstract
Inference-time techniques, such as repeated sampling or iterative revisions, are emerging as powerful ways to enhance large-language models (LLMs) at test time. However, best practices for developing systems that combine these techniques remain underdeveloped due to our limited understanding of the utility of each technique across models and tasks, the interactions between them, and the massive search space for combining them. To address these challenges, we introduce Archon, a modular and automated framework for optimizing the process of selecting and combining inference-time techniques and LLMs. Given a compute budget and a set of available LLMs, Archon explores a large design space to discover optimized configurations tailored to target benchmarks. It can design custom or general-purpose architectures that advance the Pareto frontier of accuracy vs. maximum token budget compared to top-performing baselines. Across instruction-following, reasoning, and coding tasks, we show that Archon can leverage additional inference compute budget to design systems that outperform frontier models such as OpenAI’s o1, GPT-4o, and Claude 3.5 Sonnet by an average of 15.1%.
Jon Saad-Falcon, Adrian Gamarra Lafuente, Shlok Natarajan, Nahum Maru, Hristo Todorov, Etash Kumar Guha, Estefany Kelly Buchanan, Mayee F. Chen, Neel Guha, Christopher Ré, Azalia Mirhoseini
ICML7
2025 Extracting task-relevant preserved dynamics from contrastive aligned neural recordings
abstract
Recent work indicates that low-dimensional dynamics of neural and behavioral data are often preserved across days and subjects. However, extracting these preserved dynamics remains challenging: high-dimensional neural population activity and the recorded neuron populations vary across recording sessions. While existing modeling tools can improve alignment between neural and behavioral data, they often operate on a per-subject basis or discretize behavior into categories, disrupting its natural continuity and failing to capture the underlying dynamics. We introduce $\underline{\text{C}}$ontrastive $\underline{\text{A}}$ligned $\underline{\text{N}}$eural $\underline{\text{D}}$$\underline{\text{Y}}$namics (CANDY), an end‑to‑end framework that aligns neural and behavioral data using rank-based contrastive learning, adapted for continuous behavioral variables, to project neural activity from different sessions onto a shared low-dimensional embedding space. CANDY fits a shared linear dynamical system to the aligned embeddings, enabling an interpretable model of the conserved temporal structure in the latent space. We validate CANDY on synthetic and real-world datasets spanning multiple species, behaviors, and recording modalities. Our results show that CANDY is able to learn aligned latent embeddings and preserved dynamics across neural recording sessions and subjects, and it achieves improved cross-session behavior decoding performance. We further show that the latent linear dynamical system generalizes to new sessions and subjects, achieving comparable or even superior behavior decoding performance to models trained from scratch. These advances enable robust cross‑session behavioral decoding and offer a path towards identifying shared neural dynamics that underlie behavior across individuals and recording conditions. The code and two-photon imaging data of striatal neural activity that we acquired here are available at https://github.com/schnitzer-lab/CANDY-public.git.
Yiqi Jiang, Kaiwen Sheng, Estefany Kelly Buchanan, Yu Shikano, Seung Je Woo, Yixiu Zhao, Tony Hyun Kim, Fatih Dinc, Scott W. Linderman, Mark J. Schnitzer
NeurIPS4
2025 Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for Verification
abstract
Verifiers can improve language model (LM) capabilities by providing feedback or selecting the best response from a pool of generated candidates. Currently, high-quality verifiers are either unscalable (e.g., humans) or limited in utility (e.g., tools like Lean for formal proofs). While LM judges and reward models have become broadly useful as general-purpose verifiers, a significant performance gap remains between them and oracle verifiers. To help close this gap, we introduce Weaver, a framework for designing a strong verifier by combining multiple weak, imperfect verifiers. First we find that weighted ensembles of verifiers, which typically require learning from labeled data, significantly outperform unweighted combinations due to differences in the verifiers. To reduce the dependency on labeled data, Weaver leverages weak supervision to estimate each verifier’s accuracy and combines their outputs into a unified score that better reflects true response quality. However, directly applying weak supervision algorithms poses several challenges, including inconsistent verifier output formats and handling low-quality verifiers. Weaver addresses these challenges by using dataset statistics to normalize outputs and filter specific verifiers. We study the effectiveness of Weaver in repeated sampling settings, where a model generates multiple candidate responses at test time and a verifier is used to select the correct one. Our evaluations demonstrate that Weaver significantly improves the pass@1 performance across several reasoning and math tasks, achieving o3-mini level accuracy with Llama 3.3 70B Instruct (a much cheaper non-reasoning model) as the generator, and an ensemble of smaller judge and reward models as the verifiers (86.2% average). This gain mirrors the jump achieved between GPT-4o and o3-mini (69.0% vs. 86.7%), which required extensive finetuning and post-training interventions. To make Weaver more efficient, we train a compact 400M cross-encoder using Weaver's combined output scores. This distilled model retains 98.7% of Weaver's full accuracy while reducing verification compute by up to 99.97%.
Jon Saad-Falcon, Estefany Kelly Buchanan, Mayee F. Chen, Tzu-Heng Huang, Brendan McLaughlin, Tanvir Bhathal, Shang Zhu, Ben Athiwaratkun, Frederic Sala, Scott W. Linderman, Azalia Mirhoseini, Christopher Ré
NeurIPS2
2022 Deep Ensembles Work, But Are They Necessary?
abstract
Ensembling neural networks is an effective way to increase accuracy, and can often match the performance of individual larger models. This observation poses a natural question: given the choice between a deep ensemble and a single neural network with similar accuracy, is one preferable over the other? Recent work suggests that deep ensembles may offer distinct benefits beyond predictive power: namely, uncertainty quantification and robustness to dataset shift. In this work, we demonstrate limitations to these purported benefits, and show that a single (but larger) neural network can replicate these qualities. First, we show that ensemble diversity, by any metric, does not meaningfully contribute to an ensemble's ability to detect out-of-distribution (OOD) data, but is instead highly correlated with the relative improvement of a single larger model. Second, we show that the OOD performance afforded by ensembles is strongly determined by their in-distribution (InD) performance, and - in this sense - is not indicative of any "effective robustness." While deep ensembles are a practical way to achieve improvements to predictive power, uncertainty quantification, and robustness, our results show that these improvements can be replicated by a (larger) single model.
Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel, John P. Cunningham
NeurIPS2
2021 Partitioning variability in animal behavioral videos using semi-supervised variational autoencoders
abstract
Recent neuroscience studies demonstrate that a deeper understanding of brain function requires a deeper understanding of behavior. Detailed behavioral measurements are now often collected using video cameras, resulting in an increased need for computer vision algorithms that extract useful information from video data. Here we introduce a new video analysis tool that combines the output of supervised pose estimation algorithms (e.g. DeepLabCut) with unsupervised dimensionality reduction methods to produce interpretable, low-dimensional representations of behavioral videos that extract more information than pose estimates alone. We demonstrate this tool by extracting interpretable behavioral features from videos of three different head-fixed mouse preparations, as well as a freely moving mouse in an open field arena, and show how these interpretable features can facilitate downstream behavioral and neural analyses. We also show how the behavioral features produced by our model improve the precision and interpretation of these downstream analyses compared to using the outputs of either fully supervised or fully unsupervised methods alone.
Matthew R. Whiteway, Dan Biderman, Yoni Friedman, Mario Dipoppa, Estefany Kelly Buchanan, Anqi Wu, John Zhou, Niccolò Bonacchi, Nathaniel J. Miska, Jean-Paul Noel, Erica Rodriguez, Michael Schartner, Karolina Socha, Anne E. Urai, C. Daniel Salzman, John P. Cunningham, Liam Paninski
PLoS Comput. Biol.5
2020 Deep Graph Pose: a semi-supervised deep graphical model for improved animal pose tracking
abstract
Noninvasive behavioral tracking of animals is crucial for many scientific investigations. Recent transfer learning approaches for behavioral tracking have considerably advanced the state of the art. Typically these methods treat each video frame and each object to be tracked independently. In this work, we improve on these methods (particularly in the regime of few training labels) by leveraging the rich spatiotemporal structures pervasive in behavioral video --- specifically, the spatial statistics imposed by physical constraints (e.g., paw to elbow distance), and the temporal statistics imposed by smoothness from frame to frame. We propose a probabilistic graphical model built on top of deep neural networks, Deep Graph Pose (DGP), to leverage these useful spatial and temporal constraints, and develop an efficient structured variational approach to perform inference in this model. The resulting semi-supervised model exploits both labeled and unlabeled frames to achieve significantly more accurate and robust tracking while requiring users to label fewer training frames. In turn, these tracking improvements enhance performance on downstream applications, including robust unsupervised segmentation of behavioral syllables,'' and estimation of interpretabledisentangled'' low-dimensional representations of the full behavioral video. Open source code is available at \href{\CodeLink}{https://github.com/paninski-lab/deepgraphpose}.
Anqi Wu, Estefany Kelly Buchanan, Matthew R. Whiteway, Michael Schartner, Guido Meijer, Jean-Paul Noel, Erica Rodriguez, Claire Everett, Amy Norovich, Evan Schaffer, Neeli Mishra, C. Daniel Salzman, Dora E. Angelaki, Andrés Bendesky, John P. Cunningham, Liam Paninski
NeurIPS2