Charlotte Loh

dblp:217/6481 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 30% Representation and self-supervised learning · 24% Language models and text generation · 14%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
1.322023
Towards robust and generalizable representations of extracellular data using contrastive learning · NeurIPS 2023
Multi-Symmetry Ensembles: Improving Diversity and Generalization via Opposing Symmetries · ICML 2023
Machine learning › Efficient and distributed learning
inference efficiency
0.812024
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model fine-tuning
0.812024
QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation · NeurIPS 2024
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.812024
QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation · NeurIPS 2024
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning
0.812024
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation · NeurIPS 2024
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.712023
Multi-Symmetry Ensembles: Improving Diversity and Generalization via Opposing Symmetries · ICML 2023
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.712023
Multi-Symmetry Ensembles: Improving Diversity and Generalization via Opposing Symmetries · ICML 2023
Machine learning › Learning theory › generalization
generalization analysis
0.712023
Analyzing Generalization of Neural Networks through Loss Path Kernels · NeurIPS 2023
Machine learning › Learning theory
generalization bounds
0.712023
Analyzing Generalization of Neural Networks through Loss Path Kernels · NeurIPS 2023
Machine learning › Optimization for machine learning
gradient flow
0.712023
Analyzing Generalization of Neural Networks through Loss Path Kernels · NeurIPS 2023
Machine learning › Representation and self-supervised learning › symmetry learning
symmetry-aware representation learning
0.712023
Multi-Symmetry Ensembles: Improving Diversity and Generalization via Opposing Symmetries · ICML 2023
Bioinformatics and computational biology › neuroscience › neuroinformatics
neural data analysis
0.712023
Towards robust and generalizable representations of extracellular data using contrastive learning · NeurIPS 2023
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
spike sorting
0.712023
Towards robust and generalizable representations of extracellular data using contrastive learning · NeurIPS 2023
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning
0.612022
Equivariant Self-Supervised Learning: Encouraging Equivariance in Representations · ICLR 2022
Machine learning › Trustworthy machine learning › language model interpretability
mechanistic interpretability of language models
0.212024
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step · NeurIPS 2024
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.212023
Analyzing Generalization of Neural Networks through Loss Path Kernels · NeurIPS 2023
Bioinformatics and computational biology › single-cell analysis › cell type annotation
cell type classification
0.212023
Towards robust and generalizable representations of extracellular data using contrastive learning · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

contrastive learning · 2.6data augmentation · 1.3tensor decomposition · 0.8symbolic architecture · 0.8quantum-inspired tensor adaptation · 0.8hidden state control · 0.8kernel methods · 0.7gradient flow analysis · 0.7ensemble learning · 0.7equivariance · 0.6
YearPublicationVenuePosition
2024 QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation
abstract
We propose **Quan**tum-informed **T**ensor **A**daptation (**QuanTA**), a novel, easy-to-implement, fine-tuning method with no inference overhead for large-scale pre-trained language models. By leveraging quantum-inspired methods derived from quantum circuit structures, QuanTA enables efficient *high-rank* fine-tuning, surpassing the limitations of Low-Rank Adaptation (LoRA)---low-rank approximation may fail for complicated downstream tasks. Our approach is theoretically supported by the universality theorem and the rank representation theorem to achieve efficient high-rank adaptations. Experiments demonstrate that QuanTA significantly enhances commonsense reasoning, arithmetic reasoning, and scalability compared to traditional methods. Furthermore, QuanTA shows superior performance with fewer trainable parameters compared to other approaches and can be designed to integrate with existing fine-tuning algorithms for further improvement, providing a scalable and efficient solution for fine-tuning large language models and advancing state-of-the-art in natural language processing.
Zhuo Chen 0061, Rumen Dangovski, Charlotte Loh, Owen Dugan, Marin Soljacic
NeurIPS3
2024 OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step
abstract
Despite significant advancements in text generation and reasoning, Large Language Models (LLMs) still face challenges in accurately performing complex arithmetic operations. Language model systems often enable LLMs to generate code for arithmetic operations to achieve accurate calculations. However, this approach compromises speed and security, and fine-tuning risks the language model losing prior capabilities. We propose a framework that enables exact arithmetic in *a single autoregressive step*, providing faster, more secure, and more interpretable LLM systems with arithmetic capabilities. We use the hidden states of a LLM to control a symbolic architecture that performs arithmetic. Our implementation using Llama 3 with OccamNet as a symbolic model (OccamLlama) achieves 100\% accuracy on single arithmetic operations ($+,-,\times,\div,\sin{},\cos{},\log{},\exp{},\sqrt{}$), outperforming GPT 4o with and without a code interpreter. Furthermore, OccamLlama outperforms GPT 4o with and without a code interpreter on average across a range of mathematical problem solving benchmarks, demonstrating that OccamLLMs can excel in arithmetic tasks, even surpassing much larger models. Code is available at https://github.com/druidowm/OccamLLM.
Owen Dugan, Donato Jiménez-Benetó, Charlotte Loh, Zhuo Chen 0061, Rumen Dangovski, Marin Soljacic
NeurIPS3
2023 Multi-Symmetry Ensembles: Improving Diversity and Generalization via Opposing Symmetries
abstract
Deep ensembles (DE) have been successful in improving model performance by learning diverse members via the stochasticity of random initialization. While recent works have attempted to promote further diversity in DE via hyperparameters or regularizing loss functions, these methods primarily still rely on a stochastic approach to explore the hypothesis space. In this work, we present Multi-Symmetry Ensembles (MSE), a framework for constructing diverse ensembles by capturing the multiplicity of hypotheses along symmetry axes, which explore the hypothesis space beyond stochastic perturbations of model weights and hyperparameters. We leverage recent advances in contrastive representation learning to create models that separately capture opposing hypotheses of invariant and equivariant functional classes and present a simple ensembling approach to efficiently combine appropriate hypotheses for a given task. We show that MSE effectively captures the multiplicity of conflicting hypotheses that is often required in large, diverse datasets like ImageNet. As a result of their inherent diversity, MSE improves classification performance, uncertainty quantification, and generalization across a series of transfer tasks. Our code is available at https://github.com/clott3/multi-sym-ensem
Charlotte Loh, Seungwook Han, Shivchander Sudalairaj, Rumen Dangovski, Kai Xu 0016, Florian Wenzel, Marin Soljacic, Akash Srivastava
ICML1
2023 Analyzing Generalization of Neural Networks through Loss Path Kernels
abstract
Deep neural networks have been increasingly used in real-world applications, making it critical to ensure their ability to adapt to new, unseen data. In this paper, we study the generalization capability of neural networks trained with (stochastic) gradient flow. We establish a new connection between the loss dynamics of gradient flow and general kernel machines by proposing a new kernel, called loss path kernel. This kernel measures the similarity between two data points by evaluating the agreement between loss gradients along the path determined by the gradient flow. Based on this connection, we derive a new generalization upper bound that applies to general neural network architectures. This new bound is tight and strongly correlated with the true generalization error. We apply our results to guide the design of neural architecture search (NAS) and demonstrate favorable performance compared with state-of-the-art NAS algorithms through numerical experiments.
Yilan Chen 0002, Wei Huang 0034, Charlotte Loh, Akash Srivastava, Lam M. Nguyen, Lily Weng
NeurIPS4
2023 Towards robust and generalizable representations of extracellular data using contrastive learning
abstract
Contrastive learning is quickly becoming an essential tool in neuroscience for extracting robust and meaningful representations of neural activity. Despite numerous applications to neuronal population data, there has been little exploration of how these methods can be adapted to key primary data analysis tasks such as spike sorting or cell-type classification. In this work, we propose a novel contrastive learning framework, CEED (Contrastive Embeddings for Extracellular Data), for high-density extracellular recordings. We demonstrate that through careful design of the network architecture and data augmentations, it is possible to generically extract representations that far outperform current specialized approaches. We validate our method across multiple high-density extracellular recordings. All code used to run CEED can be found at https://github.com/ankitvishnu23/CEED.
Ankit Vishnubhotla, Charlotte Loh, Akash Srivastava, Liam Paninski, Cole L. Hurwitz
NeurIPS2
2022 Equivariant Self-Supervised Learning: Encouraging Equivariance in Representations
Rumen Dangovski, Li Jing 0001, Charlotte Loh, Seungwook Han, Akash Srivastava, Brian Cheung, Pulkit Agrawal 0001, Marin Soljacic
ICLR3