Owen Dugan

dblp:330/4698 · also Owen M. Dugan · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 48% Language models and text generation · 24% Deep learning architectures and training · 14%
Theoretical computer science
1 paper
Algorithms and data structures · 50% Mathematical optimization · 50%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
least squares
0.912025
Towards Learning High-Precision Least Squares Algorithms with Sequence Models · ICLR 2025
Algorithms and data structures
numerical algorithms
0.912025
Towards Learning High-Precision Least Squares Algorithms with Sequence Models · ICLR 2025
Machine learning › Efficient and distributed learning
inference efficiency
0.812024
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model fine-tuning
0.812024
QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation · NeurIPS 2024
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.812024
QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation · NeurIPS 2024
Natural language and speech › Language models and text generation › mathematical reasoning
numerical reasoning
0.812024
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation · NeurIPS 2024
Machine learning › Generative modeling
normalizing flow
0.712023
Q-Flow: Generative Modeling for Differential Equations of Open Quantum Dynamics with Normalizing Flows · ICML 2023
Computational science and engineering › computational physics
physics simulation
0.712023
Q-Flow: Generative Modeling for Differential Equations of Open Quantum Dynamics with Normalizing Flows · ICML 2023
Emerging computing paradigms
quantum computing
0.712023
Q-Flow: Generative Modeling for Differential Equations of Open Quantum Dynamics with Normalizing Flows · ICML 2023
Machine learning › Trustworthy machine learning › language model interpretability
mechanistic interpretability of language models
0.212024
OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

euler method · 2.0transformer · 1.7polynomial architectures · 1.7linear attention · 1.7high-precision training · 1.7gated convolution · 1.7time-dependent variational principle · 1.3normalizing flow · 1.3tensor decomposition · 0.8symbolic architecture · 0.8quantum-inspired tensor adaptation · 0.8hidden state control · 0.8physics-informed neural networks · 0.7physics-informed neural network · 0.7
YearPublicationVenuePosition
2025 Towards Learning High-Precision Least Squares Algorithms with Sequence Models
abstract
This paper investigates whether sequence models can learn to perform numerical algorithms, e.g. gradient descent, on the fundamental problem of least squares. Our goal is to inherit two properties of standard algorithms from numerical analysis: (1) machine precision, i.e. we want to obtain solutions that are accurate to near floating point error, and (2) numerical generality, i.e. we want them to apply broadly across problem instances. We find that prior approaches using Transformers fail to meet these criteria, and identify limitations present in existing architectures and training procedures. First, we show that softmax Transformers struggle to perform high-precision multiplications, which prevents them from precisely learning numerical algorithms. Second, we identify an alternate class of architectures, comprised entirely of polynomials, that can efficiently represent high-precision gradient descent iterates. Finally, we investigate precision bottlenecks during training and address them via a high-precision training recipe that reduces stochastic gradient noise. Our recipe enables us to train two polynomial architectures, gated convolutions and linear attention, to perform gradient descent iterates on least squares problems. For the first time, we demonstrate the ability to train to near machine precision. Applied iteratively, our models obtain $100,000\times$ lower MSE than standard Transformers trained end-to-end and they incur a $10,000\times$ smaller generalization gap on out-of-distribution problems. We make progress towards end-to-end learning of numerical algorithms for least squares.
Jerry W. Liu, Jessica Grogan, Owen Dugan, Ashish Rao, Simran Arora, Atri Rudra, Christopher Ré
ICLR3
2024 QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor Adaptation
abstract
We propose **Quan**tum-informed **T**ensor **A**daptation (**QuanTA**), a novel, easy-to-implement, fine-tuning method with no inference overhead for large-scale pre-trained language models. By leveraging quantum-inspired methods derived from quantum circuit structures, QuanTA enables efficient *high-rank* fine-tuning, surpassing the limitations of Low-Rank Adaptation (LoRA)---low-rank approximation may fail for complicated downstream tasks. Our approach is theoretically supported by the universality theorem and the rank representation theorem to achieve efficient high-rank adaptations. Experiments demonstrate that QuanTA significantly enhances commonsense reasoning, arithmetic reasoning, and scalability compared to traditional methods. Furthermore, QuanTA shows superior performance with fewer trainable parameters compared to other approaches and can be designed to integrate with existing fine-tuning algorithms for further improvement, providing a scalable and efficient solution for fine-tuning large language models and advancing state-of-the-art in natural language processing.
Zhuo Chen 0061, Rumen Dangovski, Charlotte Loh, Owen Dugan, Marin Soljacic
NeurIPS4
2024 OccamLLM: Fast and Exact Language Model Arithmetic in a Single Step
abstract
Despite significant advancements in text generation and reasoning, Large Language Models (LLMs) still face challenges in accurately performing complex arithmetic operations. Language model systems often enable LLMs to generate code for arithmetic operations to achieve accurate calculations. However, this approach compromises speed and security, and fine-tuning risks the language model losing prior capabilities. We propose a framework that enables exact arithmetic in *a single autoregressive step*, providing faster, more secure, and more interpretable LLM systems with arithmetic capabilities. We use the hidden states of a LLM to control a symbolic architecture that performs arithmetic. Our implementation using Llama 3 with OccamNet as a symbolic model (OccamLlama) achieves 100\% accuracy on single arithmetic operations ($+,-,\times,\div,\sin{},\cos{},\log{},\exp{},\sqrt{}$), outperforming GPT 4o with and without a code interpreter. Furthermore, OccamLlama outperforms GPT 4o with and without a code interpreter on average across a range of mathematical problem solving benchmarks, demonstrating that OccamLLMs can excel in arithmetic tasks, even surpassing much larger models. Code is available at https://github.com/druidowm/OccamLLM.
Owen Dugan, Donato Jiménez-Benetó, Charlotte Loh, Zhuo Chen 0061, Rumen Dangovski, Marin Soljacic
NeurIPS1
2023 Q-Flow: Generative Modeling for Differential Equations of Open Quantum Dynamics with Normalizing Flows
abstract
Studying the dynamics of open quantum systems can enable breakthroughs both in fundamental physics and applications to quantum engineering and quantum computation. Since the density matrix $\rho$, which is the fundamental description for the dynamics of such systems, is high-dimensional, customized deep generative neural networks have been instrumental in modeling $\rho$. However, the complex-valued nature and normalization constraints of $\rho$, as well as its complicated dynamics, prohibit a seamless connection between open quantum systems and the recent advances in deep generative modeling. Here we lift that limitation by utilizing a reformulation of open quantum system dynamics to a partial differential equation (PDE) for a corresponding probability distribution $Q$, the Husimi Q function. Thus, we model the Q function seamlessly with off-the-shelf deep generative models such as normalizing flows. Additionally, we develop novel methods for learning normalizing flow evolution governed by high-dimensional PDEs based on the Euler method and the application of the time-dependent variational principle. We name the resulting approach Q-Flow and demonstrate the scalability and efficiency of Q-Flow on open quantum system simulations, including the dissipative harmonic oscillator and the dissipative bosonic model. Q-Flow is superior to conventional PDE solvers and state-of-the-art physics-informed neural network solvers, especially in high-dimensional systems.
Owen Dugan, Peter Y. Lu, Rumen Dangovski, Marin Soljacic
ICML1