VLDB 2026 Research / reviewers in the wild / expert
Isaac L. Chuang
dblp:54/4237
· DBLP profile ↗
20ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-7296-523XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 since 2021Systems, architecture and hardware · 9 · 1 since 2021Theory of computation · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Formation of Representations in Neural NetworksabstractUnderstanding neural representations will help open the black box of neural networks and advance our scientific understanding of modern AI systems. However, how complex, structured, and transferable representations emerge in modern neural networks has remained a mystery. Building on previous results, we propose the Canonical Representation Hypothesis (CRH), which posits a set of six alignment relations to universally govern the formation of representations in most hidden layers of a neural network. Under the CRH, the latent representations (R), weights (W), and neuron gradients (G) become mutually aligned during training. This alignment implies that neural networks naturally learn compact representations, where neurons and weights are invariant to task-irrelevant transformations. We then show that the breaking of CRH leads to the emergence of reciprocal power-law relations between R, W, and G, which we refer to as the Polynomial Alignment Hypothesis (PAH). We present a minimal-assumption theory proving that the balance between gradient noise and regularization is crucial for the emergence of the canonical representation. The CRH and PAH lead to an exciting possibility of unifying major key deep learning phenomena, including neural collapse and the neural feature ansatz, in a single framework. Liu Ziyin 0001, Isaac L. Chuang, Tomer Galanti, Tomaso A. Poggio |
ICLR | 2 |
| 2025 | Remove Symmetries to Control Model Expressivity and Improve OptimizationabstractWhen symmetry is present in the loss function, the model is likely to be trapped in a low-capacity state that is sometimes known as a ``collapse." Being trapped in these low-capacity states can be a major obstacle to training across many scenarios where deep learning technology is applied. We first prove two concrete mechanisms through which symmetries lead to reduced capacities and ignored features during training and inference. We then propose a simple and theoretically justified algorithm, \textit{syre}, to remove almost all symmetry-induced low-capacity states in neural networks. When this type of entrapment is especially a concern, removing symmetries with the proposed method is shown to correlate well with improved optimization or performance. A remarkable merit of the proposed method is that it is model-agnostic and does not require any knowledge of the symmetry. Liu Ziyin 0001, Isaac L. Chuang |
ICLR | 3 |
| 2025 | Neural Thermodynamics: Entropic Forces in Deep and Universal Representation LearningabstractWith the rapid discovery of emergent phenomena in deep learning and large language models, understanding their cause has become an urgent need. Here, we propose a rigorous entropic-force theory for understanding the learning dynamics of neural networks trained with stochastic gradient descent (SGD) and its variants. Building on the theory of parameter symmetries and an entropic loss landscape, we show that representation learning is crucially governed by emergent entropic forces arising from stochasticity and discrete-time updates. These forces systematically break continuous parameter symmetries and preserve discrete ones, leading to a series of gradient balance phenomena that resemble the equipartition property of thermal systems. These phenomena, in turn, (a) explain the universal alignment of neural representations between AI models and lead to a proof of the Platonic Representation Hypothesis, and (b) reconcile the seemingly contradictory observations of sharpness- and flatness-seeking behavior of deep learning optimization. Our theory and experiments demonstrate that a combination of entropic forces and symmetry breaking is key to understanding emergent phenomena in deep learning. Liu Ziyin 0001, Isaac L. Chuang |
NeurIPS | 3 |
| 2024 | Reliable Computation by Large-Alphabet Formulas in the Presence of NoiseabstractWe present two new positive results for reliable computation using formulas over physical alphabets of size$q \gt 2$. First, we show that for logical alphabets of size$\ell = q$the threshold for denoising using gates subject to q-ary symmetric noise with error probability$\varepsilon $is strictly larger than that for Boolean computation, and we show that reliable computation is possible as long as signals remain distinguishable, i.e.$\epsilon \lt (q - 1) / q$, in the limit of large fan-in$k \rightarrow \infty $. We also determine the point at which generalized majority gates with bounded fan-in fail, and show in particular that reliable computation is possible for$\epsilon \lt (q - 1) / (q (q + 1))$in the case of q prime and fan-in$k = 3$. Secondly, we provide an example where$\ell \lt q$, showing that reliable Boolean computation,$\ell = 2$, can be performed using 2-input ternary,$q = 3$, logic gates subject to symmetric ternary noise of strength$\varepsilon \lt 1/6$by using the additional alphabet element for error signaling. Andrew K. Tan, Matthew H. Ho, Isaac L. Chuang |
IEEE Trans. Inf. Theory | 3 |
| 2024 | ARQUIN: Architectures for Multinode Superconducting Quantum ComputersabstractMany proposals to scale quantum technology rely on modular or distributed designs wherein individual quantum processors, called nodes, are linked together to form one large multinode quantum computer (MNQC). One scalable method to construct an MNQC is using superconducting quantum systems with optical interconnects. However, internode gates in these systems may be two to three orders of magnitude noisier and slower than local operations. Surmounting the limitations of internode gates will require improvements in entanglement generation, use of entanglement distillation, and optimized software and compilers. Still, it remains unclear what performance is possible with current hardware and what performance algorithms require. In this article, we employ a systems analysis approach to quantify overall MNQC performance in terms of hardware models of internode links, entanglement distillation, and local architecture. We show how to navigate tradeoffs in entanglement generation and distillation in the context of algorithm performance, lay out how compilers and software should balance between local and internode gates, and discuss when noisy quantum internode links have an advantage over purely classical links. We find that a factor of 10–100× better link performance is required and introduce a research roadmap for the co-design of hardware and software towards the realization of early MNQCs. While we focus on superconducting devices with optical interconnects, our approach is general across MNQC implementations. James Ang 0001, Gabriella Carini, Yanzhu Chen, Isaac L. Chuang, Michael DeMarco, Sophia E. Economou, Alec Eickbusch, Andrei Faraon, Kai-Mei Fu, Steven M. Girvin, Michael Hatridge, Andrew A. Houck, Paul Hilaire, Kevin Krsulich, Ang Li 0006, Yuan Liu 0023, Margaret Martonosi, David C. McKay, Jim Misewich, Mark B. Ritter, Robert J. Schoelkopf, Samuel A. Stein, Sara Sussman, Teague Tomesh, Norm M. Tubman, Nathan Wiebe, Yongxin Yao, Dillon Yost, Yiyu Zhou |
ACM Trans. Quantum Comput. | 4 |
| 2023 | HetArch: Heterogeneous Microarchitectures for Superconducting Quantum SystemsabstractNoisy Intermediate-Scale Quantum Computing (NISQ) has dominated headlines in recent years, with the longer-term vision of Fault-Tolerant Quantum Computation (FTQC) offering significant potential albeit at currently intractable resource costs and quantum error correction (QEC) overheads. For problems of interest, FTQC will require millions of physical qubits with long coherence times, high-fidelity gates, and compact sizes to surpass classical systems. Just as heterogeneous specialization has offered scaling benefits in classical computing, it is likewise gaining interest in FTQC. However, systematic use of heterogeneity in either hardware or software elements of FTQC systems remains a serious challenge due to the vast design space and variable physical constraints. Samuel A. Stein, Sara Sussman, Teague Tomesh, Charlie Guinn, Esin Tureci, Sophia Fuhui Lin, James Ang 0001, Srivatsan Chakram, Ang Li 0006, Margaret Martonosi, Fred Chong, Andrew A. Houck, Isaac L. Chuang, Michael DeMarco |
MICRO | 14 |
| 2021 | Confident Learning: Estimating Uncertainty in Dataset LabelsabstractLearning exists in the context of data, yet notions of confidence typically focus on model predictions, not label quality. Confident learning (CL) is an alternative approach which focuses instead on label quality by characterizing and identifying label errors in datasets, based on the principles of pruning noisy data, counting with probabilistic thresholds to estimate noise, and ranking examples to train with confidence. Whereas numerous studies have developed these principles independently, here, we combine them, building on the assumption of a class-conditional noise process to directly estimate the joint distribution between noisy (given) labels and uncorrupted (unknown) labels. This results in a generalized CL which is provably consistent and experimentally performant. We present sufficient conditions where CL exactly finds label errors, and show CL performance exceeding seven recent competitive approaches for learning with noisy labels on the CIFAR dataset. Uniquely, the CL framework is not coupled to a specific data modality or model (e.g., we use CL to find several label errors in the presumed error-free MNIST dataset and improve sentiment classification on text data in Amazon Reviews). We also employ CL on ImageNet to quantify ontological class overlap (e.g., estimating 645 missile images are mislabeled as their parent class projectile), and moderately increase model accuracy (e.g., for ResNet) by cleaning data prior to training. These results are replicable using the open-source cleanlab release. Curtis G. Northcutt, Lu Jiang 0004, Isaac L. Chuang |
J. Artif. Intell. Res. | 3 |
| 2019 | Learnability for the Information Bottleneck
Tailin Wu, Ian S. Fischer, Isaac L. Chuang, Max Tegmark |
UAI | 3 |
| 2017 | Google BigQuery for Education: Framework for Parsing and Analyzing edX MOOC DataabstractThe size and complexity of MOOC data present overwhelming challenges to many institutions. This paper details the functionality of edx2bigquery -- an open source Python package developed by Harvard and MIT to ingest and report on hundreds of MITx and HarvardX course datasets from edX, making use of Google BigQuery to handle multiple terabytes of learner data. For this application, we find that Google BigQuery provides ease of use in loading the multi-faceted MOOC datasets and near real-time interactive querying of data, including large clickstream datasets; moreover, we are able to provide flexible research and reporting dashboards, visualizing and aggregating data, by interfacing services associated with BigQuery. This framework makes it feasible for edx2bigquery to be open source, following standards which emphasize the importance of data products that transcend a particular data science platform and allow teams with diverse backgrounds to interact with data. edx2bigquery is being adopted by other institutions with an aim toward future collaboration. Glenn Lopez, Daniel T. Seaton, Andrew M. Ang, Dustin Tingley, Isaac L. Chuang |
L@S | 5 |
| 2017 | Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels
Curtis G. Northcutt, Tailin Wu, Isaac L. Chuang |
UAI | 3 |
| 2015 | Probabilistic Use Cases: Discovering Behavioral Patterns for Predicting CertificationabstractAdvances in open-online education have led to a dramatic increase in the size, diversity, and traceability of learner populations, offering tremendous opportunities to study detailed learning behavior of users around the world. This paper adapts the topic modeling approach of Latent Dirichlet Allocation (LDA) to uncover behavioral structure from student logs in a MITx Massive Open Online Course, 8.02x: Electricity and Magnetism. LDA is typically found in the field of natural language processing, where it identifies the latent topic structure within a collection of documents. However, this framework can be adapted for analysis of user-behavioral patterns by considering user interactions with courseware as a ``bag of interactions'' equivalent to the ``bag of words'' model found in topic modeling. By employing this representation, LDA forms probabilistic use cases that clusters students based on their behavior. Through the probability distributions associated with each use case, this approach provides an interpretable representation of user access patterns, while reducing the dimensionality of the data and improving accuracy. Using only the first week of logs, we can predict whether or not a student will earn a certificate with 0.81 ± 0.01 cross-validation accuracy. Thus, the method presented in this paper is a powerful tool in understanding user behavior and predicting outcomes. Cody A. Coleman, Daniel T. Seaton, Isaac L. Chuang |
L@S | 3 |
| 2014 | Due dates in MOOCs: does stricter mean better?abstractMassive Open Online Courses (MOOCs) employ a variety of components to engage students in learning (eg. videos, forums, quizzes). Some components are graded, which means that they play a key role in a student's final grade and certificate attainment. It is not yet clear how the due date structure of graded components affects student outcomes including academic performance and alternative modes of learning of students. Using data from HarvardX and MITx, Harvard's and MIT's divisions for online learning, we study the structure of due dates on graded components for 10 completed MOOCs. We find that stricter due dates are associated with higher certificate attainment rates but fewer students who join late being able to earn a certificate. Our findings motivate further studies of how the use of graded components and deadlines affects academic and alternative learning of MOOC students, and can help inform the design of online courses. Sergiy O. Nesterko, Daniel T. Seaton, Justin Reich, Joseph McIntyre, Qiuyi Han, Isaac L. Chuang, Andrew D. Ho |
L@S | 6 |
| 2011 | Transversality Versus Universality for Additive Quantum CodesabstractLogic gates can be performed on data encoded in quantum code blocks such that errors introduced by faulty gates can be corrected. The important class of transversal gates acts bitwise between corresponding qubits of code blocks and thus limits error propagation. If any quantum gate could be implemented using transversal gates, the set would be universal. We study the structure ofGF(4)-additive quantum codes and prove that no universal set of transversal logic gates exists for these codes. This result is in stark contrast with the classical case, where universal transversal gate sets exist, and strongly supports the idea that additional quantum techniques, based, for example, on quantum teleportation or magic state distillation, are necessary to achieve universal fault-tolerant quantum computation on additive codes. Bei Zeng, Andrew W. Cross, Isaac L. Chuang |
IEEE Trans. Inf. Theory | 3 |
| 2008 | High-level interconnect model for the quantum logic array architectureabstractWe summarize the main characteristics of the quantum logic array (QLA) architecture with a careful look at the key issues not described in the original conference publications: primarily, the teleportation-based logical interconnect. The design goal of the the quantum logic array architecture is to illustrate a model for a large-scale quantum architecture that solves the primary challenges of system-level reliability and data distribution over large distances. The QLA's logical interconnect design, which employs the quantum repeater protocol, is in principle capable of supporting the communication requirements for applications as large as the factoring of a 2048-bit number using Shor's quantum factoring algorithm. Our physical-level assumptions and architectural component validations are based on the trapped ion technology for implementing quantum computing. Tzvetan S. Metodi, Darshan D. Thaker, Andrew W. Cross, Isaac L. Chuang, Fred Chong |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2007 | The quantum Schur and Clebsch-Gordan transforms: I. efficient qudit circuits
Dave Bacon, Isaac L. Chuang, Aram W. Harrow |
SODA | 2 |
| 2006 | Quantum Memory Hierarchies: Efficient Designs to Match Available Parallelism in Quantum ComputingabstractThe assumption of maximum parallelism support for the successful realization of scalable quantum computers has led to homogeneous, "sea-of-qubits" architectures. The resulting architectures overcome the primary challenges of reliability and scalability at the cost of physically unacceptable system area. We find that by exploiting the natural serialization at both the application and the physical microarchitecture level of a quantum computer, we can reduce the area requirement while improving performance. In particular we present a scalable quantum architecture design that employs specialization of the system into memory and computational regions, each individually optimized to match hardware support to the available parallelism. Through careful application and system analysis, we find that our new architecture can yield up to a factor of thirteen savings in area due to specialization. In addition, by providing a memory hierarchy design for quantum computers, we can increase time performance by a factor of eight. This result brings us closer to the realization of a quantum processor that can solve meaningful problems Darshan D. Thaker, Tzvetan S. Metodi, Andrew W. Cross, Isaac L. Chuang, Fred Chong |
ISCA | 4 |
| 2004 | Datapath and control for quantum wiresabstractAs quantum computing moves closer to reality the need for basic architectural studies becomes more pressing. Quantum wires, which transport quantum data, will be a fundamental component in all anticipated silicon quantum architectures. Since they cannot consist of a stream of electrons, as in the classical case, quantum wires must fundamentally be designed differently. In this paper, we present two quantum wire designs: a swap wire, based on swapping of adjacent qubits, and a teleportation wire, based on the quantum teleportation primitive. We characterize the latency and bandwidth of these two alternatives in a device-independent way. Furthermore, unlike classical wires, quantum wires need control signals in order to operate. We explore the complexity of the control mechanisms and the fundamental tension between the scale of quantum effects and the scale of the classical logic needed to control them. This "pitch-matching" problem imposes constraints on minimum wire lengths and wire intersections, leading us to use a SIMD approach for the control mechanisms. We ultimately show that qubit decoherence imposes a basic limit on the maximum communication distance of the swapping wire, while relatively large overhead imposes a basic limit on the minimum communication distance of the teleportation wire. Nemanja Isailovic, Mark Whitney, Yatish Patel, John Kubiatowicz, Dean Copsey, Fred Chong, Isaac L. Chuang, Mark Oskin |
ACM Trans. Archit. Code Optim. | 7 |
| 2003 | Building Quantum Wires: The Long and the Short of ItabstractAs quantum computing moves closer to reality the need for basic architectural studies becomes more pressing. Quantum wires, which transport quantum data, is a fundamental component in all anticipated silicon quantum architectures. We introduce a quantum wire architecture based upon quantum teleportation. We compare this teleportation channel with the traditional approach to transporting quantum data, which we refer to as the swapping channel. We characterize the latency and bandwidth of these two alternatives in a device-independent way and describe how the advanced architecture of the teleportation channel overcomes a basic limit to the maximum communication distance of the swapping channel. In addition, we discover a fundamental tension between the scale of quantum effects and the scale of the classical logic needed to control them. This "pitch-matching" problem imposes constraints on minimum wire lengths and wire intersections, which in turn imply a sparsely connected architecture of coarse-grained quantum computational elements. This is in direct contrast to the "sea of gates " architectures presently assumed by most quantum computing studies. Mark Oskin, Fred Chong, Isaac L. Chuang, John Kubiatowicz |
ISCA | 3 |
| 2003 | The effect of communication costs in solid-state quantum computing architecturesabstractQuantum computation has become an intriguing technology with which to attack difficult problems and to enhance system security. Quantum algorithms, however, have been analyzed under idealized assumptions without important physical constraints in mind. In this paper, we analyze two key constraints: the short spatial distance of quantum interactions and the short temporal life of quantum data.In particular, quantum computations must make use of extremely robust error correction techniques to extend the life of quantum data. We present optimized spatial layouts of quantum error correction circuits for quantum bits embedded in silicon. We analyze the complexity of error correction under the constraint that interaction between these bits is near neighbor and data must be propagated via swap operations from one part of the circuit to another.We discover two interesting results from our quantum layouts. First, the recursive nature of quantum error correction circuits requires a additional communication technique more powerful than near-neighbor swaps -- too much error accumulates if we attempt to swap over long distances. We show that quantum teleportation can be used to implement recursive structures. We also show that the reliability of the quantum swap operation is the limiting factor in solid-state quantum computation. Dean Copsey, Mark Oskin, Tzvetan S. Metodi, Fred Chong, Isaac L. Chuang, John Kubiatowicz |
SPAA | 5 |
| 2000 | Reversible arithmetic coding for quantum data compressionabstractWe study the problem of compressing a block of symbols (a block quantum state) emitted by a memoryless quantum Bernoulli source. We present a simple-to-implement quantum algorithm for projecting, with high probability, the block quantum state onto the typical subspace spanned by the lending eigenstates of its density matrix. We propose a fixed-rate quantum Shannon-Fano code to compress the projected block quantum state using a per-symbol code rate that is slightly higher than the von Neumann (1955) entropy limit. Finally, we propose quantum arithmetic codes to efficiently implement quantum Shannon-Fano (1948) codes. Our arithmetic encoder and decoder have a cubic circuit and a cubic computational complexity in the block size. Both the encoder and decoder are quantum-mechanical inverses of each other, and constitute an elegant example of reversible quantum computation. Isaac L. Chuang, Dharmendra S. Modha |
IEEE Trans. Inf. Theory | 1 |