Paul Smolensky

dblp:48/1105 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0003-2420-182XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Mechanisms of Symbol Processing for In-Context Learning in Transformer Networks
abstract
Large Language Models (LLMs) have demonstrated impressive abilities in symbol processing through in-context learning (ICL). This success flies in the face of decades of critiques asserting that artificial neural networks cannot master abstract symbol manipulation. We seek to understand the mechanisms that can enable robust symbol processing in transformer networks, illuminating both the unanticipated success, and the significant limitations, of transformers in symbol processing. Borrowing insights from symbolic AI and cognitive science on the power of Production System architectures, we develop a high-level Production System Language, PSL, that allows us to write symbolic programs to do complex, abstract symbol processing, and create compilers that precisely implement PSL programs in transformer networks which are, by construction, 100% mechanistically interpretable. The work is driven by study of a purely abstract (semantics-free) symbolic task that we develop, Templatic Generation (TGT). Although developed through study of TGT, PSL is, we demonstrate, highly general: it is Turing Universal. The new type of transformer architecture that we compile from PSL programs suggests a number of paths for enhancing transformers’ capabilities at symbol processing. We note, however, that the work we report addresses computability, and not learnability, by transformer networks.
Paul Smolensky, Roland Fernandez, Zhenghao Herbert Zhou, Mattia Opper, Adam Davies, Jianfeng Gao 0001
J. Artif. Intell. Res.1
2024 Structural Generalization of Modification in Adult Learners of an Artificial Language
Najoung Kim, Paul Smolensky
CogSci2
2024 Toward Compositional Behavior in Neural Models: A Survey of Current Views
abstract
Compositionality is a core property of natural language, and compositional behavior (CB) is a crucial goal for modern NLP systems.The research literature, however, includes conflicting perspectives on how CB should be defined, evaluated, and achieved.We propose a conceptual framework to address these questions and survey researchers active in this area.We find consensus on several key points.Researchers broadly accept our proposed definition of CB, agree that it is not solved by current models, and doubt that scale alone will achieve the target behavior.In other areas, we find the field is split on how to move forward, identifying diverse opportunities for future research.
Kate McCurdy, Paul Soulos, Paul Smolensky, Roland Fernandez, Jianfeng Gao 0001
EMNLP3
2024 Compositional Generalization Across Distributional Shifts with Sparse Tree Operations
abstract
Neural networks continue to struggle with compositional generalization, and this issue is exacerbated by a lack of massive pre-training. One successful approach for developing neural systems which exhibit human-like compositional generalization is $\textit{hybrid}$ neurosymbolic techniques. However, these techniques run into the core issues that plague symbolic approaches to AI: scalability and flexibility. The reason for this failure is that at their core, hybrid neurosymbolic models perform symbolic computation and relegate the scalable and flexible neural computation to parameterizing a symbolic system. We investigate a $\textit{unified}$ neurosymbolic system where transformations in the network can be interpreted simultaneously as both symbolic and neural computation. We extend a unified neurosymbolic architecture called the Differentiable Tree Machine in two central ways. First, we significantly increase the model’s efficiency through the use of sparse vector representations of symbolic structures. Second, we enable its application beyond the restricted set of tree2tree problems to the more general class of seq2seq problems. The improved model retains its prior generalization capabilities and, since there is a fully neural path through the network, avoids the pitfalls of other neurosymbolic techniques that elevate symbolic computation over neural computation.
Paul Soulos, Henry Conklin, Mattia Opper, Paul Smolensky, Jianfeng Gao 0001, Roland Fernandez
NeurIPS4
2023 Differentiable Tree Operations Promote Compositional Generalization
abstract
In the context of structure-to-structure transformation tasks, learning sequences of discrete symbolic operations poses significant challenges due to their non-differentiability. To facilitate the learning of these symbolic sequences, we introduce a differentiable tree interpreter that compiles high-level symbolic tree operations into subsymbolic matrix operations on tensors. We present a novel Differentiable Tree Machine (DTM) architecture that integrates our interpreter with an external memory and an agent that learns to sequentially select tree operations to execute the target transformation in an end-to-end manner. With respect to out-of-distribution compositional generalization on synthetic semantic parsing and language generation tasks, DTM achieves 100% while existing baselines such as Transformer, Tree Transformer, LSTM, and Tree2Tree LSTM achieve less than 30%. DTM remains highly interpretable in addition to its perfect performance.
Paul Soulos, Edward J. Hu, Kate McCurdy, Yunmo Chen, Roland Fernandez, Paul Smolensky, Jianfeng Gao 0001
ICML6
2023 How Much Do Language Models Copy From Their Training Data? Evaluating Linguistic Novelty in Text Generation Using RAVEN
abstract
Abstract Current language models can generate high-quality text. Are they simply copying text they have seen before, or have they learned generalizable linguistic abstractions? To tease apart these possibilities, we introduce RAVEN, a suite of analyses for assessing the novelty of generated text, focusing on sequential structure (n-grams) and syntactic structure. We apply these analyses to four neural language models trained on English (an LSTM, a Transformer, Transformer-XL, and GPT-2). For local structure—e.g., individual dependencies—text generated with a standard sampling scheme is substantially less novel than our baseline of human-generated text from each model’s test set. For larger-scale structure—e.g., overall sentence structure—model-generated text is as novel or even more novel than the human-generated baseline, but models still sometimes copy substantially, in some cases duplicating passages over 1,000 words long from the training set. We also perform extensive manual analysis, finding evidence that GPT-2 uses both compositional and analogical generalization mechanisms and showing that GPT-2’s novel text is usually well-formed morphologically and syntactically but has reasonably frequent semantic issues (e.g., being self-contradictory).
Tom McCoy 0001, Paul Smolensky, Tal Linzen, Jianfeng Gao 0001, Asli Celikyilmaz
Trans. Assoc. Comput. Linguistics2
2021 Infinite use of finite means? Evaluating the generalization of center embedding learned from an artificial grammar
Tom McCoy 0001, Jennifer Culbertson, Paul Smolensky, Geraldine Legendre
CogSci3
2021 Compositional processing emerges in neural networks solving math problems
Jacob L. Russin, Roland Fernandez, Hamid Palangi, Eric Rosen, Nebojsa Jojic, Paul Smolensky, Jianfeng Gao 0001
CogSci6
2021 Enriching Transformers with Structured Tensor-Product Representations for Abstractive Summarization
abstract
Yichen Jiang, Asli Celikyilmaz, Paul Smolensky, Paul Soulos, Sudha Rao, Hamid Palangi, Roland Fernandez, Caitlin Smith, Mohit Bansal, Jianfeng Gao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Asli Celikyilmaz, Paul Smolensky, Paul Soulos, Sudha Rao, Hamid Palangi, Roland Fernandez, Caitlin Smith, Mohit Bansal, Jianfeng Gao 0001
NAACL-HLT3
2020 Universal linguistic inductive biases via meta-learning
Tom McCoy 0001, Erin Grant, Paul Smolensky, Thomas L. Griffiths 0001, Tal Linzen
CogSci3
2020 Invertible Tree Embeddings using a Cryptographic Role Embedding Scheme
abstract
We present a novel method for embedding trees in a vector space based on Tensor-Product Representations (TPRs) which allows for inversion: the retrieval of the original tree structure and nodes from the vectorial embedding.Unlike previous attempts, this does not come at the cost of intractable representation size; we utilize a method for non-exact inversion, showing that it works well when there is sufficient randomness in the representation scheme for simple data and providing an upper bound on its error.To handle the huge number of possible tree positions without memoizing position representation vectors, we present a method (Cryptographic Role Embedding) using cryptographic hashing algorithms that allows for the representation of unboundedly many positions.Through experiments on parse tree data, we show a 30,000-dimensional Cryptographic Role Embedding of trees can provide invertibility with error < 1% that previous methods would require 8.6 × 10 57 dimensions to represent.
Coleman Haley, Paul Smolensky
COLING2
2020 Mapping natural-language problems to formal-language solutions using structured neural representations
abstract
Generating formal-language programs represented by relational tuples, such as Lisp programs or mathematical operations, to solve problems stated in natural language is a challenging task because it requires explicitly capturing discrete symbolic structural information implicit in the input. However, most general neural sequence models do not explicitly capture such structural information, limiting their performance on these tasks. In this paper, we propose a new encoder-decoder model based on a structured neural representation, Tensor Product Representations (TPRs), for mapping Natural-language problems to Formal-language solutions, called TP-N2F. The encoder of TP-N2F employs TPR ‘binding’ to encode natural-language symbolic structure in vector space and the decoder uses TPR ‘unbinding’ to generate, in symbolic space, a sequential program represented by relational tuples, each consisting of a relation (or operation) and a number of arguments. TP-N2F considerably outperforms LSTM-based seq2seq models on two benchmarks and creates new state-of-the-art results. Ablation studies show that improvements can be attributed to the use of structured TPRs explicitly in both the encoder and decoder. Analysis of the learned structures shows how TPRs enhance the interpretability of TP-N2F.
Kezhen Chen, Qiuyuan Huang, Hamid Palangi, Paul Smolensky, Kenneth D. Forbus, Jianfeng Gao 0001
ICML4
2019 Predicting the Argumenthood of English Prepositional Phrases
Najoung Kim, Kyle Rawlins, Benjamin Van Durme, Paul Smolensky
AAAI4
2019 RNNs implicitly implement tensor-product representations
Tom McCoy 0001, Tal Linzen, Ewan Dunbar, Paul Smolensky
ICLR (Poster)4
2018 Question-Answering with Grammatically-Interpretable Representations
abstract
We introduce an architecture, the Tensor Product RecurrentNetwork (TPRN). In our application of TPRN, internal representations—learned by end-to-end optimization in a deep neural network performing a textual question-answering(QA) task—can be interpreted using basic concepts from linguistic theory. No performance penalty need be paid for this increased interpretability: the proposed model performs comparably to a state-of-the-art system on the SQuAD QA task.The internal representation which is interpreted is a Tensor Product Representation: for each input word, the model selects a symbol to encode the word, and a role in which to place the symbol, and binds the two together. The selection is via soft attention. The overall interpretation is built from interpretations of the symbols, as recruited by the trained model, and interpretations of the roles as used by the model. We find support for our initial hypothesis that symbols can be interpreted as lexical-semantic word meanings, while roles can be interpreted as approximations of grammatical roles (or categories)such as subject, wh-word, determiner, etc. Fine-grained analysis reveals specific correspondences between the learned roles and parts of speech as assigned by a standard tagger(Toutanova et al. 2003), and finds several discrepancies in the model’s favor. In this sense, the model learns significant aspectsof grammar, after having been exposed solely to linguistically unannotated text, questions, and answers: no prior linguistic knowledge is given to the model. What is given is the means to build representations using symbols and roles, with an inductive bias favoring use of these in an approximately discrete manner.
Hamid Palangi, Paul Smolensky, Xiaodong He 0001, Li Deng 0001
AAAI2
2018 Tensor Product Generation Networks for Deep NLP Modeling
abstract
Qiuyuan Huang, Paul Smolensky, Xiaodong He, Li Deng, Dapeng Wu. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Qiuyuan Huang, Paul Smolensky, Xiaodong He 0001, Li Deng 0001, Dapeng Oliver Wu
NAACL-HLT2
2016 Bifurcation analysis of a Gradient Symbolic Computation model of incremental processing
Pyeong Whan Cho, Paul Smolensky
CogSci2
2012 Subsymbolic Computation Theory for the Human Intuitive Processor
Paul Smolensky
CiE1
1994 Optimality Theory: Universal Grammar, Learning and Parsing Algorithms, and Connectionist Foundations (Abstract)
abstract
No abstract available.
Paul Smolensky, Bruce Tesar
ACL1
1993 Dynamic Conflict Resolution in a Connectionist Rule-Based System
Clayton McMillan, Michael C. Mozer, Paul Smolensky
IJCAI3
1992 Harmonic Grammars for Formal Languages
Paul Smolensky
NIPS1
1991 Rule Induction through Integrated Symbolic and Subsymbolic Processing
Clayton McMillan, Michael C. Mozer, Paul Smolensky
NIPS3
1990 Distributed Recursive Structure Processing
Geraldine Legendre, Yoshiro Miyata, Paul Smolensky
NIPS3
1990 Tensor Product Variable Binding and the Representation of Symbolic Structures in Connectionist Systems
Paul Smolensky
Artif. Intell.1
1988 Skeletonization: A Technique for Trimming the Fat from a Network via Relevance Assessment
Michael C. Mozer, Paul Smolensky
NIPS2
1987 Social science and system design: interdisciplinary collaborations
abstract
Contributions from the behavioral sciences to the design of computer systems have come primarily from psychology, and have focused on individual cognition. In this symposium, we consider the applicability to system design of approaches that focus on social interaction. The participants comprise pairs of researchers engaged in projects that aim to bring together systematic studies of naturally occurring human activities with the design of computer-based technology. Each of the projects emphasizes the importance of the social organization of communities, everyday communication and practice.
Lucy A. Suchman, William O. Beeman, Michael R. Pear, Barbara A. Fox, Paul Smolensky
CHI5
1987 Analysis of Distributed Representation of Constituent Structure in Connectionist Systems
Paul Smolensky
NIPS1
1983 Schema Selection and Stochastic Inference in Modular Environments
Paul Smolensky
AAAI1
1983 A proposal for user centered system documentation
abstract
This paper outlines a set of proposals for the development of system documentation based on an analysis of user needs. It is suggested that existing documentation is not sensitive enough to the variety of levels of user expertise, nor to the variety of contexts in which on-line help is required. We outline three specific proposals for fulfilling these needs: a quick reference facility, a command-line database, and a facility for full explanation and instruction, and suggest a number of ways in which users might access these facilities. Finally, we suggest a way of combining these facilities into an integrated structured manual, offering more effective user support than is currently provided.
Claire O'Malley, Paul Smolensky, Liam J. Bannon, E. Conway, J. Graham, J. Sokolov, Melissa Lee Monty
CHI2