VLDB 2026 Research / reviewers in the wild / expert
Achim Schilling
dblp:230/3783
· DBLP profile ↗
10ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0001-9543-305XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Refusal Behavior in Large Language Models: A Nonlinear PerspectiveabstractRefusal behavior in large language models (LLMs) enables them to decline responding to harmful, unethical, or inappropriate prompts, ensuring alignment with ethical standards. This paper investigates refusal behavior across six LLMs from three architectural families. We challenge the assumption of refusal as a linear phenomenon by employing dimensionality reduction techniques, including PCA, t-SNE, and UMAP. Our results reveal that refusal mechanisms exhibit nonlinear, multidimensional characteristics that vary by model architecture and layer. These findings highlight the need for nonlinear interpretability to improve alignment research and inform safer AI deployment strategies. Fabian Hildebrandt, Andreas K. Maier, Patrick Krauss, Achim Schilling |
IJCNN | 4 |
| 2025 | Author-Specific Linguistic Patterns Unveiled: A Deep Learning Study on Word Class DistributionsabstractDeep learning methods have been increasingly applied to computational linguistics to uncover patterns in text data. This study investigates author-specific word class distributions using part-of-speech (POS) tagging and bigram analysis. By leveraging deep neural networks, we classify literary authors based on POS tag vectors and bigram frequency matrices derived from their works. We employ fully connected and convolutional neural network architectures to explore the efficacy of unigram and bigram-based representations. Our results demonstrate that while unigram features achieve moderate classification accuracy, bigram-based models significantly improve performance, suggesting that sequential word class patterns are more distinctive of authorial style. Multi-dimensional scaling (MDS) visualizations reveal meaningful clustering of authors’ works, supporting the hypothesis that stylistic nuances can be captured through computational methods. These findings highlight the potential of deep learning and linguistic feature analysis for author profiling and literary studies. Patrick Krauss, Achim Schilling |
IJCNN | 2 |
| 2025 | Multi-modal cognitive maps for language and vision based on neural successor representationsabstractCognitive maps are a proposed concept on how the brain efficiently organizes memories and retrieves context out of them. The entorhinal-hippocampal complex is heavily involved in episodic and relational memory processing, as well as spatial navigation and is thought to built cognitive maps via place and grid cells. To make use of the promising properties of cognitive maps, we set up a multi-modal neural network using successor representations which are able to model place cell dynamics and cognitive map representations. Here, we use multi-modal inputs consisting of images and word embeddings. The network learns the similarities between novel inputs and the training database and therefore the representation of the cognitive map successfully. Subsequently, the prediction of the network can be used to infer from one modality to another with over 90% accuracy. The proposed method could therefore be a building block to improve current AI systems for better understanding of the environment and the different modalities in which objects appear. The association of specific modalities with certain encounters can therefore lead to context awareness in novel situations when similar encounters with less information occur and additional information can be inferred from the learned cognitive map. Cognitive maps, as represented by the entorhinal-hippocampal complex in the brain, organize and retrieve context from memories, suggesting that large language models (LLMs) like ChatGPT could harness similar architectures to function as a high-level processing center, akin to how the hippocampus operates within the cortex hierarchy. Finally, by utilizing multi-modal inputs, LLMs can potentially bridge the gap between different forms of data (like images and words), paving the way for context-awareness and grounding of abstract concepts through learned associations, addressing the symbol grounding problem in AI. Paul Stoewer, Achim Schilling, Pegah Ramezani, Hassane Kissane, Andreas K. Maier, Patrick Krauss |
Neurocomputing | 2 |
| 2025 | Nonlinear Neural Dynamics and Classification Accuracy in Reservoir ComputingabstractReservoir computing information processing based on untrained recurrent neural networks with random connections is expected to depend on the nonlinear properties of the neurons and the resulting oscillatory, chaotic, or fixed-point dynamics of the network. However, the degree of nonlinearity required and the range of suitable dynamical regimes for a given task remain poorly understood. To clarify these issues, we study the classification accuracy of a reservoir computer in artificial tasks of varying complexity while tuning both the neuron's degree of nonlinearity and the reservoir's dynamical regime. We find that even with activation functions of extremely reduced nonlinearity, weak recurrent interactions, and small input signals, the reservoir can compute useful representations. These representations, detectable only in higher-order principal components, make complex classification tasks linearly separable for the readout layer. Increasing the recurrent coupling leads to spontaneous dynamical behavior. Nevertheless, some input-related computations can "ride on top" of oscillatory or fixed-point attractors with little loss of accuracy, whereas chaotic dynamics often reduces task performance. By tuning the system through the full range of dynamical phases, we observe in several classification tasks that accuracy peaks at both the oscillatory/chaotic and chaotic/fixed-point phase boundaries, supporting the edge of chaos hypothesis. We also present a regression task with the opposite behavior. Our findings, particularly the robust weakly nonlinear operating regime, may offer new perspectives for both technical and biological neural networks with random connectivity. Claus Metzner, Achim Schilling, Andreas K. Maier, Patrick Krauss |
Neural Comput. | 2 |
| 2024 | Quantifying and Maximizing the Information Flux in Recurrent Neural NetworksabstractFree-running recurrent neural networks (RNNs), especially probabilistic models, generate an ongoing information flux that can be quantified with the mutual information I[x→(t),x→(t+1)] between subsequent system states x→. Although previous studies have shown that I depends on the statistics of the network's connection weights, it is unclear how to maximize I systematically and how to quantify the flux in large systems where computing the mutual information becomes intractable. Here, we address these questions using Boltzmann machines as model systems. We find that in networks with moderately strong connections, the mutual information I is approximately a monotonic transformation of the root-mean-square averaged Pearson correlations between neuron pairs, a quantity that can be efficiently computed even in large systems. Furthermore, evolutionary maximization of I[x→(t),x→(t+1)] reveals a general design principle for the weight matrices enabling the systematic construction of systems with a high spontaneous information flux. Finally, we simultaneously maximize information flux and the mean period length of cyclic attractors in the state-space of these dynamical networks. Our results are potentially useful for the construction of RNNs that serve as short-time memories or pattern generators. Claus Metzner, Marius E. Yamakou, Dennis Völkl, Achim Schilling, Patrick Krauss |
Neural Comput. | 4 |
| 2023 | Word class representations spontaneously emerge in a deep neural network trained on next word predictionabstractHow do humans learn language, and can the first language be learned at all? These fundamental questions are still hotly debated. In contemporary linguistics, there are two major schools of thought that give completely opposite answers. According to Chomsky's theory of universal grammar, language cannot be learned because children are not exposed to sufficient data in their linguistic environment. In contrast, usage-based models of language assume a profound relationship between language structure and language use. In particular, contextual mental processing and mental representations are assumed to have the cognitive capacity to capture the complexity of actual language use at all levels. The prime example is syntax, i.e., the rules by which words are assembled into larger units such as sentences. Typically, syntactic rules are expressed as sequences of word classes. However, it remains unclear whether word classes are innate, as implied by universal grammar, or whether they emerge during language acquisition, as suggested by usage-based approaches. Here, we address this issue from a machine learning and natural language processing perspective. In particular, we trained an artificial deep neural network on predicting the next word, provided sequences of consecutive words as input. Subsequently, we analyzed the emerging activation patterns in the hidden layers of the neural network. Strikingly, we find that the internal representations of nine-word input sequences cluster according to the word class of the tenth word to be predicted as output, even though the neural network did not receive any explicit information about syntactic rules or word classes during training. This surprising result suggests, that also in the human brain, abstract representational categories such as word classes may naturally emerge as a consequence of predictive coding and processing during language acquisition. Kishore Surendra, Achim Schilling, Paul Stoewer, Andreas K. Maier, Patrick Krauss |
ICMLA | 2 |
| 2023 | Leaky-Integrate-and-Fire Neuron-Like Long-Short-Term-Memory Units as Model System in Computational BiologyabstractBiological neural networks encode information very efficiently, and dynamically react to sensory input on very small time scales. In contrast to most contemporary machine learning approaches which rely on rate neurons with continuous output, biological neural networks are based on spiking neurons with quasi-binary discrete output. In Artificial Intelligence (AI) Research, time series data are efficiently encoded in Long-Short- Term-Memory (LSTM) networks. Despite their strength in encoding time series data and making predictions, LSTM units are assumed to be biologically implausible. Nevertheless, recent studies show that LSTM unit networks indeed behave similar to biological neural networks. In this study, we show that a particular choice of parameters for the weights and gates in peephole LSTM units causes these units to show similar dynamic behaviour as biologically plausible leaky integrate-and-fire (LIF) neurons, which represent a simple biologically inspired spiking neuron model. We analyzed the spiking characteristics of the restricted peephole LSTM units and characterize the parameter space, in which these units show certain spiking characteristics. We conclude that tackling complex cognitive tasks with biologically plausible and explainable artificial neural networks is an important step to make progress in both fields, neuroscience and artificial intelligence. Richard Gerum, André Erpenbeck, Patrick Krauss, Achim Schilling |
IJCNN | 4 |
| 2021 | Integration of Leaky-Integrate-and-Fire Neurons in Standard Machine Learning Architectures to Generate Hybrid Networks: A Surrogate Gradient ApproachabstractUp to now, modern machine learning (ML) has been based on approximating big data sets with high-dimensional functions, taking advantage of huge computational resources. We show that biologically inspired neuron models such as the leaky-integrate-and-fire (LIF) neuron provide novel and efficient ways of information processing. They can be integrated in machine learning models and are a potential target to improve ML performance. Thus, we have derived simple update rules for LIF units to numerically integrate the differential equations. We apply a surrogate gradient approach to train the LIF units via backpropagation. We demonstrate that tuning the leak term of the LIF neurons can be used to run the neurons in different operating modes, such as simple signal integrators or coincidence detectors. Furthermore, we show that the constant surrogate gradient, in combination with tuning the leak term of the LIF units, can be used to achieve the learning dynamics of more complex surrogate gradients. To prove the validity of our method, we applied it to established image data sets (the Oxford 102 flower data set, MNIST), implemented various network architectures, used several input data encodings and demonstrated that the method is suitable to achieve state-of-the-art classification performance. We provide our method as well as further surrogate gradient methods to train spiking neural networks via backpropagation as an open-source KERAS package to make it available to the neuroscience and machine learning community. To increase the interpretability of the underlying effects and thus make a small step toward opening the black box of machine learning, we provide interactive illustrations, with the possibility of systematically monitoring the effects of parameter changes on the learning characteristics. Richard Gerum, Achim Schilling |
Neural Comput. | 2 |
| 2021 | Quantifying the separability of data classes in neural networksabstractWe introduce the Generalized Discrimination Value (GDV) that measures, in a non-invasive manner, how well different data classes separate in each given layer of an artificial neural network. It turns out that, at the end of the training period, the GDV in each given layer L attains a highly reproducible value, irrespective of the initialization of the network's connection weights. In the case of multi-layer perceptrons trained with error backpropagation, we find that classification of highly complex data sets requires a temporal reduction of class separability, marked by a characteristic 'energy barrier' in the initial part of the GDV(L) curve. Even more surprisingly, for a given data set, the GDV(L) is running through a fixed 'master curve', independently from the total number of network layers. Finally, due to its invariance with respect to dimensionality, the GDV may serve as a useful tool to compare the internal representational dynamics of artificial neural networks with different architectures for neural architecture search or network compression; or even with brain activity in order to decide between different candidate models of brain function. Achim Schilling, Andreas K. Maier, Richard Gerum, Claus Metzner, Patrick Krauss |
Neural Networks | 1 |
| 2020 | Sparsity through evolutionary pruning prevents neuronal networks from overfittingabstractModern Machine learning techniques take advantage of the exponentially rising calculation power in new generation processor units. Thus, the number of parameters which are trained to solve complex tasks was highly increased over the last decades. However, still the networks fail - in contrast to our brain - to develop general intelligence in the sense of being able to solve several complex tasks with only one network architecture. This could be the case because the brain is not a randomly initialized neural network, which has to be trained from scratch by simply investing a lot of calculation power, but has from birth some fixed hierarchical structure. To make progress in decoding the structural basis of biological neural networks we here chose a bottom-up approach, where we evolutionarily trained small neural networks in performing a maze task. This simple maze task requires dynamic decision making with delayed rewards. We were able to show that during the evolutionary optimization random severance of connections leads to better generalization performance of the networks compared to fully connected networks. We conclude that sparsity is a central property of neural networks and should be considered for modern Machine learning approaches. Richard Gerum, André Erpenbeck, Patrick Krauss, Achim Schilling |
Neural Networks | 4 |