Xiaolong Zou

dblp:135/8911 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Shaping Sequence Attractor Schema in Recurrent Neural Networks
abstract
Sequence schemas are abstract, reusable knowledge structures that facilitate rapid adaptation and generalization in novel sequential tasks. In both animals and humans, shaping is an efficient way for acquiring such schemas, particularly in complex sequential tasks. As a form of curriculum learning, shaping works by progressively advancing from simple subtasks to integrated full sequences, and ultimately enabling generalization across different task variations. Despite the importance of schemas in cognition and shaping in schema acquisition, the underlying neural dynamics at play remain poorly understood. To explore this, we train recurrent neural networks on an odor-sequence task using a shaping protocol inspired by well-established paradigms in experimental neuroscience. Our model provides the first systematic reproduction of key features of schema learning observed in the orbitofrontal cortex, including rapid adaptation to novel tasks, structured neural representation geometry, and progressive dimensionality compression during learning. Crucially, analysis of the trained RNN reveals that the learned schema is implemented through sequence attractors. These attractor dynamics emerge gradually through the shaping process: starting with isolated discrete attractors in simple tasks, evolving into linked sequences, and eventually abstracting into generalizable attractors that capture shared task structure. Moreover, applying our method to a keyword spotting task shows that shaping facilitates the rapid development of sequence attractor-like schemas, leading to enhanced learning efficiency. In summary, our work elucidates a novel attractor-based mechanism underlying schema representation and its evolution via shaping, with the potential to provide new insights into the acquisition of abstract knowledge across biological and artificial intelligence.
Zhikun Chu, Bo Ho, Xiaolong Zou, Yuanyuan Mi
NeurIPS3
2024 Continuous Rotation Group Equivariant Network Inspired by Neural Population Coding
abstract
Neural population coding can represent continuous information by neurons with a series of discrete preferred stimuli, and we find that the bell-shaped tuning curve plays an important role in this mechanism. Inspired by this, we incorporate a bell-shaped tuning curve into the discrete group convolution to achieve continuous group equivariance. Simply, we modulate group convolution kernels by Gauss functions to obtain bell-shaped tuning curves. Benefiting from the modulation, kernels also gain smooth gradients on geometric dimensions (e.g., location dimension and orientation dimension). It allows us to generate group convolution kernels from sparse weights with learnable geometric parameters, which can achieve both competitive performances and parameter efficiencies. Furthermore, we quantitatively prove that discrete group convolutions with proper tuning curves (bigger than 1x sampling step) can achieve continuous equivariance. Experimental results show that 1) our approach achieves very competitive performances on MNIST-rot with at least 75% fewer parameters compared with previous SOTA methods, which is efficient in parameter; 2) Especially with small sample sizes, our approach exhibits more pronounced performance improvements (up to 24%); 3) It also has excellent rotation generalization ability on various datasets such as MNIST, CIFAR, and ImageNet with both plain and ResNet architectures.
Zhiqiang Chen 0002, Yang Chen 0071, Xiaolong Zou
AAAI3
2024 DR-Label: Label Deconstruction and Reconstruction of GNN Models for Catalysis Systems
abstract
Attaining the equilibrium geometry of a catalyst-adsorbate system is key to fundamentally assessing its effective properties, such as adsorption energy. While machine learning methods with advanced representation or supervision strategies have been applied to boost and guide the relaxation processes of catalysis systems, existing methods that produce linearly aggregated geometry predictions are susceptible to edge representations ambiguity, and are therefore vulnerable to graph variations. In this paper, we present a novel graph neural network (GNN) supervision and prediction strategy DR-Label. Our approach mitigates the multiplicity of solutions in edge representation and encourages model predictions that are independent of graph structural variations. DR-Label first Deconstructs finer-grained equilibrium state information to the model by projecting the node-level supervision signal to each edge. Reversely, the model Reconstructs a more robust equilibrium state prediction by converting edge-level predictions to node-level via a sphere-fitting algorithm. When applied to three fundamentally different models, DR-Label consistently enhanced performance. Leveraging the graph structure invariance of the DR-Label strategy, we further propose DRFormer, which applied explicit intermediate positional update and achieves a new state-of-the-art performance on the Open Catalyst 2020 (OC20) dataset and the Cu-based single-atom alloys CO adsorption (SAA) dataset. We expect our work to highlight vital principles for advancing geometric GNN models for catalysis systems and beyond. Our code is available at https://github.com/bowenwang77/DR-Label
Bowen Wang 0017, Jiezhong Qiu, Furui Liu, Shaogang Hao, Dong Li 0016, Guangyong Chen, Xiaolong Zou, Pheng-Ann Heng
AAAI9
2024 Leveraging Attractor Dynamics in Spatial Navigation for Better Language Parsing
abstract
Increasing experimental evidence suggests that the human hippocampus, evolutionarily shaped by spatial navigation tasks, also plays an important role in language comprehension, indicating a shared computational mechanism for both functions. However, the specific relationship between the hippocampal formation's computational mechanism in spatial navigation and its role in language processing remains elusive. To investigate this question, we develop a prefrontal-hippocampal-entorhinal model (which called PHE-trinity) that features two key aspects: 1) the use of a modular continuous attractor neural network to represent syntactic structure, akin to the grid network in the entorhinal cortex; 2) the creation of two separate input streams, mirroring the factorized structure-content representation found in the hippocampal formation. We evaluate our model on language command parsing tasks, specifically using the SCAN dataset. Our findings include: 1) attractor dynamics can facilitate systematic generalization and efficient learning from limited data; 2) through visualization and reverse engineering, we unravel a potential dynamic mechanism for grid network representing syntactic structure. Our research takes an initial step in uncovering the dynamic mechanism shared by spatial navigation and language information processing.
Xiaolong Zou, Xingxing Cao, Xiaojiao Yang
ICML1
2023 Learning and processing the ordinal information of temporal sequences in recurrent neural circuits
abstract
Temporal sequence processing is fundamental in brain cognitive functions. Experimental data has indicated that the representations of ordinal information and contents of temporal sequences are disentangled in the brain, but the neural mechanism underlying this disentanglement remains largely unclear. Here, we investigate how recurrent neural circuits learn to represent the abstract order structure of temporal sequences, and how this disentangled representation of order structure from that of contents facilitates the processing of temporal sequences. We show that with an appropriate learn protocol, a recurrent neural circuit can learn a set of tree-structured attractor states to encode the corresponding tree-structured orders of given temporal sequences. This abstract temporal order template can then be bound with different contents, allowing for flexible and robust temporal sequence processing. Using a transfer learning task, we demonstrate that the reuse of a temporal order template facilitates the acquisition of new temporal sequences of the same or similar ordinal structure. Using a key-word spotting task, we demonstrate that the attractor representation of order structure improves the robustness of temporal sequence discrimination, if the ordinal information is the key to differentiate different sequences. We hope this study gives us insights into the neural mechanism of representing the ordinal information of temporal sequences in the brain, and helps us to develop brain-inspired temporal sequence processing algorithms.
Xiaolong Zou, Zhikun Chu, Qinghai Guo, Bo Ho, Si Wu 0001, Yuanyuan Mi
NeurIPS1
2023 Visual information processing through the interplay between fine and coarse signal pathways
abstract
Object recognition is often viewed as a feedforward, bottom-up process in machine learning, but in real neural systems, object recognition is a complicated process which involves the interplay between two signal pathways. One is the parvocellular pathway (P-pathway), which is slow and extracts fine features of objects; the other is the magnocellular pathway (M-pathway), which is fast and extracts coarse features of objects. It has been suggested that the interplay between the two pathways endows the neural system with the capacity of processing visual information rapidly, adaptively, and robustly. However, the underlying computational mechanism remains largely unknown. In this study, we build a two-pathway model to elucidate the computational properties associated with the interactions between two visual pathways. Specifically, we model two visual pathways using two convolution neural networks: one mimics the P-pathway, referred to as FineNet, which is deep, has small-size kernels, and receives detailed visual inputs; the other mimics the M-pathway, referred to as CoarseNet, which is shallow, has large-size kernels, and receives blurred visual inputs. We show that CoarseNet can learn from FineNet through imitation to improve its performance, FineNet can benefit from the feedback of CoarseNet to improve its robustness to noise; and the two pathways interact with each other to achieve rough-to-fine information processing. Using visual backward masking as an example, we further demonstrate that our model can explain visual cognitive behaviors that involve the interplay between two pathways. We hope that this study gives us insight into understanding the interaction principles between two visual pathways.
Xiaolong Zou, Zilong Ji, Tianqiu Zhang, Tiejun Huang 0001, Si Wu 0001
Neural Networks1
2022 Neural feedback facilitates rough-to-fine information retrieval
abstract
Categorical relationships between objects are encoded as overlapped neural representations in the brain, where the more similar the objects are, the larger the correlations between their evoked neuronal responses. These representation correlations, however, inevitably incur interference when memories are retrieved. Here, we propose that neural feedback, which is widely observed in the brain but whose function remains largely unknown, contributes to disentangle neural correlations to improve information retrieval. We study a hierarchical neural network storing the hierarchical categorical information of objects, and information retrieval goes from rough-to-fine, aided by the push-pull neural feedback. We elucidate that the push and the pull components of the feedback suppress the interferences due to the representation correlations between objects from different and the same categories, respectively. Our model reproduces the push-pull phenomenon observed in neural data and sheds light on our understanding of the role of feedback in neural information processing.
Xiaolong Zou, Zilong Ji, Gengshuo Tian, Yuanyuan Mi, Tiejun Huang 0001, K. Y. Michael Wong, Si Wu 0001
Neural Networks2
2021 A Just-In-Time Compilation Approach for Neural Dynamics Simulation
Chaoming Wang, Yingqian Jiang, Xiaohan Lin, Xiaolong Zou, Zilong Ji, Si Wu 0001
ICONIP (3)5
2021 A brain-inspired computational model for spatio-temporal information processing
abstract
Spatio-temporal information processing is fundamental in both brain functions and AI applications. Current strategies for spatio-temporal pattern recognition usually involve explicit feature extraction followed by feature aggregation, which requires a large amount of labeled data. In the present study, motivated by the subcortical visual pathway and early stages of the auditory pathway for motion and sound processing, we propose a novel brain-inspired computational model for generic spatio-temporal pattern recognition. The model consists of two modules, a reservoir module and a decision-making module. The former projects complex spatio-temporal patterns into spatially separated neural representations via its recurrent dynamics, the latter reads out neural representations via integrating information over time, and the two modules are linked together using known examples. Using synthetic data, we demonstrate that the model can extract the frequency and order information of temporal inputs. We apply the model to reproduce the looming pattern discrimination behavior as observed in experiments successfully. Furthermore, we apply the model to the gait recognition task, and demonstrate that our model accomplishes the recognition in an event-based manner and outperforms deep learning counterparts when training data is limited.
Xiaohan Lin, Xiaolong Zou, Zilong Ji, Tiejun Huang 0001, Si Wu 0001, Yuanyuan Mi
Neural Networks2
2020 An Attention-Driven Two-Stage Clustering Method for Unsupervised Person Re-identification
Zilong Ji, Xiaolong Zou, Xiaohan Lin, Tiejun Huang 0001, Si Wu 0001
ECCV (28)2
2019 Push-pull Feedback Implements Hierarchical Information Retrieval Efficiently
abstract
Experimental data has revealed that in addition to feedforward connections, there exist abundant feedback connections in a neural pathway. Although the importance of feedback in neural information processing has been widely recognized in the field, the detailed mechanism of how it works remains largely unknown. Here, we investigate the role of feedback in hierarchical information retrieval. Specifically, we consider a hierarchical network storing the hierarchical categorical information of objects, and information retrieval goes from rough to fine, aided by dynamical push-pull feedback from higher to lower layers. We elucidate that the push (positive) and pull (negative) feedbacks suppress the interferences due to neural correlations between different and the same categories, respectively, and their joint effect improves retrieval performance significantly. Our model agrees with the push-pull phenomenon observed in neural data and sheds light on our understanding of the role of feedback in neural information processing.
Xiaolong Zou, Zilong Ji, Gengshuo Tian, Yuanyuan Mi, Tiejun Huang 0001, K. Y. Michael Wong, Si Wu 0001
NeurIPS2
2018 Neural Information Processing in Hierarchical Prototypical Networks
Zilong Ji, Xiaolong Zou, Tiejun Huang 0001, Yuanyuan Mi, Si Wu 0001
ICONIP (3)2
2018 Learning, Storing, and Disentangling Correlated Patterns in Neural Networks
Xiaolong Zou, Zilong Ji, Tiejun Huang 0001, Yuanyuan Mi, Dahui Wang, Si Wu 0001
ICONIP (3)1
2017 Learning a Continuous Attractor Neural Network from Real Images
Xiaolong Zou, Zilong Ji, Yuanyuan Mi, K. Y. Michael Wong, Si Wu 0001
ICONIP (4)1