Zixuan Gong

dblp:363/7481 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0008-8961-5437ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 26% Generative modeling · 17% Efficient and distributed learning · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
brain decoding
1.932025
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction · AAAI 2025
NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction · NeurIPS 2024
Lite-Mind: Towards Efficient and Robust Brain Representation Learning · ACM Multimedia 2024
Machine learning › Learning theory
generalization bounds
1.722025
Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization · ICLR 2025
Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology · AAAI 2025
Computer vision › 3D vision › brain decoding
brain visual decoding
0.912025
Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding · AAAI 2025
Computer vision › 3D vision › brain decoding
fMRI-to-image reconstruction
0.912025
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction · AAAI 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization · ICLR 2025
Machine learning › Transfer learning and domain adaptation
meta-learning
0.912025
Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding · AAAI 2025
Machine learning › Learning theory › PAC-Bayesian analysis
PAC-Bayesian generalization bound
0.912025
Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization · ICLR 2025
Machine learning › Optimization for machine learning › black-box optimization
zeroth-order optimization
0.912025
Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology · AAAI 2025
Medical and health informatics
neuroimaging
0.912025
MindSimulator: Exploring Brain Concept Localization via Synthetic fMRI · ICLR 2025
Computer vision › 3D vision › brain decoding
brain-to-image retrieval
0.812024
Lite-Mind: Towards Efficient and Robust Brain Representation Learning · ACM Multimedia 2024
Computer vision › Vision and language › vision-language model
CLIP
0.812024
MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization · ICML 2024
Machine learning › Generative modeling
diffusion model
0.812024
NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction · NeurIPS 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization · ICML 2024
Machine learning › Generative modeling › video generation
text-to-video generation
0.812024
NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › token compression
token merging
0.812024
MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization · ICML 2024
Computer vision › Vision and language
vision-language pretraining
0.812024
MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization · ICML 2024
Computer vision › Vision and language
cross-modal retrieval
0.312025
Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding · AAAI 2025
Machine learning › Efficient and distributed learning › distributed training
decentralized learning
0.312025
Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology · AAAI 2025
Machine learning › Efficient and distributed learning › distributed training
distributed stochastic gradient descent
0.312025
Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology · AAAI 2025
Machine learning › Generative modeling
image reconstruction
0.312025
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction · AAAI 2025
Machine learning › Generative modeling › autoregressive model
next-token prediction
0.312025
Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization · ICLR 2025
Image and video processing
frequency domain analysis
0.212024
MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization · ICML 2024

Methods — techniques the papers use, named apart from their topics

generative modeling · 1.7zeroth-order optimization · 0.9stochastic gradient descent · 0.9mixture of experts · 0.9meta-learning · 0.9generalization analysis · 0.9fMRI alignment · 0.9contrastive alignment · 0.9Skip-LoRA · 0.9LoRA fine-tuning · 0.9token merging · 0.8frequency transform · 0.8contrastive learning · 0.8
YearPublicationVenuePosition
2025 Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding
abstract
Decoding visual information from human brain activity has seen remarkable advancements in recent research. However, the diversity in cortical parcellation and fMRI patterns across individuals has prompted the development of deep learning models tailored to each subject. The personalization limits the broader applicability of brain visual decoding in real-world scenarios. To address this issue, we introduce Wills Aligner, a novel approach designed to achieve multi-subject collaborative brain visual decoding. Wills Aligner begins by aligning the fMRI data from different subjects at the anatomical level. It then employs delicate mixture-of-brain-expert adapters and a meta-learning strategy to account for individual fMRI pattern differences. Additionally, Wills Aligner leverages the semantic relation of visual stimuli to guide the learning of inter-subject commonality, enabling visual decoding for each subject to draw insights from other subjects' data. We rigorously evaluate our Wills Aligner across various visual decoding tasks, including classification, cross-modal retrieval, and image reconstruction. The experimental results demonstrate that Wills Aligner achieves promising performance.
Guangyin Bao, Qi Zhang 0020, Zixuan Gong, Jialei Zhou, Wei Fan 0010, Kun Yi 0001, Usman Naseem, Liang Hu 0004, Duoqian Miao 0001
AAAI3
2025 MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
abstract
Decoding natural visual scenes from brain activity has flourished, with extensive research in single-subject tasks and, however, less in cross-subject tasks. Reconstructing high-quality images in cross-subject tasks is a challenging problem due to profound individual differences between subjects and the scarcity of data annotation. In this work, we proposed MindTuner for cross-subject visual decoding, which achieves high-quality and rich semantic reconstructions using only 1 hour of fMRI training data benefiting from the phenomena of visual fingerprint in the human visual system and a novel fMRI-to-text alignment paradigm. Firstly, we pre-train a multi-subject model among 7 subjects and fine-tune it with scarce data on new subjects, where LoRAs with Skip-LoRAs are utilized to learn the visual fingerprint. Then, we take the image modality as the intermediate pivot modality to achieve fMRI-to-text alignment, which achieves impressive fMRI-to-text retrieval performance and corrects fMRI-to-image reconstruction with fine-tuned semantics. The results of both qualitative and quantitative analyses demonstrate that MindTuner surpasses state-of-the-art cross-subject visual decoding models on the Natural Scenes Dataset (NSD), whether using training data of 1 hour or 40 hours.
Zixuan Gong, Qi Zhang 0020, Guangyin Bao, Rongtao Xu, Liang Hu 0004, Duoqian Miao 0001
AAAI1
2025 Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology
abstract
Zeroth-order (ZO) optimization as the gradient-free method has become a powerful tool when the first-order gradient is unavailable or expensive to obtain, especially in decentralized learning scenarios where data and computational resources are distributed across multiple clients. There have been many efforts to analyze the optimization convergence rate of zeroth-order decentralized stochastic gradient descent (ZO-DSGD) algorithms. However, the generalization of these methods has not been well studied. In this paper, we provide a generalization analysis of ZO-DSGD with changing topology, where the clients run zeroth-order SGD with local data and communicate with each other according to time-varying topology. We systematically analyze the generalization error in convex, strongly convex, and non-convex cases. The obtained results in the convex and strongly convex cases with zeroth-order oracles recover the results of SGD. Moreover, the generalization bounds derived in non-convex cases align with that of DSGD. To capture the influence of communication topology on the generalization performance, we analyze local generalization bounds concerning local models held at different clients. The obtained results reflect the influence of the number of clients, local sample size, and topology on the generalization error. To the best of our knowledge, this is the first work that provides a generalization analysis of zeroth-order decentralized stochastic gradient descent methods and recovers the results of SGD.
Xiaolin Hu 0001, Zixuan Gong, Gengze Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020
AAAI2
2025 MindSimulator: Exploring Brain Concept Localization via Synthetic fMRI
abstract
Concept-selective regions within the human cerebral cortex exhibit significant activation in response to specific visual stimuli associated with particular concepts. Precisely localizing these regions stands as a crucial long-term goal in neuroscience to grasp essential brain functions and mechanisms. Conventional experiment-driven approaches hinge on manually constructed visual stimulus collections and corresponding brain activity recordings, constraining the support and coverage of concept localization. Additionally, these stimuli often consist of concept objects in unnatural contexts and are potentially biased by subjective preferences, thus prompting concerns about the validity and generalizability of the identified regions. To address these limitations, we propose a data-driven exploration approach. By synthesizing extensive brain activity recordings, we statistically localize various concept-selective regions. Our proposed MindSimulator leverages advanced generative technologies to learn the probability distribution of brain activity conditioned on concept-oriented visual stimuli. This enables the creation of simulated brain recordings that reflect real neural response patterns. Using the synthetic recordings, we successfully localize several well-studied concept-selective regions and validate them against empirical findings, achieving promising prediction accuracy. The feasibility opens avenues for exploring novel concept-selective regions and provides prior hypotheses for future neuroscience research.
Guangyin Bao, Qi Zhang 0020, Zixuan Gong, Zhuojia Wu, Duoqian Miao 0001
ICLR3
2025 Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization
abstract
Large language models (LLMs) have demonstrated remarkable in-context learning (ICL) abilities. However, existing theoretical analysis of ICL primarily exhibits two limitations: \textbf{(a) Limited \textit{i.i.d.} Setting.} Most studies focus on supervised function learning tasks where prompts are constructed with \textit{i.i.d.} input-label pairs. This \textit{i.i.d.} assumption diverges significantly from real language learning scenarios where prompt tokens are interdependent. \textbf{(b) Lack of Emergence Explanation.} Most literature answers \textbf{\textit{what}} ICL does from an implicit optimization perspective but falls short in elucidating \textbf{\textit{how}} ICL emerges and the impact of pre-training phase on ICL. In our paper, to extend (a), we adopt a more practical paradigm, \textbf{\textit{auto-regressive next-token prediction (AR-NTP)}}, which closely aligns with the actual training of language models. Specifically, within AR-NTP, we emphasize prompt token-dependency, which involves predicting each subsequent token based on the preceding sequence. To address (b), we formalize a systematic pre-training and ICL framework, highlighting the layer-wise structure of sequences and topics, alongside a two-level expectation. In conclusion, we present data-dependent, topic-dependent and optimization-dependent PAC-Bayesian generalization bounds for pre-trained LLMs, investigating that \textbf{\textit{ICL emerges from the generalization of sequences and topics}}. Our theory is supported by experiments on numerical linear dynamic systems, synthetic GINC and real-world language datasets.
Zixuan Gong, Xiaolin Hu 0001, Huayi Tang, Yong Liu 0018
ICLR1
2024 MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization
abstract
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, leading to rapid advancements in multimodal studies. However, CLIP faces a notable challenge in terms of *inefficient data utilization*. It relies on a single contrastive supervision for each image-text pair during representation learning, disregarding a substantial amount of valuable information that could offer richer supervision. Additionally, the retention of non-informative tokens leads to increased computational demands and time costs, particularly in CLIP's ViT image encoder. To address these issues, we propose **M**ulti-Perspective **L**anguage-**I**mage **P**retraining (**MLIP**). In MLIP, we leverage the frequency transform's sensitivity to both high and low-frequency variations, which complements the spatial domain's sensitivity limited to low-frequency variations only. By incorporating frequency transforms and token-level alignment, we expand CILP's single supervision into multi-domain and multi-level supervision, enabling a more thorough exploration of informative image features. Additionally, we introduce a token merging method guided by comprehensive semantics from the frequency and spatial domains. This allows us to merge tokens to multi-granularity tokens with a controllable compression rate to accelerate CLIP. Extensive experiments validate the effectiveness of our design.
Yu Zhang 0133, Qi Zhang 0020, Zixuan Gong, Yiwei Shi, Duoqian Miao 0001, Kun Yi 0001, Wei Fan 0010, Liang Hu 0004, Changwei Wang 0001
ICML3
2024 Lite-Mind: Towards Efficient and Robust Brain Representation Learning
abstract
The limited data availability and the low signal-to-noise ratio of fMRI signals lead to the challenging task of fMRI-to-image retrieval. State-of-the-art MindEye remarkably improves fMRI-to-image retrieval performance by leveraging a large model, i.e., a 996M MLP Backbone per subject, to align fMRI embeddings to the final hidden layer of CLIP's Vision Transformer (ViT). However, significant individual variations exist among subjects, even under identical experimental setups, mandating the training of large subject-specific models. The substantial parameters pose significant challenges in deploying fMRI decoding on practical devices. To this end, we propose Lite-Mind, a lightweight, efficient, and robust brain representation learning paradigm based on Discrete Fourier Transform (DFT), which efficiently aligns fMRI voxels to fine-grained information of CLIP. We elaborately design a DFT backbone with Spectrum Compression and Frequency Projector modules to learn informative and robust voxel embeddings. Our experiments demonstrate that Lite-Mind achieves an impressive 94.6% fMRI-to-image retrieval accuracy on the NSD dataset for Subject 1, with 98.7% fewer parameters than MindEye. Lite-Mind is also proven to be able to be migrated to smaller fMRI datasets and establishes a new state-of-the-art for zero-shot classification on the GOD dataset.
Zixuan Gong, Qi Zhang 0020, Guangyin Bao, Lei Zhu 0002, Yu Zhang 0133, Liang Hu 0004, Duoqian Miao 0001
ACM Multimedia1
2024 NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction
abstract
Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e.g., a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https://github.com/gongzix/NeuroClips.
Zixuan Gong, Guangyin Bao, Qi Zhang 0020, Zhongwei Wan, Duoqian Miao 0001, Shoujin Wang, Lei Zhu 0003, Changwei Wang 0001, Rongtao Xu, Liang Hu 0004, Yu Zhang 0133
NeurIPS1