VLDB 2026 Research / reviewers in the wild / expert
Yuetong Fang
dblp:330/5208
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-0228-9082ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Deep learning architectures and training · 32% Efficient and distributed learning · 25% Generative modeling · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Emerging computing paradigms · 96% Energy-efficient computing · 4% |
Topics — the 25 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Emerging computing paradigms
neuromorphic computing |
2.7 | 3 | 2026 | TDSNNs: Competitive Topographic Deep Spiking Neural Networks for Visual Cortex Modeling · AAAI 2026 Spiking Neural Networks Need High-Frequency Information · NeurIPS 2025 Adaptive Calibration: A Unified Conversion Framework of Spiking Neural Networks · AAAI 2025 |
Emerging computing paradigms › neuromorphic computing
spiking neural network |
2.7 | 3 | 2026 | TDSNNs: Competitive Topographic Deep Spiking Neural Networks for Visual Cortex Modeling · AAAI 2026 Spiking Neural Networks Need High-Frequency Information · NeurIPS 2025 Adaptive Calibration: A Unified Conversion Framework of Spiking Neural Networks · AAAI 2025 |
Machine learning › Deep learning architectures and training
spiking neural network |
2.4 | 3 | 2026 | Dynamic Token Masking in Spiking Neural Network · Int. J. Comput. Vis. 2026 Spiking Wavelet Transformer · ECCV (76) 2024 Masked Spiking Transformer · ICCV 2023 |
Machine learning › Efficient and distributed learning
model compression |
1.9 | 2 | 2026 | Dynamic Token Masking in Spiking Neural Network · Int. J. Comput. Vis. 2026 Optimal Brain Apoptosis · ICLR 2025 |
Machine learning › Deep learning architectures and training › spiking neural network
spiking transformer |
1.5 | 2 | 2025 | Spiking Neural Networks Need High-Frequency Information · NeurIPS 2025 Masked Spiking Transformer · ICCV 2023 |
Machine learning › Deep learning architectures and training
biologically inspired neural network |
1.0 | 1 | 2026 | TDSNNs: Competitive Topographic Deep Spiking Neural Networks for Visual Cortex Modeling · AAAI 2026 |
Machine learning › Generative modeling
diffusion model |
1.0 | 1 | 2026 | Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and Refine · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Generative modeling › diffusion model
diffusion model acceleration |
1.0 | 1 | 2026 | Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and Refine · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
1.0 | 1 | 2026 | Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and Refine · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Natural language and speech › Language models and text generation
token masking |
1.0 | 1 | 2026 | Dynamic Token Masking in Spiking Neural Network · Int. J. Comput. Vis. 2026 |
Computer vision › 3D vision › biological vision modeling
visual cortex modeling |
1.0 | 1 | 2026 | TDSNNs: Competitive Topographic Deep Spiking Neural Networks for Visual Cortex Modeling · AAAI 2026 |
Computer vision › Video understanding and tracking
action recognition |
0.9 | 1 | 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Deep learning architectures and training › spiking neural network
ANN-to-SNN conversion |
0.9 | 1 | 2025 | Adaptive Calibration: A Unified Conversion Framework of Spiking Neural Networks · AAAI 2025 |
Computer vision › Video understanding and tracking › action recognition
event-based action recognition |
0.9 | 1 | 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Efficient and distributed learning › model deployment
model conversion |
0.9 | 1 | 2025 | Adaptive Calibration: A Unified Conversion Framework of Spiking Neural Networks · AAAI 2025 |
Computer vision › Image recognition and object detection › point set representation
point cloud representation |
0.9 | 1 | 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.9 | 1 | 2025 | Optimal Brain Apoptosis · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression › pruning
second-order pruning |
0.9 | 1 | 2025 | Optimal Brain Apoptosis · ICLR 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Spiking Wavelet Transformer · ECCV (76) 2024 |
Machine learning › Efficient and distributed learning › energy-efficient learning
energy-efficient neural network |
0.7 | 1 | 2023 | Masked Spiking Transformer · ICCV 2023 |
Machine learning › Generative modeling › autoregressive model
autoregressive image generation |
0.3 | 1 | 2026 | Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and Refine · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Generative modeling
autoregressive model |
0.3 | 1 | 2026 | Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and Refine · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › 3D vision › visual localization
camera relocalization |
0.3 | 1 | 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Image recognition and object detection
image classification |
0.3 | 1 | 2025 | Spiking Neural Networks Need High-Frequency Information · NeurIPS 2025 |
Energy-efficient computing › energy-efficient machine learning
energy-efficient inference |
0.3 | 1 | 2025 | Adaptive Calibration: A Unified Conversion Framework of Spiking Neural Networks · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
topographic organization · 2.0spatio-temporal constraints loss · 2.0spike compression · 1.7adaptive firing neuron model · 1.7weak-to-strong guidance · 1.0tuning-based inference · 1.0dynamic token masking · 1.0classifier-free guidance · 1.0max-pooling · 0.9hessian-vector products · 0.9depthwise convolution · 0.9adaptive timesteps · 0.9adaptive time step · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TDSNNs: Competitive Topographic Deep Spiking Neural Networks for Visual Cortex ModelingabstractThe primate visual cortex exhibits topographic organization, where functionally similar neurons are spatially clustered, a structure widely believed to enhance neural processing efficiency. While prior works have demonstrated that conventional deep ANNs can develop topographic representations, these models largely neglect crucial temporal dynamics. This oversight often leads to significant performance degradation in tasks like object recognition and compromises their biological fidelity. To address this, we leverage spiking neural networks (SNNs), which inherently capture spike-based temporal dynamics and offer enhanced biological plausibility. We propose a novel Spatio-Temporal Constraints (STC) loss function for topographic deep spiking neural networks (TDSNNs), successfully replicating the hierarchical spatial functional organization observed in the primate visual cortex from low-level sensory input to high-level abstract representations. Our results show that STC effectively generates representative topographic features across simulated visual cortical areas. While introducing topography typically leads to significant performance degradation in ANNs, our spiking architecture exhibits a remarkably small performance drop (No drop in ImageNet top-1 accuracy, compared to a 3% drop observed in TopoNet, which is the best-performing topographic ANN so far) and outperforms topographic ANNs in brain-likeness. We also reveal that topographic organization facilitates efficient and stable temporal information processing via the spike mechanism in TDSNNs, contributing to model robustness. These findings suggest that TDSNNs offer a compelling balance between computational performance and brain-like features, providing not only a framework for interpreting neural science phenomena but also novel insights for designing more efficient and robust deep learning models. Deming Zhou, Yuetong Fang, Zhaorui Wang 0006, Renjing Xu |
AAAI | 2 |
| 2026 | Dynamic Token Masking in Spiking Neural Network
Yuetong Fang, Deming Zhou, Shibo Zhou, Renjing Xu |
Int. J. Comput. Vis. | 1 |
| 2026 | Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and RefineabstractRecently, enhancing the generative capability of text-to-image (T2I) models has become a promising direction in both academia and industry. Prior studies often focused on either improving generative quality or reducing inference latency, but typically failed to improve both quality and speed simultaneously. Moreover, existing inference-enhancement methods do not achieve significant improvements simultaneously across both diffusion models (DMs) and autoregressive models (ARMs). In this paper, we introduce a general tuning-based inference-enhancement framework, named CoRe$^{2}$2, which is the first to simultaneously achieve significant generative quality and reduced inference overhead across DMs and ARMs, to the best of our knowledge. CoRe$^{2}$2 comprises three stages: Collect, Reflect, and Refine. During the Collect stage, classifier-free guidance (CFG) trajectories are collected and subsequently used in the Reflect stage to train a weak model capable of reflecting the "easy-to-learn" content. Finally, during the Refine stage, CoRe$^{2}$2 can utilize the trained weak model to achieve speedup and performance gain in inference. Specifically, in the early sampling steps, CoRe$^{2}$2 employs weak-to-strong guidance to refine the "difficult-to-learn" and realistic content, thereby improving generative quality. In the later sampling steps, CoRe$^{2}$2 can use the weak model to generate "easy-to-learn" content instead of CFG, dramatically reducing inference time. Experimental outcomes substantiates CoRe$^{2}$2 achieve significant performance improvements on HPD v2, Pick-of-Pic, Drawbench, GenEval, and T2I-Compbench across SDXL, SD3.5, FLUX and LlamaGen. Notably, for SD3.5, CoRe$^{2}$2 can be seamlessly integrated with the state-of-the-art inference-enhancement algorithm Z-Sampling, outperforming it even with less time. Shitong Shao, Zikai Zhou, Dian Xie, Yuetong Fang, Tian Ye 0001, Lichen Bai, Bo Han 0003, Zeke Xie |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Adaptive Calibration: A Unified Conversion Framework of Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are seen as an energy-efficient alternative to traditional Artificial Neural Networks (ANNs), but the performance gap remains a challenge. While this gap is narrowing through ANN-to-SNN conversion, substantial computational resources are still needed, and the energy efficiency of converted SNNs cannot be ensured. To address this, we present a unified training-free conversion framework that significantly enhances both the performance and efficiency of converted SNNs. Inspired by the biological nervous system, we propose a novel Adaptive-Firing Neuron Model (AdaFire), which dynamically adjusts firing patterns across different layers to substantially reduce the Unevenness Error - the primary source of error of converted SNNs within limited inference timesteps. We further introduce two efficiency-enhancing techniques: the Sensitivity Spike Compression (SSC) technique for reducing spike operations, and the Input-aware Adaptive Timesteps (IAT) technique for decreasing latency. These methods collectively enable our approach to achieve state-of-the-art performance with significant energy savings of up to 70.1%, 60.3%, and 43.1% on CIFAR-10, CIFAR-100, and ImageNet datasets, respectively. Extensive experiments across 2D, 3D, event-driven classification tasks, object detection, and segmentation tasks, demonstrate the effectiveness of our method in various domains. Yuetong Fang, Jiahang Cao, Renjing Xu |
AAAI | 2 |
| 2025 | Reuse and Blend: A Weight-Sharing Energy-Efficient Optical Neural NetworkabstractOptical neural networks (ONN) based on micro-ring resonators (MRR) have emerged as a promising alternative to significantly accelerating the massive matrix-vector multiplication (MVM) operations in artificial intelligence (AI) applications. However, the limited scale of MRR arrays presents a challenge for AI acceleration. The disparity between the small MRR arrays and the large weight matrices in AI necessitates extensive MRR writings, including reprogramming and calibration, resulting in considerable latency and energy overheads. To address this problem, we propose a novel design methodology to lessen the need for frequent weight reloading. Specifically, we propose a reuse and blend (R&B) architecture to support efficient layer-wise and block-wise weight sharing, which allows weights to be reused several times between layers/ blocks. Experiment results demonstrate the R&B system can maintain comparable accuracy with 69% energy savings and 57% latency improvement. These results highlight the promise of the R&B to enable efficient deployment of advanced deep learning models on photonic accelerators. Yuetong Fang, Shaoliang Yu, Renjing Xu |
ASP-DAC | 2 |
| 2025 | Optimal Brain ApoptosisabstractThe increasing complexity and parameter count of Convolutional Neural Networks (CNNs) and Transformers pose challenges in terms of computational efficiency and resource demands. Pruning has been identified as an effective strategy to address these challenges by removing redundant elements such as neurons, channels, or connections, thereby enhancing computational efficiency without heavily compromising performance. This paper builds on the foundational work of Optimal Brain Damage (OBD) by advancing the methodology of parameter importance estimation using the Hessian matrix. Unlike previous approaches that rely on approximations, we introduce Optimal Brain Apoptosis (OBA), a novel pruning method that calculates the Hessian-vector product value directly for each parameter. By decomposing the Hessian matrix across network layers and identifying conditions under which inter-layer Hessian submatrices are non-zero, we propose a highly efficient technique for computing the second-order Taylor expansion of parameters. This approach allows for a more precise pruning process, particularly in the context of CNNs and Transformers, as validated in our experiments including VGG19, ResNet32, ResNet50, and ViT-B/16 on CIFAR10, CIFAR100 and Imagenet datasets. Our code is available at https://github.com/NEU-REAL/OBA. Zheng Fang 0001, Delei Kong, Chenming Hu, Yuetong Fang, Renjing Xu |
ICLR | 7 |
| 2025 | Spiking Neural Networks Need High-Frequency InformationabstractSpiking Neural Networks promise brain-inspired and energy-efficient computation by transmitting information through binary (0/1) spikes. Yet, their performance still lags behind that of artificial neural networks, often assumed to result from information loss caused by sparse and binary activations. In this work, we challenge this long-standing assumption and reveal a previously overlooked frequency bias: **spiking neurons inherently suppress high-frequency components and preferentially propagate low-frequency information.** This frequency-domain imbalance, we argue, is the root cause of degraded feature representation in SNNs. Empirically, on Spiking Transformers, adopting Avg-Pooling (low-pass) for token mixing lowers performance to 76.73% on Cifar-100, whereas replacing it with Max-Pool (high-pass) pushes the top-1 accuracy to 79.12%. Accordingly, we introduce **Max-Former** that restores high-frequency signals through two frequency-enhancing operators: (1) extra Max-Pool in patch embedding, and (2) Depth-Wise Convolution in place of self-attention. Notably, **Max-Former** attains 82.39% top-1 accuracy on ImageNet using only 63.99M parameters, surpassing Spikformer (74.81%, 66.34M) by +7.58%. Extending our insight beyond transformers, our **Max-ResNet-18** achieves state-of-the-art performance on convolution-based benchmarks: 97.17% on CIFAR-10 and 83.06% on CIFAR-100. We hope this simple yet effective solution inspires future research to explore the distinctive nature of spiking neural networks. Code is available: https://github.com/bic-L/MaxFormer. Yuetong Fang, Deming Zhou, ZeCui Zeng, Lusong Li, Shibo Zhou, Renjing Xu |
NeurIPS | 1 |
| 2025 | Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and RegressionabstractEvent cameras draw inspiration from biological systems, boasting low latency and high dynamic range while consuming minimal power. The most current approach to processing Event Cloud often involves converting it into frame-based representations, which neglects the sparsity of events, loses fine-grained temporal information, and increases the computational burden. In contrast, Point Cloud is a popular representation for processing 3-dimensional data and serves as an alternative method to exploit local and global spatial features. Nevertheless, previous point-based methods show an unsatisfactory performance compared to the frame-based method in dealing with spatio-temporal event streams. In order to bridge the gap, we propose EventMamba, an efficient and effective framework based on Point Cloud representation by rethinking the distinction between Event Cloud and Point Cloud, emphasizing vital temporal information. The Event Cloud is subsequently fed into a hierarchical structure with staged modules to process both implicit and explicit temporal features. Specifically, we redesign the global extractor to enhance explicit temporal extraction among a long sequence of events with temporal aggregation and State Space Model (SSM) based Mamba. Our model consumes minimal computational resources in the experiments and still exhibits SOTA point-based performance on six different scales of action recognition datasets. It even outperformed all frame-based methods on both Camera Pose Relocalization (CPR) and eye-tracking regression tasks. Yue Zhou 0010, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang 0001, Yuetong Fang, Fei Ma 0006, Hao Yu 0001, Bojun Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | Spiking Wavelet Transformer
Yuetong Fang, Jiahang Cao, Honglei Chen, Renjing Xu |
ECCV (76) | 1 |
| 2023 | Masked Spiking TransformerabstractThe combination of Spiking Neural Networks (SNNs) and Transformers has attracted significant attention due to their potential for high energy efficiency and high-performance nature. However, existing works on this topic typically rely on direct training, which can lead to suboptimal performance. To address this issue, we propose to leverage the benefits of the ANN-to-SNN conversion method to combine SNNs and Transformers, resulting in significantly improved performance over existing state-of-the-art SNN models. Furthermore, inspired by the quantal synaptic failures observed in the nervous system, which reduce the number of spikes transmitted across synapses, we introduce a novel Masked Spiking Transformer (MST) framework. This incorporates a Random Spike Masking (RSM) method to prune redundant spikes and reduce energy consumption without sacrificing performance. Our experimental results demonstrate that the proposed MST model achieves a significant reduction of 26.8% in power consumption when the masking ratio is 75% while maintaining the same level of performance as the unmasked model. The code is available at: https://github.com/bic-L/Masked-Spiking-Transformer. Yuetong Fang, Jiahang Cao, Qiang Zhang 0029, Renjing Xu |
ICCV | 2 |