VLDB 2026 Research / reviewers in the wild / expert
Shiyan Chen
dblp:227/7255
· DBLP profile ↗
15ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpikeCV: open a continuous computer vision era
Yajing Zheng, Jiyuan Zhang 0005, Rui Zhao 0010, Jianhao Ding, Shiyan Chen, Weijian Wu, Ruiqin Xiong, Zhaofei Yu, Tiejun Huang 0001 |
Sci. China Inf. Sci. | 5 |
| 2024 | Transient Glimpses: Unveiling Occluded Backgrounds through the Spike CameraabstractThe de-occlusion problem, involving extracting clear background images by removing foreground occlusions, holds significant practical importance but poses considerable challenges. Most current research predominantly focuses on generating discrete images from calibrated camera arrays, but this approach often struggles with dense occlusions and fast motions due to limited perspectives and motion blur. To overcome these limitations, an effective solution requires the integration of multi-view visual information. The spike camera, as an innovative neuromorphic sensor, shows promise with its ultra-high temporal resolution and dynamic range. In this study, we propose a novel approach that utilizes a single spike camera for continuous multi-view imaging to address occlusion removal. By rapidly moving the spike camera, we capture a dense stream of spikes from occluded scenes. Our model, SpkOccNet, processes these spikes by integrating multi-view spatial-temporal information via long-short-window feature extractor (LSW) and employs a novel cross-view mutual attention-based module (CVA) for effective fusion and refinement. Additionally, to facilitate research in occlusion removal, we introduce the S-OCC dataset, which consists of real-world spike-based data. Experimental results demonstrate the efficiency and generalization capabilities of our model in effectively removing dense occlusions across diverse scenes. Public project page: https://github.com/Leozhangjiyuan/SpikeDeOcclusion. Jiyuan Zhang 0005, Shiyan Chen, Yajing Zheng, Zhaofei Yu, Tiejun Huang 0001 |
AAAI | 2 |
| 2024 | Exploring Efficient Asymmetric Blind-Spots for Self-Supervised Denoising in Real-World ScenariosabstractSelf-supervised denoising has attracted widespread at-tention due to its ability to train without clean images. How-ever, noise in real-world scenarios is often spatially cor-related, which causes many self-supervised algorithms that assume pixel-wise independent noise to perform poorly. Re-cent works have attempted to break noise correlation with downsampling or neighborhood masking. However, denoising on downsampled subgraphs can lead to aliasing effects and loss of details due to a lower sampling rate. Further-more, the neighborhood masking methods either come with high computational complexity or do not consider local spatial preservation during inference. Through the analy-sis of existing methods, we point out that the key to obtaining high-quality and texture-rich results in real-world self-supervised denoising tasks is to train at the original input resolution structure and use asymmetric operations during training and inference. Based on this, we propose Asymmet-ric Tunable Blind-Spot Network (AT-BSN), where the blind-spot size can be freely adjusted, thus better balancing noise correlation suppression and image local spatial destruction during training and inference. In addition, we regard the pre-trained AT-BSN as a meta-teacher network capable of generating various teacher networks by sampling different blind-spots. We propose a blind-spot based multi-teacher distillation strategy to distill a lightweight network, signif-icantly improving performance. Experimental results on multiple datasets prove that our method achieves state-of-the-art, and is superior to other self-supervised algorithms in terms of computational overhead and visual effects. Shiyan Chen, Jiyuan Zhang 0005, Zhaofei Yu, Tiejun Huang 0001 |
CVPR | 1 |
| 2024 | Spike-guided Motion Deblurring with Unknown Modal Spatiotemporal AlignmentabstractThe traditional frame-based cameras that rely on exposure windows for imaging experience motion blur in high-speed scenarios. Frame-based deblurring methods lack reliable motion cues to restore sharp images under extreme blur conditions. The spike camera is a novel neuromorphic visual sensor that outputs spike streams with ultra-high temporal resolution. It can supplement the temporal information lost in traditional cameras and guide motion deblurring. However, in real-world scenarios, aligning discrete RGB images and continuous spike streams along both temporal and spatial axes is challenging due to the complexity of calibrating their coordinates, device displacements in vibrations, and time deviations. Misalignment of pixels leads to severe degradation of deblurring. We introduce the first framework for spike-guided motion deblurring without knowing the spatiotemporal alignment between spikes and images. To address the problem, we first propose a novel three-stage network containing a basic deblurring net, a carefully designed bi-directional deformable aligning module, and a flow-based multi-scale fusion net. Experimental results demonstrate that our approach can effectively guide the image deblurring with unknown alignment, surpassing the performance of other methods. Public project page: https://github.com/Leozhangjiyuan/UaSDN. Jiyuan Zhang 0005, Shiyan Chen, Yajing Zheng, Zhaofei Yu, Tiejun Huang 0001 |
CVPR | 2 |
| 2024 | SpikeGS: 3D Gaussian Splatting from Spike Streams with High-Speed Camera Motion
Jiyuan Zhang 0005, Shiyan Chen, Yajing Zheng, Tiejun Huang 0001, Zhaofei Yu |
ACM Multimedia | 3 |
| 2024 | SpikeReveal: Unlocking Temporal Sequences from Real Blurry Inputs with Spike StreamsabstractReconstructing a sequence of sharp images from the blurry input is crucial for enhancing our insights into the captured scene and poses a significant challenge due to the limited temporal features embedded in the image. Spike cameras, sampling at rates up to 40,000 Hz, have proven effective in capturing motion features and beneficial for solving this ill-posed problem. Nonetheless, existing methods fall into the supervised learning paradigm, which suffers from notable performance degradation when applied to real-world scenarios that diverge from the synthetic training data domain. To address these challenges, we propose the first self-supervised framework for the task of spike-guided motion deblurring. Our approach begins with the formulation of a spike-guided deblurring model that explores the theoretical relationships among spike streams, blurry images, and their corresponding sharp sequences. We subsequently develop a self-supervised cascaded framework to alleviate the issues of spike noise and spatial-resolution mismatching encountered in the deblurring model. With knowledge distillation and re-blurring loss, we further design a lightweight deblur network to generate high-quality sequences with brightness and texture consistency with the original input. Quantitative and qualitative experiments conducted on our real-world and synthetic datasets with spikes validate the superior generalization of the proposed framework. Our code, data and trained models are available at \url{https://github.com/chenkang455/S-SDM}. Shiyan Chen, Jiyuan Zhang 0005, Baoyue Zhang, Yajing Zheng, Tiejun Huang 0001, Zhaofei Yu |
NeurIPS | 2 |
| 2024 | Automatic quantitative stroke severity assessment based on Chinese clinical named entity recognition with domain-adaptive pre-trained large language modelabstractBACKGROUND: Stroke is a prevalent disease with a significant global impact. Effective assessment of stroke severity is vital for an accurate diagnosis, appropriate treatment, and optimal clinical outcomes. The National Institutes of Health Stroke Scale (NIHSS) is a widely used scale for quantitatively assessing stroke severity. However, the current manual scoring of NIHSS is labor-intensive, time-consuming, and sometimes unreliable. Applying artificial intelligence (AI) techniques to automate the quantitative assessment of stroke on vast amounts of electronic health records (EHRs) has attracted much interest. OBJECTIVE: This study aims to develop an automatic, quantitative stroke severity assessment framework through automating the entire NIHSS scoring process on Chinese clinical EHRs. METHODS: Our approach consists of two major parts: Chinese clinical named entity recognition (CNER) with a domain-adaptive pre-trained large language model (LLM) and automated NIHSS scoring. To build a high-performing CNER model, we first construct a stroke-specific, densely annotated dataset "Chinese Stroke Clinical Records" (CSCR) from EHRs provided by our partner hospital, based on a stroke ontology that defines semantically related entities for stroke assessment. We then pre-train a Chinese clinical LLM coined "CliRoberta" through domain-adaptive transfer learning and construct a deep learning-based CNER model that can accurately extract entities directly from Chinese EHRs. Finally, an automated, end-to-end NIHSS scoring pipeline is proposed by mapping the extracted entities to relevant NIHSS items and values, to quantitatively assess the stroke severity. RESULTS: Results obtained on a benchmark dataset CCKS2019 and our newly created CSCR dataset demonstrate the superior performance of our domain-adaptive pre-trained LLM and the CNER model, compared with the existing benchmark LLMs and CNER models. The high F1 score of 0.990 ensures the reliability of our model in accurately extracting the entities for the subsequent automatic NIHSS scoring. Subsequently, our automated, end-to-end NIHSS scoring approach achieved excellent inter-rater agreement (0.823) and intraclass consistency (0.986) with the ground truth and significantly reduced the processing time from minutes to a few seconds. CONCLUSION: Our proposed automatic and quantitative framework for assessing stroke severity demonstrates exceptional performance and reliability through directly scoring the NIHSS from diagnostic notes in Chinese clinical EHRs. Moreover, this study also contributes a new clinical dataset, a pre-trained clinical LLM, and an effective deep learning-based CNER model. The deployment of these advanced algorithms can improve the accuracy and efficiency of clinical assessment, and help improve the quality, affordability and productivity of healthcare services. Zhanzhong Gu, Xiangjian He, Ping Yu 0004, Wenjing Jia, Xiguang Yang, Penghui Hu, Shiyan Chen, Yiguang Lin |
Artif. Intell. Medicine | 8 |
| 2023 | Self-Supervised Joint Dynamic Scene Reconstruction and Optical Flow Estimation for Spiking CameraabstractSpiking camera, a novel retina-inspired vision sensor, has shown its great potential for capturing high-speed dynamic scenes with a sampling rate of 40,000 Hz. The spiking camera abandons the concept of exposure window, with each of its photosensitive units continuously capturing photons and firing spikes asynchronously. However, the special sampling mechanism prevents the frame-based algorithm from being used to spiking camera. It remains to be a challenge to reconstruct dynamic scenes and perform common computer vision tasks for spiking camera. In this paper, we propose a self-supervised joint learning framework for optical flow estimation and reconstruction of spiking camera. The framework reconstructs clean frame-based spiking representations in a self-supervised manner, and then uses them to train the optical flow networks. We also propose an optical flow based inverse rendering process to achieve self-supervision by minimizing the difference with respect to the original spiking temporal aggregation image. The experimental results demonstrate that our method bridges the gap between synthetic and real-world scenes and achieves desired results in real-world scenarios. To the best of our knowledge, this is the first attempt to jointly reconstruct dynamic scenes and estimate optical flow for spiking camera from a self-supervised learning perspective. Shiyan Chen, Zhaofei Yu, Tiejun Huang 0001 |
AAAI | 1 |
| 2023 | Multi-View Super Resolution for Underwater Images Utilizing Atmospheric Light Scattering ModelabstractThe underwater environment is complex and the underwater light propagation undergoes absorption, scattering and reflection. This leads to the fact that the underwater light imaging cannot be generalized from land-based. How to use these imaging features to work better with super-resolution tasks for underwater imagery applications is still rarely studied. In this paper, we introduce the medium transmission (MT) maps to advance super-resolution tasks for underwater images. A multi-view network is designed to fuse information from the original underwater images and the MT maps, which provides information on the underlying physical properties of the water, such as the attenuation coefficients in different parts of water. By integrating information from multiple views, the proposed network can capture more of the underlying structure and features of the scene, leading to higher-quality super-resolved images. Besides, a new loss function, namely MT Loss, is developed according to the lack of details in special region of the underwater images. This loss function emphasizes the regions with less influence from the underwater environment during the underwater imaging process and therefore the network outputs a more detailed image. Finally, we compare our algorithm with state-of-the-art methods, and extensive results show that our network achieves better qualitative and quantitative performance. Jin Hao, Wenli Duan, Guangfei Li, Shiyan Chen, Wenhui Wu 0001, Hua Li 0012 |
ICPADS | 4 |
| 2023 | Enhancing Motion Deblurring in High-Speed Scenes with Spike StreamsabstractTraditional cameras produce desirable vision results but struggle with motion blur in high-speed scenes due to long exposure windows. Existing frame-based deblurring algorithms face challenges in extracting useful motion cues from severely blurred images. Recently, an emerging bio-inspired vision sensor known as the spike camera has achieved an extremely high frame rate while preserving rich spatial details, owing to its novel sampling mechanism. However, typical binary spike streams are relatively low-resolution, degraded image signals devoid of color information, making them unfriendly to human vision. In this paper, we propose a novel approach that integrates the two modalities from two branches, leveraging spike streams as auxiliary visual cues for guiding deblurring in high-speed motion scenes.
We propose the first spike-based motion deblurring model with bidirectional information complementarity. We introduce a content-aware motion magnitude attention module that utilizes learnable mask to extract relevant information from blurry images effectively, and we incorporate a transposed cross-attention fusion module to efficiently combine features from both spike data and blurry RGB images.
Furthermore, we build two extensive synthesized datasets for training and validation purposes, encompassing high-temporal-resolution spikes, blurry images, and corresponding sharp images. The experimental results demonstrate that our method effectively recovers clear RGB images from highly blurry scenes and outperforms state-of-the-art deblurring algorithms in multiple settings. Shiyan Chen, Jiyuan Zhang 0005, Yajing Zheng, Tiejun Huang 0001, Zhaofei Yu |
NeurIPS | 1 |
| 2022 | Self-Supervised Mutual Learning for Dynamic Scene Reconstruction of Spiking CameraabstractMimicking the sampling mechanism of the primate fovea, a retina-inspired vision sensor named spiking camera has been developed, which has shown great potential for capturing high-speed dynamic scenes with a sampling rate of 40,000 Hz. Unlike conventional digital cameras, the spiking camera continuously captures photons and outputs asynchronous binary spikes with various inter-spike intervals to record dynamic scenes. However, how to reconstruct dynamic scenes from asynchronous spike streams remains challenging. In this work, we propose a novel pretext task to build a self-supervised reconstruction framework for spiking cameras. Specifically, we utilize the blind-spot network commonly used in self-supervised denoising tasks as our backbone, and perform self-supervised learning by constructing proper pseudo-labels. In addition, in view of the poor scalability and insufficient information utilization of the blind-spot network, we present a mutual learning framework to improve the overall performance of the network through mutual distillation between a non-blind-spot network and a blind-spot network. This also enables the network to bypass constraints of the blind-spot network, allowing state-of-the-art modules to be used to further improve performance. The experimental results demonstrate that our methods evidently outperform previous unsupervised spiking camera reconstruction methods and achieve desirable results compared with supervised methods. Shiyan Chen, Chaoteng Duan, Zhaofei Yu, Ruiqin Xiong, Tiejun Huang 0001 |
IJCAI | 1 |
| 2022 | Temporal Effective Batch Normalization in Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are promising in neuromorphic hardware owing to utilizing spatio-temporal information and sparse event-driven signal processing. However, it is challenging to train SNNs due to the non-differentiable nature of the binary firing function. The surrogate gradients alleviate the training problem and make SNNs obtain comparable performance as Artificial Neural Networks (ANNs) with the same structure. Unfortunately, batch normalization, contributing to the success of ANNs, does not play a prominent role in SNNs because of the additional temporal dimension. To this end, we propose an effective normalization method called temporal effective batch normalization (TEBN). By rescaling the presynaptic inputs with different weights at every time-step, temporal distributions become smoother and uniform. Theoretical analysis shows that TEBN can be viewed as a smoother of SNN's optimization landscape and could help stabilize the gradient norm. Experimental results on both static and neuromorphic datasets show that SNNs with TEBN outperform the state-of-the-art accuracy with fewer time-steps, and achieve better robustness to hyper-parameters than other normalizations. Chaoteng Duan, Jianhao Ding, Shiyan Chen, Zhaofei Yu, Tiejun Huang 0001 |
NeurIPS | 3 |
| 2020 | A benchmark for clothes variation in person re-identificationabstractPerson re-identification (re-ID) has drawn attention significantly in the computer vision society due to its application and research significance. It aims to retrieve a person of interest across different camera views. However, there are still several factors that hinder the applications of person re-ID. In fact, most common data sets either assume that pedestrians do not change their clothing across different camera views or are taken under constrained environments. Those constraints simplify the person re-ID task and contribute to early development of person re-ID, yet a person has a great possibility to change clothes in real life. To facilitate the research toward conquering those issues, this paper mainly introduces a new benchmark data set for person re-identification. To the best of our knowledge, this data set is currently the most diverse for person re-identification. It contains 107 persons with 9,738 images, captured in 15 indoor/outdoor scenes from September 2019 to December 2019, varying according to viewpoints, lighting, resolutions, human pose, seasons, backgrounds, and clothes especially. We hope that this benchmark data set will encourage further research on person re-identification with clothes variation. Moreover, we also perform extensive analyses on this data set using several state-of-the-art methods. Our dataset is available at https://github.com/nkicsl/NKUP-dataset. Kai Wang 0001, Shiyan Chen, Jinni Yang, Keke Zhou, Tao Li 0022 |
Int. J. Intell. Syst. | 3 |
| 2018 | Storage-Aware Network Stack for NVM-Assisted Key-Value StoreabstractThis paper describes the design of a new software zero-copy network framework for NVM-assisted key-value stores, which directly stores and persists transactions from network into raw non-volatile memory used as write-ahead cache for data consistency. NVM is fast and bit-addressable which makes it the perfect choice for transient transaction log persistency than hard disks or even Flash drives, but its limited write cycle requires wear-leveling during direct access. However, popular RDMA-based zero-copy transmission normally needs to have the remote memory address beforehand and cannot cope with the address changing caused by wear-leveling easily. The software zero-copy solution proposed in this paper is designed with the awareness of NVM wear-leveling and log metadata management. Simulation results show that the new network framework improves performance by over 200× in throughput and decreases latency by more than 20× comparing to the traditional socket and hard disk based solution. When both equipped with NVM, the zero-copy network stack improves performance by 18 to 62% in throughput and 40 to 81% in latency comparing to the standard socket and with the lowest CPU consumption. Shiyan Chen, Dagang Li 0001, Wenbing Han, Deze Zeng |
ICCCN | 1 |
| 2018 | A novel non-volatile memory storage system for I/O-intensive applicationsabstractThe emerging memory technologies, such as phase change memory (PCM), provide chances for highperformance storage of I/O-intensive applications. However, traditional software stack and hardware architecture need to be optimized to enhance I/O efficiency. In addition, narrowing the distance between computation and storage reduces the number of I/O requests and has become a popular research direction. This paper presents a novel PCMbased storage system. It consists of the in-storage processing enabled file system (ISPFS) and the configurable parallel computation fabric in storage, which is called an in-storage processing (ISP) engine. On one hand, ISPFS takes full advantage of non-volatile memory (NVM)’s characteristics, and reduces software overhead and data copies to provide low-latency high-performance random access. On the other hand, ISPFS passes ISP instructions through a command file and invokes the ISP engine to deal with I/O-intensive tasks. Extensive experiments are performed on the prototype system. The results indicate that ISPFS achieves 2 to 10 times throughput compared to EXT4. Our ISP solution also reduces the number of I/O requests by 97% and is 19 times more efficient than software implementation for I/O-intensive applications. Wenbing Han, Shunfen Li, Gezi Li, Zhitang Song, Dagang Li 0001, Shiyan Chen |
Frontiers Inf. Technol. Electron. Eng. | 7 |