Yizhou Jiang

dblp:201/8247 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 42% Efficient and distributed learning · 14% Deep learning architectures and training · 14%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 68% Hardware accelerators and domain-specific architectures · 25% Embedded and real-time systems · 7%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
neuromorphic computing
1.622025
Adaptive Fission: Post-training Encoding for Low-latency Spike Neural Networks · NeurIPS 2025
Spatio-Temporal Approximation: A Training-Free SNN Conversion for Transformers · ICLR 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
spiking neural network accelerator
0.912025
Adaptive Fission: Post-training Encoding for Low-latency Spike Neural Networks · NeurIPS 2025
Machine learning › Deep learning architectures and training › spiking neural network
ANN-to-SNN conversion
0.812024
Spatio-Temporal Approximation: A Training-Free SNN Conversion for Transformers · ICLR 2024
Machine learning › Trustworthy machine learning › robustness › noisy data
feature noise
0.812024
Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to Generalize · ICML 2024
Machine learning › Efficient and distributed learning › model deployment
model conversion
0.812024
Spatio-Temporal Approximation: A Training-Free SNN Conversion for Transformers · ICLR 2024
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.812024
Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to Generalize · ICML 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to Generalize · ICML 2024
Emerging computing paradigms › neuromorphic computing
spiking neural network
0.812024
Spatio-Temporal Approximation: A Training-Free SNN Conversion for Transformers · ICLR 2024
Machine learning › Learning paradigms
multi-task learning
0.412020
Task Understanding from Confusing Multi-task Data · ICML 2020
Computer vision › Video understanding and tracking › activity recognition
task recognition
0.412020
Task Understanding from Confusing Multi-task Data · ICML 2020
Embedded and real-time systems
low-latency inference
0.312025
Adaptive Fission: Post-training Encoding for Low-latency Spike Neural Networks · NeurIPS 2025
Machine learning › Learning theory
inductive bias
0.212024
Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to Generalize · ICML 2024

Methods — techniques the papers use, named apart from their topics

temporal approximation · 1.5spatial approximation · 1.5self-attention · 1.5post-training encoding · 0.9population coding · 0.9two-layer ReLU network · 0.8teacher-student network · 0.8stochastic gradient descent · 0.8iterative algorithm · 0.4confusing supervised learning · 0.4
YearPublicationVenuePosition
2025 An Adaptive Watermark Embedding Method for Multi-modal Data in Power Systems
Yizhen Sun, Yizhou Jiang, Wen-Xiao Zhao, Jingyuan Xue, Shigeng Zhang
ICA3PP (8)2
2025 Adaptive Fission: Post-training Encoding for Low-latency Spike Neural Networks
abstract
Spiking Neural Networks (SNNs) often rely on rate coding, where high-precision inference depends on long time-steps, leading to significant latency and energy cost—especially for ANN-to-SNN conversions. To address this, we propose Adaptive Fission, a post-training encoding technique that selectively splits high-sensitivity neurons into groups with varying scales and weights. This enables neuron-specific, on-demand precision and threshold allocation while introducing minimal spatial overhead. As a generalized form of population coding, it seamlessly applies to a wide range of pretrained SNN architectures without requiring additional training or fine-tuning. Experiments on neuromorphic hardware demonstrate up to 80\% reductions in latency and power consumption without degrading accuracy.
Yizhou Jiang, Feng Chen 0007, Yuqian Liu, Haichuan Gao
NeurIPS1
2025 Causal dreamer for partially observable model-based reinforcement learning
Haichuan Gao, Tianrun Xu, Chujie Zhao, Jinsheng Ren, Yizhou Jiang, Shangqi Guo, Feng Chen 0007
Neurocomputing7
2025 Iterative compression towards in-distribution features in domain generalization
Yizhou Jiang, Feng Chen 0007
Neurocomputing1
2024 Spatio-Temporal Approximation: A Training-Free SNN Conversion for Transformers
abstract
Spiking neural networks (SNNs) are energy-efficient and hold great potential for large-scale inference. Since training SNNs from scratch is costly and has limited performance, converting pretrained artificial neural networks (ANNs) to SNNs is an attractive approach that retains robust performance without additional training data and resources. However, while existing conversion methods work well on convolution networks, emerging Transformer models introduce unique mechanisms like self-attention and test-time normalization, leading to non-causal non-linear interactions unachievable by current SNNs. To address this, we approximate these operations in both temporal and spatial dimensions, thereby providing the first SNN conversion pipeline for Transformers. We propose \textit{Universal Group Operators} to approximate non-linear operations spatially and a \textit{Temporal-Corrective Self-Attention Layer} that approximates spike multiplications at inference through an estimation-correction approach. Our algorithm is implemented on a pretrained ViT-B/32 from CLIP, inheriting its zero-shot classification capabilities, while improving control over conversion losses. To our knowledge, this is the first direct training-free conversion of a pretrained Transformer to a purely event-driven SNN, promising for neuromorphic hardware deployment.
Yizhou Jiang, Kunlin Hu, Haichuan Gao, Yuqian Liu, Feng Chen 0007
ICLR1
2024 Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to Generalize
abstract
Learning representations that generalize under distribution shifts is critical for building robust machine learning models. However, despite significant efforts in recent years, algorithmic advances in this direction have been limited. In this work, we seek to understand the fundamental difficulty of out-of-distribution generalization with deep neural networks. We first empirically show that perhaps surprisingly, even allowing a neural network to explicitly fit the representations obtained from a teacher network that can generalize out-of-distribution is insufficient for the generalization of the student network. Then, by a theoretical study of two-layer ReLU networks optimized by stochastic gradient descent (SGD) under a structured feature model, we identify a fundamental yet unexplored feature learning proclivity of neural networks, feature contamination: neural networks can learn uncorrelated features together with predictive features, resulting in generalization failure under distribution shifts. Notably, this mechanism essentially differs from the prevailing narrative in the literature that attributes the generalization failure to spurious correlations. Overall, our results offer new insights into the non-linear feature learning dynamics of neural networks and highlight the necessity of considering inductive biases in out-of-distribution generalization.
Chujie Zhao, Yizhou Jiang, Feng Chen 0007
ICML4
2023 IEMS: An IoT-Empowered Wearable Multimodal Monitoring System in Neurocritical Care
abstract
IoT-empowered wearable multimodal monitoring system (IEMS), an IEMS for neurocritical care is developed to perform simultaneous monitoring of 8-channel electroencephalogram (EEG), 2-channel regional cerebral oxygen saturation (rSO2) based on near-infrared spectrum (NIRS), body surface temperature, electrocardiogram (ECG), photoplethysmography (PPG), and bioimpedance (Bio-Z). The IoT platform and wireless devices enable the patients’ signals available for remote diagnosis. Besides, analysis functions and artificial intelligence (AI) algorithms could be embedded in both the bedside platform and the cloud server to support physicians with clinical decisions. In the multimodal neural monitoring device, the following designs are adopted to face the neurological intensive care unit (NICU) application. Active electrodes and preamplifying free topology provide better signal quality and a larger dynamic range (DR). Dedicated low-power designs ensure the device lasts 10 h of operation. A nonwoven headset improves long-term wearability, which is also quick and easy to install. In the cardiovascular patch, the disposable electrode patch based on elastic materials ensures tight and comfortable contact with skin. Besides, the reusable wireless sensing module is tiny (20 mm$\times16$mm$\times9$mm) but could measure ECG, PPG, and Bio-Z simultaneously. Electrical tests and human subject (healthy volunteers and NICU patients) studies were conducted to examine the performance. The EEG channels show 130.75-dB DR and 0.84-$\mu {}\text{V}_{\mathrm{ RMS}}$input-referred noise, which also yields signals with high quality during human EEG monitoring. The NIRS channels exhibit good linearity and are able to operate under severe ambient interference. The temperature sensors show ±0.2 °C accuracy. Moreover, the system complies with mandatory standards for medical equipment. Overall, an IEMS can meet the requirements for NICU applications and could provide better comfort during long-term wearing.
Yizhou Jiang, Jianzheng Li, Cehui Tan, Chongyuan Ren, Jiuqing Feng, Yichen Cai 0003, Jianpeng Gao, Ye Gong, Yajie Qin
IEEE Internet Things J.1
2020 Task Understanding from Confusing Multi-task Data
abstract
Beyond machine learning’s success in the specific tasks, research for learning multiple tasks simultaneously is referred to as multi-task learning. However, existing multi-task learning needs manual definition of tasks and manual task annotation. A crucial problem for advanced intelligence is how to understand the human task concept using basic input-output pairs. Without task definition, samples from multiple tasks are mixed together and result in a confusing mapping challenge. We propose Confusing Supervised Learning (CSL) that takes these confusing samples and extracts task concepts by differentiating between these samples. We theoretically proved the feasibility of the CSL framework and designed an iterative algorithm to distinguish between tasks. The experiments demonstrate that our CSL methods could achieve a human-like task understanding without task labeling in multi-function regression problems and multi-task recognition problems.
Yizhou Jiang, Shangqi Guo, Feng Chen 0007
ICML2
2020 A 37.37μW-Per-Cell Multifunctional Automated Nanopore Sequencing CMOS Platform with 16∗8 Biosensor Array
abstract
Nanopore-based DNA sequencing technology has become one of the most promising sequencing approaches with its advantages of label-free and low cost. However, most of the biosensor systems for nanopore sequencing only perform passive detection which is merely part of the overall function of a practical DNA sequencing platform. In this paper, a multifunctional automated integrated CMOS platform for nanopore-based DNA sequencing is presented. The platform equipped with 16*8 biosensor array for nanopore detection is also able to perform bilayer lipid membrane capacitance detection and nanopore insertion pulse generation, realizing the whole process automated auxiliary function from transducer preparation to DNA sequencing. Post layout simulation shows that each cell consumes only 37.366μW while the whole system occupying 1.633mm2.
Chenjie Dong, Yizhou Jiang, Yumei Huang, Yajie Qin
ISCAS2