Jun Zhou 0014

dblp:99/3847-14 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-7899-2909ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical Notes
abstract
Effective clinical history taking is a foundational yet underexplored component of clinical reasoning. While large language models (LLMs) have shown promise on static benchmarks, they often fall short in dynamic, multi-turn diagnostic settings that require iterative questioning and hypothesis refinement. To address this gap, we propose Note2Chat, a note-driven framework that trains LLMs to conduct structured history taking and diagnosis by learning from widely available medical notes. Instead of relying on scarce and sensitive dialogue data, we convert real-world medical notes into high-quality doctor-patient dialogues using a decision tree-guided generation and refinement pipeline. We then propose a three-stage fine-tuning strategy combining supervised learning, simulated data augmentation, and preference learning. Furthermore, we propose a novel single-turn reasoning paradigm that reframes history taking as a sequence of single-turn reasoning problems. This design enhances interpretability and enables local supervision, dynamic adaptation, and greater sample efficiency. Experimental results show that our method substantially improves clinical reasoning, achieving gains of +16.9 F1 and +21.0 Top-1 diagnostic accuracy over GPT-4o.
Yang Zhou 0017, Zhenting Sheng, Mingrui Tan, Yuting Song, Jun Zhou 0014, Yu Heng Kwan, Lian Leng Low, Yang Bai 0011, Yong Liu 0026
AAAI5
2025 AdvMIM: Adversarial Masked Image Modeling for Semi-supervised Medical Image Segmentation
Lei Zhu 0003, Jun Zhou 0014, Rick Siow Mong Goh, Yong Liu 0026
MICCAI (16)2
2025 Low Latency Conversion of Artificial Neural Network Models to Rate-Encoded Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are well suited for resource-constrained applications as they do not need expensive multipliers. In a typical rate-encoded SNN, a series of binary spikes within a globally fixed time window is used to fire the neurons. The time window size is also the latency of the network in performing a single inference, as well as determining the overall energy efficiency of the model. The aim of this article is to reduce this while maintaining accuracy when converting artificial neural networks (ANNs) to their equivalent SNNs. The state-of-the-art conversion schemes yield SNNs with accuracies comparable with ANNs only for large window sizes. In this article, we start with understanding the information loss when converting from preexisting ANN models to standard rate-encoded SNN models. From these insights, we propose a suite of techniques that includes a novel SNN encoding scheme, a new spike generation model, an input channel expansion strategy, and a threshold training technique. Together, these methods enabled us to achieve state-of-the-art accuracies using the lowest latencies reported in the literature. In particular, our method achieved a top-1 SNN accuracy of 98.73% (using a single time step) on the MNIST dataset, 76.38% (with eight time steps) on the CIFAR-100 dataset, and 93.71% (eight time steps) on the CIFAR-10 dataset. On ImageNet, an SNN accuracy of 81.9% was achieved using 40 time steps.
Zhanglu Yan, Kaiwen Tang, Jun Zhou 0014, Weng-Fai Wong
IEEE Trans. Neural Networks Learn. Syst.3
2024 1.63 pJ/SOP Neuromorphic Processor With Integrated Partial Sum Routers for In-Network Computing
abstract
Neuromorphic computing is promising to achieve unprecedented energy efficiency by emulating the human brain’s mechanism. Conventional neuromorphic accelerators employ split-and-merge method to map spiking neural networks’ inputs to surpass the fan-in capabilities of a single neuron core. However, this approach gives rise to the risk of accuracy compromise and extra core usage for the merging process. Moreover, it requires excessive data movement and clock cycles to aggregate spikes generated by partial sums instead of total sums obtained from different cores with substantial power and energy overhead. This work presents a novel approach to addressing the challenges imposed by the split-and-merge method. We propose an energy-efficient, reconfigurable neuromorphic processor that leverages several key techniques to mitigate the above issues. First, we introduce a partial sum router circuitry that enables in-network computing (INC), eliminating the need for extra merge cores. Second, we adopt software-defined Networks-on-Chip (NoCs) by leveraging predefined, efficient routing, eliminating power-hungry routing computation. At last, we incorporate fine-grained power gating and clock gating techniques for further power reduction. Experimental results from our test chip demonstrate the lossless mapping of the algorithm and exceptional energy efficiency, achieving an energy consumption of 1.63 pJ/SOP at 0.48 V. This energy efficiency represents a 22.4% improvement compared to the state-of-the-art results. Our proposed neuromorphic processor provides an efficient and flexible solution for neural network processing, mitigating the limitations of the traditional split-and-merge approach while delivering superior energy efficiency.
Dongrui Li, Ming Ming Wong, Yi Sheng Chong, Jun Zhou 0014, Mohit Upadhyay, Ananta Narayanan Balaji, Aarthy Mani, Weng-Fai Wong, Li-Shiuan Peh, Anh-Tuan Do, Bo Wang 0020
IEEE Trans. Very Large Scale Integr. Syst.4
2023 1.7pJ/SOP Neuromorphic Processor with Integrated Partial Sum Routers for In-Network Computing
abstract
Conventional neuromorphic accelerators primarily leverage split-merge method to accommodate a neural network that is beyond a single core's size, leading to possible accuracy loss, extra core usage and significant power and energy overhead. This work presents an energy-efficient, reconfigurable neuro-morphic processor to address the problem by (i) a partial sum router circuitry that enables in-network computing to remove the need of extra merge cores; (ii) software-defined Networks-on-Chip that eliminates the power-hungry routing compute and (iii) fine-grained power gating and clock gating technique for power reduction. Our test chip achieves lossless mapping as the algorithm and an energy efficiency of 1.7pJ/SOP at 0.5V, 19% lower than state-of-the-art result.
Bo Wang 0020, Ming Ming Wong, Dongrui Li, Yi Sheng Chong, Jun Zhou 0014, Weng-Fai Wong, Li-Shiuan Peh, Aarthy Mani, Mohit Upadhyay, Ananta Narayanan Balaji, Anh-Tuan Do
ISCAS5
2023 CQ$^{+}$+ Training: Minimizing Accuracy Loss in Conversion From Convolutional Neural Networks to Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are attractive for energy-constrained use-cases due to their binarized activation, eliminating the need for weight multiplication. However, its lag in accuracy compared to traditional convolutional network networks (CNNs) has limited its deployment. In this paper, we propose CQ+ training (extended "clamped" and "quantized" training), an SNN-compatible CNN training algorithm that achieves state-of-the-art accuracy for both CIFAR-10 and CIFAR-100 datasets. Using a 7-layer modified VGG model (VGG-*), we achieved 95.06% accuracy on the CIFAR-10 dataset for equivalent SNNs. The accuracy drop from converting the CNN solution to an SNN is only 0.09% when using a time step of 600. To reduce the latency, we propose a parameterized input encoding method and a threshold training method, which further reduces the time window size to 64 while still achieving an accuracy of 94.09%. For the CIFAR-100 dataset, we achieved an accuracy of 77.27% using the same VGG-* structure and a time window of 500. We also demonstrate the transformation of popular CNNs, including ResNet (basic, bottleneck, and shortcut block), MobileNet v1/2, and Densenet, to SNNs with near-zero conversion accuracy loss and a time window size smaller than 60. The framework was developed in PyTorch and is publicly available.
Zhanglu Yan, Jun Zhou 0014, Weng-Fai Wong
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Benchmarking Emergency Department Triage Prediction Models with Machine Learning and Large Public Electronic Health Records
Feng Xie 0004, Jun Zhou 0014, Jin Wee Lee, Mingrui Tan, Siqi Li 0004, Logasan S/O Rajnthern, Marcel Lucas Chee, Bibhas Chakraborty, An-Kwok Ian Wong, Alon Dagan, Marcus Eng Hock Ong, Nan Liu 0003
AMIA2
2022 REACT: a heterogeneous reconfigurable neural network accelerator with software-configurable NoCs for training and inference on wearables
abstract
On-chip training improves model accuracy on personalised user data and preserves privacy. This work proposes REACT, an AI accelerator for wearables that has heterogeneous cores supporting both training and inference. REACT's architecture is NoC-centric, with weights, features and gradients distributed across cores, accessed and computed efficiently through software-configurable NoCs. Unlike conventional dynamic NoCs, REACT's NoCs have no buffer queues, flow control or routing, as they are entirely configured by software for each neural network. REACT's online learning realises upto 75% accuracy improvement, and is upto 25× faster and 520× more energy-efficient than state-of-the-art accelerators with similar memory and computation footprint.
Mohit Upadhyay, Rohan Juneja, Bo Wang 0020, Jun Zhou 0014, Weng-Fai Wong, Li-Shiuan Peh
DAC4
2022 Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong
Neurocomputing7
2022 Corrigendum to "Coreset: Hierarchical neuromorphic computing supporting large-scale neural networks with improved resource efficiency" [Neurocomputing (2022) 128-140]
Huaipeng Zhang, Tao Luo 0014, Chuping Qu, Myat Thu Linn Aung, Yingnan Cui, Jun Zhou 0014, Ming Ming Wong, Junran Pu, Anh-Tuan Do, Rick Siow Mong Goh, Weng-Fai Wong
Neurocomputing7
2021 Near Lossless Transfer Learning for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) significantly reduce energy consumption by replacing weight multiplications with additions. This makes SNNs suitable for energy-constrained platforms. However, due to its discrete activation, training of SNNs remains a challenge. A popular approach is to first train an equivalent CNN using traditional backpropagation, and then transfer the weights to the intended SNN. Unfortunately, this often results in significant accuracy loss, especially in deeper networks. In this paper, we propose CQ training (Clamped and Quantized training), an SNN-compatible CNN training algorithm with clamp and quantization that achieves near-zero conversion accuracy loss. Essentially, CNN training in CQ training accounts for certain SNN characteristics. Using a 7 layer VGG-* and a 21 layer VGG-19, running on the CIFAR-10 dataset, we achieved 94.16% and 93.44% accuracy in the respective equivalent SNNs. It outperforms other existing comparable works that we know of. We also demonstrate the low-precision weight compatibility for the VGG-19 structure. Without retraining, an accuracy of 93.43% and 92.82% using quantized 9-bit and 8-bit weights, respectively, was achieved. The framework was developed in PyTorch and is publicly available.
Zhanglu Yan, Jun Zhou 0014, Weng-Fai Wong
AAAI2
2021 An optimal global algorithm for route guidance in advanced traveler information systems
Bokui Chen, Zhong-Jun Ding, Jun Zhou 0014, Yongquan Chen
Inf. Sci.4
2021 Synthesis of the Dynamical Properties of Feedback Loops in Bio-Pathways
abstract
Feedback loops regulate various biological functions such as oscillations, bistability, and robustness. They play a significant role in developmental signalling and failure of feedback can lead to disease. Systematic analysis of feedback loops could be useful in understanding their properties and biological effects. We propose here a method to automatically analyze feedback loops in bio-pathways and synthesize temporal logic properties which describe their dynamics. Starting with an ordinary differential equations (ODEs) based model of a bio-pathway, for a chosen feedback loop present in the pathway, we use a convolutional neural network to classify the behaviours of the key components of the feedback according to templates specified in bounded linear temporal logic (BLTL). Once a template has been identified, we instantiate the symbolic variables appearing in the template and synthesize properties using a parameter estimation procedure based on sequential hypothesis testing. We have applied this framework to a number of bio-pathway models and validated that the synthesized properties faithfully describe the behaviours of the feedback loops.
Jun Zhou 0014, R. Ramanathan 0002, Weng-Fai Wong
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 Shenjing: A low power reconfigurable neuromorphic accelerator with partial-sum and spike networks-on-chip
abstract
The next wave of on-device AI will likely require energy-efficient deep neural networks. Brain-inspired spiking neural networks (SNN) has been identified to be a promising candidate. Doing away with the need for multipliers significantly reduces energy. For on-device applications, besides computation, communication also incurs a significant amount of energy and time. In this paper, we propose Shenjing, a configurable SNN architecture which fully exposes all on-chip communications to software, enabling software mapping of SNN models with high accuracy at low power. Unlike prior SNN architectures like TrueNorth, Shenjing does not require any model modification and retraining for the mapping. We show that conventional artificial neural networks (ANN) such as multilayer perceptron, convolutional neural networks, as well as the latest residual neural networks can be mapped successfully onto Shenjing, realizing ANNs with SNN’s energy efficiency. For the MNIST inference problem using a multilayer perceptron, we were able to achieve an accuracy of 96% while consuming just 1.26 mW using 10 Shenjing cores.
Bo Wang 0020, Jun Zhou 0014, Weng-Fai Wong, Li-Shiuan Peh
DATE2
2020 A future intelligent traffic system with mixed autonomous vehicles and human-driven vehicles
Bokui Chen, Duo Sun, Jun Zhou 0014, Weng-Fai Wong, Zhong-Jun Ding
Inf. Sci.3
2019 Resource Efficient Personalized ECG Beat Classification via Temporal Logic Synthesis
abstract
According to the World Health Organization, cardio-vascular diseases accounts for 31% of all deaths worldwide in 2016. Detecting the onset of heart irregularities can potentially save many lives. The ubiquity of wearable devices opens up the possibility of having a heart disease detection at everyone's disposal. To enable this, low energy ECG classification is needed. Unlike previous methods of using signal processing or even deep learning networks, this paper is the first to propose a low cost means to detect abnormal ECG beat signals by first synthesizing temporal logic formulas from training signals, and then checking if the synthesized formulas by the input signal at runtime. Our results show that the method has a high accuracy in detecting abnormal ECG beats while requiring significantly lower computation resource. Compared to It takes only a state-of-the-art convolutional neural network approach, our method achieves a comparable accuracy but with 0.3% of memory, and millions of computation operations, hence energy, saved.
Jun Zhou 0014, Weng-Fai Wong
BIBE1
2019 A System-Level Simulator for RRAM-Based Neuromorphic Computing Chips
abstract
Advances in non-volatile resistive switching random access memory (RRAM) have made it a promising memory technology with potential applications in low-power and embedded in-memory computing devices owing to a number of advantages such as low-energy consumption, low area cost and good scaling. There have been proposals to employ RRAM in architecting chips for neuromorphic computing and artificial neural networks where matrix-vector multiplication can be computed in the analog domain in a single timestep. However, it is challenging to employ RRAM devices in neuromorphic chips owing to the non-ideal behavior of RRAM. In this article, we propose a cycle-accurate and scalable system-level simulator that can be used to study the effects of using RRAM devices in neuromorphic computing chips. The simulator models a spatial neuromorphic chip architecture containing many neural cores with RRAM crossbars connected via a Network-on-Chip (NoC). We focus on system-level simulation and demonstrate the effectiveness of our simulator in understanding how non-linear RRAM effects such as stuck-at-faults (SAFs), write variability, and random telegraph noise (RTN) can impact an application’s behavior. By using our simulator, we show that RTN and write variability can have adverse effects on an application. Nevertheless, we show that these effects can be mitigated through proper design choices and the implementation of a write-verify scheme.
Matthew Kay Fei Lee, Yingnan Cui, Thannirmalai Somu, Tao Luo 0014, Jun Zhou 0014, Wai Teng Tang, Weng-Fai Wong, Rick Siow Mong Goh
ACM Trans. Archit. Code Optim.5
2019 Fault Tolerant Stencil Computation on Cloud-Based GPU Spot Instances
abstract
This paper describes a fault tolerant framework for distributed stencil computation on cloud-based GPU clusters. It uses pipelining to overlap the data movement with computation in the halo region as well as parallelises data movement within the GPUs. Instead of running stencil codes on traditional clusters and supercomputers, the computation is performed on the Amazon Web Service GPU cloud, and utilizes its spot instances to improve cost-efficiency. The implementation is based on a low-cost fault-tolerant mechanism to handle the possible termination of the spot instances. Coupled with a price bidding module, our stencil framework not only optimizes for performance but also for cost. Experimental results show that our framework outperforms the state-of-the-art solutions achieving a peak of 25 TFLOPS for 2-D decomposition running on 512 nodes. We also show that the use of spot instances yields good cost-efficiency, increasing the average TFLOPS/USD from 132 to 360.
Jun Zhou 0014, Yan Zhang 0027, Weng-Fai Wong
IEEE Trans. Cloud Comput.1