Na Gong

dblp:60/1284 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-3297-7436ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Hardware-Early-in-the-Loop for Edge AI
abstract
Designing edge AI systems for resource-constrained hardware requires balancing accuracy, latency, and power simultaneously. These three objectives conflict at every stage of development. Current practice follows an algorithm-first sequence: a model is selected and trained first, and hardware is chosen afterward to accommodate it. This approach defers hardware constraints to the end of the design process, where they are most expensive to resolve and hardest to justify to the wider project team.
Isaac Arnold, Na Gong
ACM Great Lakes Symposium on VLSI2
2026 Portable Breath Acetone Sensing System with Embedded Machine Learning for Non-Invasive Diabetes Monitoring
abstract
Breath acetone is a promising non-invasive biomarker for diabetes monitoring. This work presents a portable breath acetone sensing system based on a MgWO chemiresistive gas sensor integrated with a high-impedance measurement circuit and embedded machine learning. To accurately measure the high resistance of the sensor, a voltage divider with buffer amplification and adaptive resistor network is implemented, enabling reliable acquisition over a wide dynamic range. To address the nonlinear response of metal oxide sensors, a machine learning approach using Support Vector Machine (SVM) is employed for acetone concentration prediction. Experimental validation is conducted using controlled acetone concentrations in the range of 0.5–1.5 ppm. The proposed system achieves a high prediction accuracy with an R² value of ∼0.98, demonstrating strong agreement with reference measurements. The results highlight the potential of the developed device for real-time, low-cost, and non-invasive diabetes monitoring applications.
Md Hasib Fakir, Md Fuyad Al Masud, Na Gong, Danling Wang
ACM Great Lakes Symposium on VLSI3
2026 Adaptive Post-Decoder Memory for Low-Power 360° Video
abstract
360° video processing is rapidly gaining popularity in applications such as entertainment, medical visualization, and immersive training. However, compared to conventional video, 360° video requires significantly higher resources, including bandwidth, power, and memory capacity, due to the large volume of panoramic data. Importantly, not all regions of a 360° frame carry equal visual importance. In this work, we propose an adaptive post-decoder memory architecture that dynamically optimizes video data based on the region of interest. The decoded frame is divided into three regions: Field-of-View (FoV), border, and background. While FoV data is preserved without modification, border and background regions are compressed using flexible bit truncation during memory read operations. This approach reduces redundant data storage and memory bandwidth requirements. Experimental results demonstrate approximately 50% frame size reduction while maintaining about 42 dB WS-PSNR, achieving nearly 19.5% read power savings compared to conventional 360° video processing systems.
Md Humaun Kabir, Ricardo Mulino, Ali Ahmad Haidous, Soumya Swaraj Mondal, Dakota Hudson, Md. Sajjad Hossain, Na Gong, Hritom Das
ACM Great Lakes Symposium on VLSI7
2026 Reliability-Driven Sneak Path Current Modeling and Optimization for Passive Memristor Crossbar Arrays
Zhenlin Pei, Shah Zayed Riam, Kyle Mooney, Chenyun Pan, Na Gong
ACM Great Lakes Symposium on VLSI5
2026 Breaking Silos: Integrating Computational Thinking Across Elementary Subjects for Future Educators
abstract
This study examined the role of preservice teachers' (PSTs') teaching self-efficacy in a training program, which focused on integrating computational thinking (CT) in elementary subjects. Results showed that PSTs' CT knowledge significantly predicted their teaching self-efficacy though the explanatory power was weak. Evidence from the qualitative analysis on PSTs' post-lesson reflections showed different patterns between PSTs with positive and non-positive teaching self-efficacy. These results not only provided insights into future research but also highlighted the need to develop PSTs' pedagogical skills alongside their CT knowledge during the training.
Shenghua Zha, Lauren Brannan, Na Gong, Kelly Byrd, Drew Gossen, Todd Johnson, Jennifer Simpson, Karen Morrison
SIGCSE (2)3
2025 Low-Cost Wearable Edge-AI Device for Diabetes Management
Luke Young, Danling Wang, Na Gong
ACM Great Lakes Symposium on VLSI3
2025 Predicting Students' Interest from Small Group Conversational Characteristics: Insights from an AI Literacy Education with High School Students
abstract
Recent years have seen developments in AI instructional practices for K-12 students. In literature, students' interest in AI is shown to correlate with gaining AI knowledge; however, little is known about how AI interest manifests in classroom discourses during AI literacy lessons. This study examined students' participation in an integrated AI curriculum delivered to a cognitive science class in a high school in the southern US. Students worked in small groups and built a supervised machine learning model to recognize kids' drawings at different stages of artistic development. Our analysis showed that semantic features extracted from students' small group conversations significantly predicted their interest in learning AI. However, we found no significant relationship between students' social construction of knowledge and their interests. This study sheds light on the relationship between the learning process and interest; when further developed, this analysis may be developed into a classroom activity analytics tool that may provide real-time feedback to teachers engaged in AI literacy education to enhance teaching effectiveness in this nascent content area.
Shenghua Zha, Lujie Karen Chen, Woei Hung, Na Gong, Pamela Moore, Bethany Klemetsrud
SIGCSE (2)4
2024 Two Birds With One Stone: Differential Privacy by Low-Power SRAM Memory
abstract
The software-based implementation of differential privacy mechanisms has been shown to be neither friendly for lightweight devices nor secure against side-channel attacks. In this work, we aim to develop a hardware-based technique to achieve differential privacy by design. In contrary to the conventional software-based noise generation and injection process, our design realizes local differential privacy (LDP) by harnessing the inherent hardware noise into controlled LDP noise when data is stored in the memory. Specifically, the noise is tamed through a novel memory design and power downscaling technique, which leads to double-faceted gains in privacy and power efficiency. A well-round study that consists of theoretical design and analysis and chip implementation and experiments is presented. The results confirm that the developed technique is differentially private, saves 88.58% system power, speeds up software-based DP mechanisms by more than$10^{6}$times, while only incurring 2.46% chip overhead and 7.81% estimation errors in data recovery.
Jianqing Liu, Na Gong, Hritom Das
IEEE Trans. Dependable Secur. Comput.2
2023 LiteCCLKNet: A lightweight criss-cross large kernel convolutional neural network for hyperspectral image classification
abstract
Abstract High‐performance convolutional neural networks (CNNs) stack many convolutional layers to obtain powerful feature extraction capability, which leads to huge storing and computational costs. The authors focus on lightweight models for hyperspectral image (HSI) classification, so a novel lightweight criss‐cross large kernel convolutional neural network (LiteCCLKNet) is proposed. Specifically, a lightweight module containing two 1D convolutions with self‐attention mechanisms in orthogonal directions is presented. By setting large kernels within the 1D convolutional layers, the proposed module can efficiently aggregate long‐range contextual features. In addition, the authors effectively obtain a global receptive field by stacking only two of the proposed modules. Compared with traditional lightweight CNNs, LiteCCLKNet reduces the number of parameters for easy deployment to resource‐limited platforms. Experimental results on three HSI datasets demonstrate that the proposed LiteCCLKNet outperforms the previous lightweight CNNs and has higher storage efficiency.
Chengcheng Zhong, Na Gong, Zitong Zhang 0001, Kai Zhang 0060
IET Comput. Vis.2
2022 Classification of hyperspectral images via improved cycle-MLP
abstract
Abstract Pixel‐wise classification of hyperspectral image (HSI) is a hot spot in the field of remote sensing. The classification of HSI requires the model to be more sensitive to dense features, which is quite different from the modelling requirements of traditional image classification tasks. Cycle‐Multilayer Perceptron (MLP) has achieved satisfactory results in dense feature prediction tasks because it is an expert in extracting high‐resolution features. In order to obtain a more stable receptive field and enhance the effect of feature extraction in multiple directions, we propose an MLP‐like model called DriftNet for HSI classification inspired by Cycle‐MLP and deformable convolution. Besides, the proposed DriftNet uses a unique ladder‐like fully connected structure to achieve progressive activation of neurons and facilitates the fusion of spatial and spectral information, thereby obtaining more sensitive location information for better classification results. Experimental results on three public hyperspectral datasets demonstrate the effectiveness and generalisation of DriftNet.
Na Gong, Heng Zhou 0007, Kai Zhang 0060, Zhongyuan Wu, Xin Zhang 0115
IET Comput. Vis.1
2021 Application-Aware Quality-Energy Optimization: Mathematical Models Enabled Simultaneous Quality and Energy-Sensitive Optimal Memory Design
abstract
Energy efficiency is nowadays a well-known principal design goal across all layers of computing systems (e.g., sensors, mobile, cloud). With diminishing benefits from CMOS technology scaling and increasing demands from unprecedented data size, new memory hardware innovation is greatly needed to enhance energy efficiency of computing systems. Recently, quality-aware hardware design techniques have been developed from different stack layers (device/circuit/architecture/system) to enable near-threshold/sub-threshold voltage operation by trading off between quality and power efficiency. Specifically, based on the energy-quality trade-off and application requirement, memory hardware is designed for the work mode with maximum quality first (step1-design for the work mode) and then the supply voltage will be adjusted in the sleep mode with just-enough quality to achieve maximum efficiency (step2-adjust in the sleep mode). We first propose mathematical models for this two-step design method to avoid time-consuming and laborious ASIC design iterations in traditional hardware design process. However, such a two-step design method focuses more at the application quality than the energy efficiency in all cases. In addition, as the supply voltage in the second step depends on the design from the first step, the solution space of the second step may be greatly limited, far from the true minimum supply voltage. To handle these issues, we propose a new design concept, the simultaneous quality and energy-sensitive optimal design (SQEOD), in which the two objectives are considered simultaneously rather than by two separate steps. By introducing a system-wised importance weight parameter in the modeling process, our method demonstrates system-specific SQEOD mathematical models for different memory designs with various requirements on the application quality and/or energy efficiency. The results of the numerical studies on embedded memory design show that the proposed models provide a useful and fast tool to enable the optimal hardware designs.
Hritom Das, Na Gong
IEEE Trans. Sustain. Comput.3
2021 Flexible Low-Cost Power-Efficient Video Memory With ECC-Adaptation
abstract
In this article, a flexible power-efficient video memory is presented that can dynamically adjust the strength of error correction code (ECC), thereby enabling power-quality tradeoff based on application requirements. Specifically, we utilize the bit significance characteristics of video data to develop a low-cost parity storage scheme that supports both hamming code-74 (ECC74) and hamming code-1511 (ECC1511). Based on this, we propose a flexible memory with three dynamic power-quality adaptation schemes (i.e., ECC74, ECC1511, and no ECC) to meet different video application requirements. Our simulation results in 45-nm CMOS technology show that the proposed memory can enable up to 35.37% power savings without a noticeable degradation in video quality, as compared to the conventional design. We also design an integrated ECC encoder/decoder that handles both ECC74 and ECC1511, which reduces area overhead. To evaluate the effectiveness of the proposed technique, we further develop a system-level video storage embedded test platform based on a commercial 65-nm SRAM chip, which shows that the proposed technique results in significant supply voltage reduction without noticeable video quality degradation.
Hritom Das, Ali Ahmad Haidous, Scott C. Smith, Na Gong
IEEE Trans. Very Large Scale Integr. Syst.4
2020 On Mathematical Models of Optimal Video Memory Design
abstract
The big video data size today imposes huge pressure on storage. The variation and aging induced memory failures significantly influence the video output quality. Recently, researchers have developed different memory designs for videos, deep learning, and other data-intensive applications, which enables better energy-quality tradeoffs with design constraints. Unfortunately, designing memory has been proven to be a very challenging problem due to: 1) various design constraints; 2) multiple memory bitcell design options; and 3) challenging layout integration and cost analysis using different memory technologies. In this paper, we develop novel mathematical models for optimizing embedded video memory design without applying a time-consuming and laborious ASIC design process. The problems are formulated as nonlinear programs and integer linear programs. Different SRAM designs and hybrid SRAM and DRAM designs are considered in our models. The results of the numerical studies show that by applying our proposed method the average mean square error of the video storage can be greatly reduced, even by more than 90% in many cases.
Hritom Das, Yifu Gong, Na Gong
IEEE Trans. Circuits Syst. Video Technol.4
2020 Memory Optimization for Energy-Efficient Differentially Private Deep Learning
abstract
With the advent of Internet of Things (IoT) technologies and availability of a large amount of data, deep learning has been applied in a variety of artificial intelligence (AI) applications. However, sharing personal data using IoT edge devices carries inherent risks to individual privacy. Meanwhile, the energy and memory resources needed during the inference process become a constraint to the resource-limited IoT edge devices. This article brings memory hardware optimization to meet the tight power budget in IoT edge devices by considering the privacy, accuracy, and power efficiency tradeoff in differentially efficient deep learning systems. Based on a detailed analysis on these characteristics, an integer linear programs (ILP) model is developed to minimize mean square error (MSE), thereby enabling optimal input data memory design. Our simulation results in 45-nm CMOS technology show that the proposed technique can enable near-threshold energy-efficient memory operation for different privacy requirements, with less than 1% degradation in classification accuracy.
Jonathon Edstrom, Hritom Das, Na Gong
IEEE Trans. Very Large Scale Integr. Syst.4
2019 Data-Pattern Enabled Self-Recovery Low-Power Storage System for Big Video Data
abstract
The growing popularity of powerful mobile devices such as smart phones and tablet devices has resulted in the exponential growth of demand for video applications. However, due to the large video data size and intensive computation, mobile video applications require frequent embedded memory access, which consumes a large amount of power and limits battery life. In this paper, we present a low-cost self-recovery video storage system by investigating meaningful data patterns hidden in big video data, by introducing data mining techniques to the hardware design process. We propose a two-dimensional data-pattern approach to explore horizontal data-association and vertical data-correlation characteristics. Such data relationship discovery and pattern identification enable a new dimension for the hardware design space and bring self-recovery ability to memories in the presence of bitcell failures. Based on the identified optimal data patterns, we present a low-cost and efficient SRAM design to enable data self-recovery at low voltages. A 45nm 32 kb SRAM is implemented that delivers good video quality at near-threshold voltage (0.5 V) with negligible area overhead (7.94 percent).
Jonathon Edstrom, Yifu Gong, Na Gong
IEEE Trans. Big Data5
2018 Viewer-Aware Intelligent Efficient Mobile Video Embedded Memory
abstract
Embedded memory is a critical component in today's mobile video processing systems, increasingly dominating power consumption and shortening battery life of mobile devices. Traditional hardware-level power optimization techniques usually come with significant implementation overhead to solve the memory failure problem during low-voltage operations. This paper presents a novel mobile video memory to exploit the power saving opportunities provided by a viewer experience under environmental visual interference. The viewing contexts, in particular the ambient luminance, significantly influence the quality of the viewer experience, and in the context with higher luminance levels, mobile users have higher tolerance to the video degradation. Accordingly, the memory failures can be introduced adaptively to achieve power savings without influencing the viewer experience. To meet the silicon area constraint in mobile devices, a simple but an efficient hardware implementation scheme is developed to minimize area overhead. The experimental results based on a 45-nm CMOS technology show that, as compared with the conventional memory design, the proposed technique can achieve up to 48% power savings with good perceivable quality and negligible implementation overhead.
Jonathon Edstrom, Yifu Gong, Mark E. McCourt, Na Gong
IEEE Trans. Very Large Scale Integr. Syst.8
2018 A Novel Hybrid Delay Unit Based on Dummy TSVs for 3-D On-Chip Memory
Xiaowei Chen 0004, Seyed Alireza Pourbakhsh, Jingyan Fu, Na Gong
IEEE Trans. Very Large Scale Integr. Syst.4
2017 SPIDER: Sizing-Priority-Based Application-Driven Memory for Mobile Video Applications
abstract
Recently, mobile devices such as smartphones and tablets have become the most important medium for delivering internet traffic, especially multimedia content, to end users. However, mobile embedded memory incurs large power consumption owing to the highly frequent access and extensive computation. This paper presents an sizing-priority-based application-driven memory (SPIDER) design methodology for low-power mobile video applications. We investigate the size dependent memory failure characteristics and effectively reduce the memory failure rate with low area overhead. Also, we develop a model for the influence of the memory failure on video output, connecting the hardware design process and application requirement. Based on this, we design the SPIDER algorithms for area-priority and quality-priority mobile video applications. During this process, we also consider the contribution of both Luma and Chroma to output quality, avoiding over-optimization issue. We also develop a hardware-based python-assisted SPIDER simulator to apply our proposed design in one leading edge video compression system, the H.264 decoder. Our simulation results in 45-nm CMOS technology show that SPIDER supports mobile videos successfully as voltage downs to 500 mV from 1 V, enabling over 70% power savings in memory arrays.
Na Gong, Seyed Alireza Pourbakhsh, Xiaowei Chen 0004
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Data-Pattern enabled Self-Recovery multimedia storage system for near-threshold computing
abstract
The growing popularity of powerful mobile devices such as smart phones and tablet devices has resulted in the exponential growth of demand for video applications. However, due to the intensive computation of the video decoding process, mobile video applications require frequent embedded memory access, which consumes a large amount of power and limits battery life. Various low-voltage memory techniques have been investigated to enhance the energy efficiency of multimedia processing system. Unfortunately, the existing research suffers from high implementation complexity and large area overhead. In this paper, we present a low-cost self-recovery video storage system by investigating meaningful data patterns hidden in mobile video data. Specifically, we propose a two-dimensional data-pattern approach to explore horizontal data-association and vertical data-correlation characteristics. Based on the identified optimal data patterns, we present a simple circuit-level SRAM design to enable self-recovery at low voltages. A 45nm 32kb SRAM is designed that delivers good video quality at near-threshold voltage (0.5 V) with negligible area overhead (3.97%).
Na Gong, Jonathon Edstrom
ICCD1
2016 Luminance-adaptive smart video storage system
abstract
The use of mobile devices is rapidly growing each year due to their portability and convenience. User experience and battery life are both crucial topics in the advancement of these devices. Sensing the light environment of a mobile device and adapting the output of the device accordingly can create positive results for both user experience and battery life. In this paper, we concentrate on different levels of ambient luminance and its influence on the viewing experience of the user. Since the human visual system is less sensitive to video noise at higher levels of ambient luminance, under those conditions we allow more noise to achieve power consumption reduction. Through subjective testing we show that in general our method works well for attaining a good quality of experience to the user. The hardware simulation for our design results in 32.6% power savings.
Jonathon Edstrom, Huan Gu, Enrique Alvarez Vazquez, Mark E. McCourt, Na Gong
ISCAS7
2016 Data-Driven Low-Cost On-Chip Memory with Adaptive Power-Quality Trade-off for Mobile Video Streaming
abstract
Nowadays, people enjoy watching mobile videos more than ever and mobile video streaming contributes to the majority of the total mobile data traffic. However, due to the high power consumption of mobile video decoders, especially the on-chip memories, short battery life represents one of the biggest contributors to user dissatisfaction. Various mobile embedded memory techniques have been investigated to reduce power consumption and prolong battery life. Unfortunately, the existing hardware-level research suffers from high implementation complexity and large overhead. In this paper, by introducing advanced data-mining techniques, we investigate meaningful data patterns hidden in mobile video data and apply the identified patterns to implement a low-power flexible hardware design with dynamic power-quality trade-off. A 45nm 32kb SRAM is presented that enables three levels of power-quality trade-off (up to 43.7% power savings) with negligible area overhead (0.06%).
Jonathon Edstrom, Xiaowei Chen 0004, Wei Jin 0006, Na Gong
ISLPED6
2016 cNV SRAM: CMOS Technology Compatible Non-Volatile SRAM Based Ultra-Low Leakage Energy Hybrid Memory System
abstract
A CMOS technology compatible non-volatile SRAM (cNV SRAM) is proposed in this paper to achieve energy efficient on-chip memory. cNV SRAM works as conventional 8T SRAM to keep high speed in work mode; in sleep mode, it backs up the data in its NV component and switches off the power supply, thereby minimizing the leakage energy without data loss. The circuit- and architectural- level implementation schemes of cNV SRAM are developed considering multiple key performance parameters including energy dissipation, access time, write time, noise margin, layout area, restoration time, and injection charges. Simulation results on SPEC 2000 benchmark suite demonstrate that cNV SRAM realizes 86 percent energy savings on average with negligible performance impact and small hardware overhead as compared to conventional SRAM. Finally, the impact of the sleep time and memory size on the effectiveness of cNV SRAM is analyzed in detail and it shows that cNV SRAM is particularly effective to implement large on-chip memories with long idle time.
Haibin Yin, Zikui Wei, Zezhong Yang, Na Gong
IEEE Trans. Computers6
2016 PNS-FCR: Flexible Charge Recycling Dynamic Circuit Technique for Low-Power Microprocessors
abstract
Due to the superior speed and area characteristics, dynamic circuits are widely applied in data paths and other time critical components in modern microprocessors. The high switching activity of dynamic circuits, however, consumes significant power. In this paper, a p-type/n-type dynamic circuit selection (PNS) algorithm and a flexible charge recycling (FCR) design methodology are proposed to achieve high power efficiency in data paths. The effects of technology scaling, data path width, design complexity, clock skew, and environmental conditions are discussed. Simulation results show that the power consumption of an arithmetic and logic unit (ALU) with the proposed PNS-FCR can be reduced by up to 60% as compared with a conventional ALU. An 8-bit ALU test circuit has also been manufactured based on a 0.35-μm Global Foundries technology, demonstrating the power and area efficiency of the proposed methodology.
Na Gong, Eby G. Friedman
IEEE Trans. Very Large Scale Integr. Syst.2
2015 TM-RF: Aging-Aware Power-Efficient Register File Design for Modern Microprocessors
abstract
Modern microprocessors employ register files (RFs) for performance enhancement and achieving instruction level parallelism simultaneously. However, RF incurs large power consumption owing to the highly frequent access. Meanwhile, as technology scales, bias temperature instability has become a major reliability concern for RF designers. This paper presents an aging-aware trimodal register file (TM-RF) design to enhance the power efficiency. As instructions pass through the pipeline, TM-RF places the bit-cells in different modes based on the register activity, thereby achieving significant power reduction. To meet design constraints of different applications, we present four schemes to implement the proposed design, providing design flexibility. Additionally, with device selection and worst case sizing methodology, we mitigate aging-effect-induced RF reliability degradation. Simulation results on SPEC 2000 benchmarks demonstrate that TM-RF achieves up to 81.4% power savings and 17% reliability improvement on average, with minimal impact on performance.
Na Gong, Shixiong Jiang, Ramalingam Sridhar
IEEE Trans. Very Large Scale Integr. Syst.1
2009 Switching and leakage power modeling for multiple-supply dynamic gate with delay constraining based on wavelet neural networks
abstract
A model for forecasting the switching power, leakage power and delay of the multiple-supply dynamic OR gates based on wavelet neural networks in 45 nm technology is proposed. By studying the impact of the multiple-supply technique (MST) on the power and delay characteristics, the proposed forecasting model could forecast the nonlinear changing of the switching power, leakage power and delay of the different inputs dynamic OR gates with fast speed convergence and high precision. At last, the trend of the forecasting curve is discussed.
Wuchen Wu, Na Gong, Ligang Hou, Shuqin Geng, Daming Gao
IJCNN3
2009 Estimation for Speed and Leakage Power of Dual Threshold Domino OR Based on Wavelet Neural Networks
Na Gong, Daming Gao, Shuqin Geng, Ligang Hou, Xiaohong Peng, Wuchen Wu
ISNN (1)3