VLDB 2026 Research / reviewers in the wild / expert
Ning Lin
dblp:00/218
· DBLP profile ↗
36ranked-venue papers
15as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 11 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D-S evidence theory-driven attribute reduction for partial label heterogeneous data using label confidence and weighted neighborhood rough set model
Zhaowen Li, Yumei Nong, Guangming Xue, Yonghua Lin, Ning Lin |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | AgeBalance: Low-Cost Lifetime Extension for SRAM-Based PIM AcceleratorsabstractAlthough processing-in-memory (PIM) techniques have widely been used for deep neural networks (DNNs) acceleration, the inference performance of aged PIM-based accelerators remains to be investigated. This paper makes the first attempt to study Hot Carrier Injection (HCI) and Negative Bias Temperature Instability (NBTI) aging impacts on SRAM-based DNN accelerators, which provides a novel and unified framework, termedAgeBalancefor aging detection, analysis and mitigation. First, we discuss a convenient aging detection scheme. Then, we benchmark the inference accuracy drops of DNNs running on aged SRAM-based PIM accelerators. Finally, we propose a low-cost anti-aging training method without incurring additional hardware overhead on SRAM-based DNN accelerators. Extensive experimental results on MNIST, CIFAR10 and AG News datasets show that aging can cause the inference accuracy of shallow or deep DNNs to drop to about 10%, close to random guessing. The aging mitigation scheme proposed in this paper can largely restore the accuracy to the original. Moreover, the SRAM write overhead of our method is much reduced thanks to a score-based training approach, leading to a reduction of 5× to 10× writing energy compared to the traditional training method. Ning Lin, Shaocong Wang 0001, Yangu He, Songqi Wang, Kwunhang Wong, Rongliang Fu, Wenxing Li, Tsung-Yi Ho, Dashan Shang, Xiaojuan Qi 0001, Xiaoming Chen 0003 |
IEEE Trans. Computers | 1 |
| 2026 | Real-Time Digital-Twin for Synergistic Interaction of SMRs and Sustainable Power Systems
Ning Lin, Venkata Dinavahi |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | DANN: Diffractive Acoustic Neural Network for in-sensor computing system target at multi-biomarker diagnosisabstractAnalog machine learning hardware platforms, such as those using wave physics, present potential for edge artificial intelligence (AI) applications due to in-sensor computing architecture, offering superior energy efficiency compared to digital circuits. While the diffractive neural network has been implemented in optical systems, its deployment on integrated acoustic systems has not been achieved due to the challenges associated with hardware optimization. In this paper, we propose the Diffractive Acoustic Neural Network (DANN), a novel approach that applies diffractive neural network algorithms to surface acoustic wave (SAW) systems for in-sensor multibiomarker diagnosis. To address optimization challenges, we introduce a novel training methodology that combines Finite Element Analysis (FEA) with gradient descent. We validate our method on Major Depressive Disorder (MDD) and prostate cancer, achieving accuracies of $74.07 \%$ and $86.0 \%$, respectively, nearly reaching the accuracy levels of clinical diagnoses. By comparing the co-training method with the traditional gradient descent training method and direct training on the FEA model, the co-training method demonstrates its advantages in balancing training efficiency and accuracy. Furthermore, a comparison of power consumption is conducted between the traditional method and the in-sensor computing system, indicating $66 \%$ energy savings attributed to its high level of integration. Lewei He, Ning Lin, Binbin Cui |
DAC | 2 |
| 2025 | Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-MemoryabstractVision Transformers (ViTs) are new foundation models for vision applications. Edge-deploying ViTs to realize energy-saving, low-latency, and high-performance dense predictions have wide applications, such as autonomous driving and surveillance image analysis. However, the quadratic complexity of the self-attention mechanism renders ViTs slow and resource-intensive, particularly for pixel-level dense predictions that involve long contexts. Additionally, the pyramid-like architecture of modern ViT variants leads to an unbalanced workload, further reducing hardware utilization and decreasing the throughput of conventional edge devices. To this end, we propose an algorithm-hardware co-optimized edge ViT accelerator tailored for efficient dense predictions. At the algorithm level, we propose a decoupled chunk attention (DCA) mechanism implemented in a pipelined manner to reduce off-chip memory access, thereby enabling efficient dense predictions within limited on-chip memory. At the architecture level, we introduce a hybrid architecture that combines SRAM-based computing-in-memory (CIM) and nonvolatile RRAM storage to eliminate extensive off-chip memory access, with a fusion scheduling to balance workloads and minimize intermediate on-chip memory access. At the circuit level, a bit/element two-way-reconfigurable CIM macro is proposed to improve hardware utilization across pyramidal ViT blocks with varied matrix sizes. The experimental results on object detection, semantic segmentation, and depth estimation tasks demonstrate that our design can efficiently process patch lengths up to 16384 with a speedup of 18.5×-217.1×, a reduction in memory accesses of 1.7×-7.4×, and an improvement in energy efficiency of 1.8×, under less than 1% performance degradation. Yi Li 0049, Zijian Ye, Xiangqu Fu, Songqi Wang, Shucheng Du, Ning Lin, Dashan Shang, Jinshan Yue, Xiaojuan Qi 0001, Feng Zhang 0014 |
DAC | 6 |
| 2025 | Re4PUF: A Reliable, Reconfigurable ReRAM-based PUF Resilient to DNN and Side Channel AttacksabstractResistive random-access memory (ReRAM) based Physical Unclonable Functions (PUFs) have emerged as an attractive hardware security primitive due to their low energy consumption and compact footprint. However, the reliability of existing ReRAM-based PUFs is challenged by read noise and temperature variations, as well as their resistance to Deep Neural Network (DNN) modeling attacks and Side Channel Attacks (SCAs). In this paper, we propose a novel 3T2R ReRAM-based reconfigurable PUF to address these challenges. By adopting the digital 3T2R voltage division cell design, we improve its reliability against ReRAM read noise and temperature variations, while the adjustable analog supply voltage of inverters enables quick, low-cost reconfigurability without reprogramming ReRAMs, effectively mitigating DNN modeling and SCA vulnerabilities. Our $\mathrm{Re}^{4}$ PUF chip has been experimentally validated, achieving a low Bit Error Rate (BER) of $1 \%$ at $85^{\circ} \mathrm{C}$, a 7.59 -fold reduction compared to existing ReRAM-based PUFs. It also demonstrates robust resistance to both DNN modeling attacks (MLP and Transformer) and SCAs, with success rates of approximately 50% and less than $70 \%$, respectively. Ning Lin, Yangu He, Songqi Wang, Hegan Chen, Kwunhang Wong, Chuxin Li, Jichang Yang, Yongkang Han, Xiaoxin Xu, Dashan Shang |
DAC | 1 |
| 2025 | Guarder: A Stable and Lightweight Reconfigurable RRAM-based PIM Accelerator for DNN IP ProtectionabstractDeploying deep neural networks (DNNs) on conventional digital edge devices faces significant challenges due to high energy consumption. A promising solution is the processing-inmemory (PIM) architecture with resistive random-access memory (RRAM), but RRAM-based systems suffer from imprecise weights due to programming stochasticity and cannot effectively utilize conventional weight encryption/decryption intellectual property (IP) protection schemes. To address these issues, we propose a novel software-hardware co-design Guarder. On the hardware side, we introduce 3T2R cells to achieve reliable multiply-accumulate (MAC) operations and use reconfigurable inverter operating voltages to encode keys for encrypting DNNs on RRAM. On the software side, we implement a contrastive training method that ensures high model accuracy on authorized chips while degrading performance on unauthorized ones. This approach protects DNN IP with minimal hardware overhead while significantly mitigating the effects of RRAM programming stochasticity. Extensive experiments on tasks such as image classification (using MLP, ResNet, and ViT), segmentation (using SegFormer), and image generation (using DiT) validate the effectiveness of our method. The proposed contrastive training ensures negligible performance degradation on authorized chips, while performance on unauthorized chips drops to random guessing or generation. Compared to traditional RRAM accelerators, the 3T2R-based accelerator achieves a $1.41 \times$ reduction in area overhead and a $2.28 \times$ reduction in energy consumption. Ning Lin, Yi Li 0049, Jiankun Li, Jichang Yang, Yangu He, Yukui Luo, Dashan Shang, Xiaoming Chen 0003, Xiaojuan Qi 0001 |
DAC | 1 |
| 2025 | SeDA: Secure and Efficient DNN Accelerators with Hardware/Software SynergyabstractEnsuring the confidentiality and integrity of DNN accelerators is paramount across various scenarios spanning autonomous driving, healthcare, and finance. However, current security approaches typically require extensive hardware resources, and incur significant off-chip memory access overheads. This paper introduces SeDA, which utilizes 1) a bandwidth-aware encryption mechanism to improve hardware resource efficiency, 2) optimal block granularity through intra-layer and inter-layer tiling patterns, and 3) a multi-level integrity verification mechanism that minimizes, or even eliminates, memory access overheads. Experimental results show that SeDA decreases performance overhead by over 12% for both server and edge neural processing units (NPUs), while ensuring robust scalability.11SeDA source code:https://github.com/wayne4s/seda.git Lang Feng 0001, Ning Lin, Zihao Xuan, Rongliang Fu, Tsung-Yi Ho, Yuzhong Jiao, Luhong Liang |
DAC | 4 |
| 2025 | When Pipelined In-Memory Accelerators Meet Spiking Direct Feedback Alignment: A Co-Design for Neuromorphic Edge ComputingabstractSpiking Neural Networks (SNNs) are increasingly favored for deployment on resource-constrained edge devices due to their energy-efficient and event-driven processing capabilities. However, training SNNs remains challenging because of the computational intensity of traditional backpropagation algorithms adapted for spike-based systems. In this paper, we propose a novel software-hardware co-design that introduces a hardware-friendly training algorithm, Spiking Direct Feedback Alignment (SDFA) and implement it on a Resistive Random Access Memory (RRAM)-based In-Memory Computing (IMC) architecture, referred to as PipeSDFA, to accelerate SNN training. Software-wise, the computational complexity of SNN training is reduced by the SDFA through the elimination of sequential error propagation. Hardware-wise, a three-level pipelined dataflow is designed based on IMC architecture to parallelize the training process. Experimental results demonstrate that the PipeSDFA training accelerator incurs less than 2% accuracy loss on five datasets compared to baselines, while achieving 1.1× ~10.5×and 1.37×~2.1× reductions in training time and energy consumption, respectively compared to PipeLayer. Haoxiong Ren, Yangu He, Kwunhang Wong, Rui Bao, Ning Lin, Dashan Shang |
ICCAD | 5 |
| 2025 | SlotPi: Physics-informed Object-centric Reasoning ModelsabstractUnderstanding and reasoning about dynamics governed by physical laws through visual observation, akin to human capabilities in the real world, poses significant challenges. Currently, object-centric dynamic simulation methods, which emulate human behavior, have achieved notable progress but overlook two critical aspects: 1) the integration of physical knowledge into models. Humans gain physical insights by observing the world and apply this knowledge to accurately reason about various dynamic scenarios; 2) the validation of model adaptability across diverse scenarios. Real-world dynamics, especially those involving fluids and objects, demand models that not only capture object interactions but also simulate fluid flow characteristics. To address these gaps, we introduce SlotPi, a slot-based physics-informed object-centric reasoning model. SlotPi integrates a physical module based on Hamiltonian principles with a spatio-temporal prediction module for dynamic forecasting. Our experiments highlight the model's strengths in tasks such as prediction and Visual Question Answering (VQA) on benchmark and fluid datasets. Furthermore, we have created a real-world dataset encompassing object interactions, fluid dynamics, and fluid-object interactions, on which we validated our model's capabilities. The model's robust performance across all datasets underscores its strong adaptability, laying a foundation for developing more advanced world models. Jian Li 0064, Han Wan, Ning Lin, Yuliang Zhan, Ruizhi Chengze, Yi Zhang 0164, Hongsheng Liu 0002, Zidong Wang 0010, Fan Yu 0004, Hao Sun 0002 |
KDD (2) | 3 |
| 2025 | Universally Invariant Learning in Equivariant GNNsabstractEquivariant Graph Neural Networks (GNNs) have demonstrated significant success across various applications. To achieve completeness---that is, the universal approximation property over the space of equivariant functions---the network must effectively capture the intricate multi-body interactions among different nodes. Prior methods attain this via deeper architectures, augmented body orders, or increased degrees of steerable features, often at high computational cost and without polynomial-time solutions. In this work, we present a theoretically grounded framework for constructing complete equivariant GNNs that is both efficient and practical. We prove that a complete equivariant GNN can be achieved through two key components: 1) a complete scalar function, referred to as the canonical form of the geometric graph; and 2) a full-rank steerable basis set. Leveraging this finding, we propose an efficient algorithm for constructing complete equivariant GNNs based on two common models: EGNN and TFN. Empirical results demonstrate that our model demonstrates superior completeness and excellent performance with only a few layers, thereby significantly reducing computational overhead while maintaining strong practical efficacy. Jiacheng Cen, Anyi Li, Ning Lin, Tingyang Xu, Yu Rong 0001, Deli Zhao, Wenbing Huang 0001 |
NeurIPS | 3 |
| 2025 | Local fuzzy rough attribute reduction for large-scale mixed data with limited missing labels based on local fuzzy self information
Zhaowen Li, Run Guo, Ning Lin |
Inf. Sci. | 3 |
| 2024 | Older and Wiser: The Marriage of Device Aging and Intellectual Property Protection of DNNsabstractDeep neural networks (DNNs), such as the widely-used GPT-3 with billions of parameters, are often kept secret due to high training costs and privacy concerns surrounding the data used to train them. Previous approaches to securing DNNs typically require expensive circuit redesign, resulting in additional overheads such as increased area, energy consumption, and latency. To address these issues, we propose a novel hardware-software co-design approach for DNN intellectual property (IP) protection that capitalizes on the inherent aging characteristics of circuits and a novel differential orientation fine-tuning (DOFT) to ensure effective protection. Ning Lin, Shaocong Wang 0001, Yue Zhang 0011, Yangu He, Kwunhang Wong, Arindam Basu, Dashan Shang, Xiaoming Chen 0003 |
DAC | 1 |
| 2024 | LSMR: Synergy Randomness in Liquid State Machine and RRAM-based Analog-digital AcceleratorabstractBio-inspired event sensors are gaining popularity at the edge, such as in robots and wearable electronics. This trend necessitates learning vast amounts of sensory data on the edge, often in few-shot or even zero-shot scenarios, posing challenges in both software and hardware. This paper presents a novel software-hardware co-design to address these issues. Software-wise, we develop an SNN-ANN model, where the SNN encoder is a liquid state machine (LSM) that naturally processes events and significantly reduces learning complexity at the edge due to fixed random weights. The lightweight trainable ANN projection heads are optimized through contrastive learning, enabling zero-shot learning of multimodal events. Hardware-wise, we propose a hybrid analog (RRAM)-digital (CMOS) accelerator - LSMR. The analog in-memory computing core physically implements the LSM by leveraging RRAM stochasticity to generate fixed random weights. The digital core utilizes innovative reconfigurable systolic arrays to accelerate the contrastive learning of ANN projection heads. Extensive experimental outcomes from six neuromorphic datasets, encompassing visual, tactile, and auditory modalities, demonstrate that LSMR considerably improves energy efficiency by a range of 1.65× to 23.70×, in comparison to state-of-the-art edge devices. Simultaneously, it reduces training complexity by a range of 152.83× to 20,587.77× across various edge learning tasks. Ning Lin, Songqi Wang, Xinyuan Zhang 0008, Shaocong Wang 0001, Yangu He, Woyu Zhang, Bo Wang 0153, Jiankun Li, Mingzi Li, Binbin Cui, Yi Li 0049, Jia Chen 0032, Chunwei Xia, Xiaoming Chen 0003, Dashan Shang |
ICCAD | 1 |
| 2024 | SNNGX: Securing Spiking Neural Networks with Genetic XOR Encryption on RRAM-based Neuromorphic AcceleratorabstractBiologically plausible Spiking Neural Networks (SNNs), characterized by spike sparsity, are growing tremendous attention over intellectual edge devices and critical bio-medical applications as compared to artificial neural networks (ANNs). However, there is a considerable risk from malicious attempts to extract white-box information (i.e., weights) from SNNs, as attackers could exploit well-trained SNNs for profit and white-box adversarial concerns. There is a dire need for intellectual property (IP) protective measures. In this paper, we present a novel secure software-hardware co-designed RRAM-based neuromorphic accelerator for protecting the IP of SNNs. Software-wise, we design a tailored genetic algorithm with classic XOR encryption to target the least number of weights that need encryption. From a hardware perspective, we develop a low-energy decryption module, meticulously designed to provide zero decryption latency. Extensive results from various datasets, including NMNIST, DVSGesture, EEGMMIDB, Braille Letter, and SHD, demonstrate that our proposed method effectively secures SNNs by encrypting a minimal fraction of stealthy weights, only 0.00005% to 0.016% weight bits. Additionally, it achieves a substantial reduction in energy consumption, ranging from ×59 to ×6780, and significantly lowers decryption latency, ranging from ×175 to ×4250. Moreover, our method requires as little as one sample per class in dataset for encryption and addresses hessian/gradient-based search insensitive problems. This strategy offers a highly efficient and flexible solution for securing SNNs in diverse applications1. Kwunhang Wong, Songqi Wang, Wei Huang 0042, Xinyuan Zhang 0008, Yangu He, Karl M. H. Lai, Yuzhong Jiao, Ning Lin, Xiaojuan Qi 0001, Xiaoming Chen 0003 |
ICCAD | 8 |
| 2024 | RNC: Efficient RRAM-aware NAS and Compilation for DNNs on Resource-Constrained Edge DevicesabstractComputing-in-memory (CIM) is an emerging computing paradigm, offering noteworthy potential for accelerating neural networks with high parallelism, low latency, and energy efficiency compared to conventional von Neumann architectures. However, existing research has primarily focused on hardware architecture and network co-design for large-scale neural networks, without considering resource constraints. In this study, we aim to develop edge-friendly deep neural networks (DNNs) for accelerators based on resistive random-access memory (RRAM). To achieve this, we propose an edge compilation and resource-constrained RRAM-aware neural architecture search (NAS) framework to search for optimized neural networks meeting specific hardware constraints. Our compilation approach integrates layer partitioning, duplication, and network packing to maximize the utilization of computation units. The resulting network architecture can be optimized for either high accuracy or low latency using a one-shot neural network approach with Pareto optimality achieved through the Non-dominated Sorted Genetic Algorithm II (NSGA-II). The compilation of mobile-friendly networks, like Squeezenet and MobilenetV3 small can achieve over 80% of utilization and over 6x speedup compared to ISAAC-like framework with different crossbar resources. The resulting model from NAS optimized for speed achieved 5x-30x speedup. The code for this paper is available at https://github.com/ArChiiii/rram_nas_comp_pack. Kam Chi Loong, Shihao Han, Sishuo Liu, Ning Lin |
ICCD | 4 |
| 2024 | Are High-Degree Representations Really Unnecessary in Equivariant Graph Neural Networks?abstractEquivariant Graph Neural Networks (GNNs) that incorporate E(3) symmetry have achieved significant success in various scientific applications. As one of the most successful models, EGNN leverages a simple scalarization technique to perform equivariant message passing over only Cartesian vectors (i.e., 1st-degree steerable vectors), enjoying greater efficiency and efficacy compared to equivariant GNNs using higher-degree steerable vectors. This success suggests that higher-degree representations might be unnecessary. In this paper, we disprove this hypothesis by exploring the expressivity of equivariant GNNs on symmetric structures, including $k$-fold rotations and regular polyhedra. We theoretically demonstrate that equivariant GNNs will always degenerate to a zero function if the degree of the output representations is fixed to 1 or other specific values. Based on this theoretical insight, we propose HEGNN, a high-degree version of EGNN to increase the expressivity by incorporating high-degree steerable vectors while maintaining EGNN's efficiency through the scalarization trick. Our extensive experiments demonstrate that HEGNN not only aligns with our theoretical analyses on toy datasets consisting of symmetric structures, but also shows substantial improvements on more complicated datasets such as $N$-body and MD17. Our theoretical findings and empirical results potentially open up new possibilities for the research of equivariant GNNs. Jiacheng Cen, Anyi Li, Ning Lin, Yuxiang Ren, Wenbing Huang 0001 |
NeurIPS | 3 |
| 2024 | DTC: Real-Time and Accurate Distributed Triangle Counting in Fully Dynamic Graph StreamsabstractTriangle counting is a fundamental problem in graph mining, essential for analyzing graph streams with arbitrary edge orders. However, exact counting becomes impractical due to the massive size of real-world graph streams. To address this, approximate algorithms have been developed, but existing distributed streaming algorithms lack adaptability and struggle with edge deletions. In this article, we propose DTC, a novel family of single-pass distributed streaming algorithms for global and local triangle counting in fully dynamic graph streams. Our DTC-AR algorithm accurately estimates triangle counts without prior knowledge of graph size, leveraging multi-machine resources. Additionally, we introduce DTC-FD, an algorithm tailored for fully dynamic graph streams, incorporating edge insertions and deletions. Using Random Pairing and future edge insertion compensation, DTC-FD achieves unbiased and accurate approximations across multiple machines. Experimental results demonstrate significant improvements over baselines. DTC-AR achieves up to 2029.4× and 27.1× more accuracy, while maintaining the best trade-off between accuracy and storage space. DTC-FD reduces estimation errors by up to 32.5× and 19.3×, scaling linearly with graph stream size. These findings highlight the effectiveness of our proposed algorithms in tackling triangle counting in real-world scenarios. The source code and datasets are released and available at https://github.com/Anonymousview/Real-Time-and-Accurate-Distributed-Triangle-Counting-in-Fully-Dynamic-Graph-Streams. Huawei Cao, Ning Lin, Xiaochun Ye, Dongrui Fan |
SRDS | 4 |
| 2023 | Zero-shot Skeleton-based Action Recognition via Mutual Information Estimation and MaximizationabstractZero-shot skeleton-based action recognition aims to recognize actions of unseen categories after training on data of seen categories. The key is to build the connection between visual and semantic space from seen to unseen classes. Previous studies have primarily focused on encoding sequences into a singular feature vector, with subsequent mapping the features to an identical anchor point within the embedded space. Their performance is hindered by 1) the ignorance of the global visual/semantic distribution alignment, which results in a limitation to capture the true interdependence between the two spaces. 2) the negligence of temporal information since the frame-wise features with rich action clues are directly pooled into a single feature vector. We propose a new zero-shot skeleton-based action recognition method via mutual information (MI) estimation and maximization. Specifically, 1) we maximize the MI between visual and semantic space for distribution alignment; 2) we leverage the temporal information for estimating the MI by encouraging MI to increase as more frames are observed. Extensive experiments on three large-scale skeleton action datasets confirm the effectiveness of our method. Wenwen Qiang, Anyi Rao, Ning Lin, Bing Su 0001, Jiaqi Wang 0003 |
ACM Multimedia | 4 |
| 2022 | TB-LNPs: A Web Server for Access to Lung Nodule Prediction Models
Huaichao Luo, Ning Lin, Ziru Huang, Ruiling Zu, Jian Huang 0004 |
ICIC (2) | 2 |
| 2022 | VNet: a versatile network to train real-time semantic segmentation models on a single GPU
Wenxing Li, Ning Lin, Mingzhe Zhang 0005, Xiaoming Chen 0003, Xiaowei Li 0001 |
Sci. China Inf. Sci. | 2 |
| 2021 | ChaoPIM: A PIM-based Protection Framework for DNN Accelerators Using Chaotic EncryptionabstractAlthough deep neural networks (DNNs) have been widely used, DNN models running on ASIC- or FPGA-based accelerators still lack effective and efficient protection. Once DNN models are stolen by attackers, it will not only infringe the intellectual property of model providers but also lead to security issues. The existing parameter encryption method brings greater power consumption, which is difficult to apply to resource-constrained edge devices. This paper proposes an effective and efficient framework –ChaoPIM to protect the security of DNN models by utilizing the chaotic encryption and the Processing-In-Memory (PIM) technology. Detailed experimental results show that our framework can effectively prevent attackers from using DNN models normally, as the accuracy of stolen models is quite low. Compared with the powerful Cortex-A53, Kryo-280, Intel-i5-8265U CPUs and TITAN V GPU, ChaoPIM achieves considerable performance improvements on various DNN models. Ning Lin, Xiaoming Chen 0003, Chunwei Xia, Jing Ye 0001, Xiaowei Li 0001 |
ATS | 1 |
| 2021 | Chaotic Weights: A Novel Approach to Protect Intellectual Property of Deep Neural NetworksabstractDespite the high accuracy achieved by the deep neural network (DNN) technique, there is still a lack of satisfying methodologies to protect the intellectual property (IP) of DNNs, which involves extensive valuable training data, abundant hardware training resources, and fine-tuning skills of experienced experts. Existing solutions based on watermarking cannot prevent malicious/unauthorized users from using well-trained DNNs. This paper proposes chaotic weights (ChaoWs), a novel framework based on the Chaotic Map theory, to protect the IP of DNN providers with very low overhead. Specifically, in order to alleviate the storage overhead and abridge the decryption time, our method makes convolutional or fully connected kernels chaotic by exchanging the weight positions to obtain a satisfying encryption effect, instead of using the conventional idea of encrypting the weight values. Comprehensive experimental evaluations on image classification, semantic segmentation, and name generation demonstrate that ChaoW can effectively protect the IP of DNNs without damaging the inference accuracy, and the impact on the inference speed is negligible. Ning Lin, Xiaoming Chen 0003, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | Redeeming chip-level power efficiency by collaborative management of the computation and communicationabstractPower consumption is the first order design constraint in future many-core processors. Conventional power management approaches usually focus on certain functional components, either computation or communication hardware resources, trying to optimize its power consumption as much as possible, while leave the other part untouched. However, such unilateral power control concept, though has some potentials to contribute overall power reduction, cannot guarantee the optimal power efficiency of the chip. In this paper, we propose a novel Collaborative management approach, coordinating both Computation and Communication infrastructure in tandem, termed as CoCom. Apart from prior work that deals with power control separately, it leverages the correlations between the two parts, as the "key chain" to guide their respective power state coordination to the appropriate direction. Besides, it uses dedicated hybrid on-chip/off-chip mechanisms to minimize the control cost and simultaneously guarantee the effectiveness. Experimental results show that, compared with the conventional unilateral baselines, CoCom is able to achieve abundant power reduction with minimal performance degradation at the same time. Ning Lin, Xiaowei Li 0001 |
ASP-DAC | 1 |
| 2019 | HeadStart: Enforcing Optimal Inceptions in Pruning Deep Neural Networks for Efficient Inference on GPGPUsabstractDeep convolutional neural networks are well-known for the extensive parameters and computation intensity. Structured pruning is an effective solution to obtain a more compact model for the efficient inference on GPGPUs, without designing specific hardware accelerators. However, previous works resort to certain metrics in channel/filter pruning and count on labor intensive fine-tunings to recover the accuracy loss. The "inception" of the pruned model, as another form factor, has indispensable impact to the final accuracy but its importance is often ignored in these works. In this paper, we prove that optimal inception will be more likely to induce a satisfied performance and shortened fine-tuning iterations. We also propose a reinforcement learning based solution, termed as HeadStart, seeking to learn the best way of pruning aiming at the optimal inception. With the help of the specialized head-start network, it could automatically balance the tradeoff between the final accuracy and the preset speedup rather than tilting to one of them, which makes it differentiated from existing works as well. Experimental results show that HeadStart could attain up to 2.25x inference speedup with only 1.16% accuracy loss tested with large scale images on various GPGPUs, and could be well generalized to various cutting-edge DCNN models. Ning Lin, Xiaowei Li 0001 |
DAC | 1 |
| 2019 | VNet: A Versatile Network for Efficient Real-Time Semantic SegmentationabstractMany recent excellent methods for efficient real-time semantic segmentation are of low precision and heavily rely on multiple GPUs for training. In this paper, we rethink the critical factors affecting the accuracy of efficient segmentation models. The previous works usually reduce the input resolution prior to training the parameters of models by cropping or resizing the images. On the contrary, our empirical study shows that the reduced images lose the important content information and details, which are vital to the high precision. However, the previous methods are unable to train the original high-resolution images due to the memory-limited GPUs. To tackle this problem, we propose a novel versatile network (VNet), which employs reversible mechanism and asymmetric convolution to achieve highly efficient and extremely low memory consumption in backward propagation. In particular, we keep all the detailed spatial information of the input images without cropping or resizing to pursue decent prediction accuracy. It is worth noting that VNet can train multiple 1024×2048 high-resolution images on only one standard GPU card. Under the same conditions, our model achieves a new state-of-the-art result on Cityscapes datasets. Specifically, it can process the 1024×2048 high-resolution inputs at a rate of 37.4 and 15.5 frames per second (fps) on a standard GPU and an edge device, respectively, with only 0.16 million parameters. Ning Lin, Jingliang Gao, Shunjie Qiao, Xiaowei Li 0001 |
ICCD | 1 |
| 2019 | When Deep Learning Meets the Edge: Auto-Masking Deep Neural Networks for Efficient Machine Learning on Edge DevicesabstractDeep neural network (DNN) has demonstrated promising performance in various machine learning tasks. Due to the privacy issue and the unpredictable transmission latency, inferring DNN models directly on edge devices trends the development of intelligent systems, like self-driving cars, smart Internet-of-Things (IoTs) and autonomous robotics. The on-device DNN model is obtained by expensive training via vast volumes of high-quality training data in the cloud datacenter, and then deployed into these devices, expecting it to work effectively at the edge. However, edge device always deals with low-quality images caused by compression or environmental noise pollutions. The well-trained model, though could work perfectly on the cloud, cannot adapt to these edge-specific conditions without remarkable accuracy drop. In this paper, we propose an automated strategy, called "AutoMask", to embrace effective machine learning and accelerate DNN inference on edge devices. AutoMask comprises end-to-end trainable software strategies and cost-effective hardware accelerator architecture to improve the adaptability of the device without compromising the constrained computation and storage resources. Extensive experiments, over ImageNet dataset and various state-of-the-art DNNs, show that AutoMask achieves significant inference acceleration and storage reduction while maintains comparable accuracy level on embedded Xilinx Z7020 FPGA, as well as NVIDIA Jetson TX2. Ning Lin, Xing Hu 0001, Jingliang Gao, Mingzhe Zhang 0005, Xiaowei Li 0001 |
ICCD | 1 |
| 2019 | ShuttleNoC: Power-Adaptable Communication Infrastructure for Many-Core ProcessorsabstractNetworks-on-chip (NoCs), as the communication infrastructure in many-core processors, has demonstrated remarkable power consumption along with the technology scaling. However, due to the temporal and spatial heterogeneity of the on-chip traffic, one critical problem is that the NoC power consumption cannot effectively adapt to the variation of its traffic intensity, also known as localized power adaptation, hence yielding a suboptimal power efficiency. Prior approaches either resort to the over-provisioned NoC design or coarse-grained bandwidth scaling to partially alleviate excessive power consumption brought by the traffic temporal or spatial heterogeneity. While in this paper, we propose a novel NoC architecture called Shuttle NoC (ShuttleNoC) to address this challenge. It leverages the link reconfiguration to enable flexible packet traversing between multiple subnetworks, and specialized punch lines to accelerate latency sensitive traffic. With the support of the dedicated power adaptation mechanisms, it is shown in the evaluation that the proposed ShuttleNoC architecture could effectively tackle the power and performance tradeoff and significantly boost the power efficiency compared with the state-of-the-art baselines. Yisong Chang, Guihai Yan, Ning Lin, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Tetris: re-architecting convolutional neural network computation for machine learning acceleratorsabstractInference efficiency is the predominant consideration in designing deep learning accelerators. Previous work mainly focuses on skipping zero values to deal with remarkable ineffectual computation, while zero bits in non-zero values, as another major source of ineffectual computation, is often ignored. The reason lies on the difficulty of extracting essential bits during operating multiply-and-accumulate (MAC) in the processing element. Based on the fact that zero bits occupy as high as 68.9% fraction in the overall weights of modern deep convolutional neural network models, this paper firstly proposes a weight kneading technique that could eliminate ineffectual computation caused by either zero value weights or zero bits in non-zero weights, simultaneously. Besides, a split-and-accumulate (SAC) computing pattern in replacement of conventional MAC, as well as the corresponding hardware accelerator design called Tetris are proposed to support weight kneading at the hardware level. Experimental results prove that Tetris could speed up inference up to 1.50x, and improve power efficiency up to 5.33x compared with the state-of-the-art baselines. Ning Lin, Guihai Yan, Xiaowei Li 0001 |
ICCAD | 3 |
| 2018 | Reduced-Reference Image Quality Assessment Based on Free-Energy Principle with Multi-Channel DecompositionabstractThe free-energy principle studied in brain theory and neuroscience accounts for the mechanism of perception and understanding in human brain, which is highly adapted for measuring the visual quality of perceptions. On the other hand, psychologists and neurologists report that different frequency and orientation components of one stimulus arouse different neurons in striate cortex. In this paper, a novel reduce-reference (RR) image quality assessment (IQA) metric based on free-energy principle in multi-channel is proposed, which is called MCFEM (Multi-Channel Free-Energy principle Metric). We first decompose the input reference image and distorted image via a two-level discrete Haar wavelet transform (DHWT). Next, the free-energy features of each subband images are computed based on sparse representation. Finally, an overall quality index is received through the support vector regressor (SVR). Extensive experimental comparisons on four (LIVE, CSIQ, TID2008 and TID2013) benchmark image databases demonstrate that the proposed method is highly competitive with the representative RR and no-reference models as well as full-reference ones. Wenhan Zhu, Guangtao Zhai, Yutao Liu 0002, Ning Lin, Xiaokang Yang 0001 |
MMSP | 4 |
| 2016 | A Multi-agent Approach for the Newsvendor Problem with Word-of-Mouth Marketing Strategies
Ning Lin |
ICIC (2) | 2 |
| 2014 | Music recommendation based on artist novelty and similarityabstractMost existing systems recommend songs to the user based on the popularity of songs and singers. However, the system proposed in this paper is driven by an emerging and somewhat different need in the music industry-promoting new talents. The system recommends songs based on the novelty of singers (or artists) and their similarity to the user's favorite artists. Novel artists whose popularity is on the rise have a higher priority to be recommended. Specifically, given a user's favorite artists, the system first determines the candidate artists based on their similarity with the favorite artists and then selects those who have a higher novelty score than the favorite artists. Then, the system outputs a playlist composed of the most popular songs of the selected artists. The proposed system can be integrated into most existing systems. Its performance is evaluated using the Spotify Radio Recommender as a reference and a pool of 100 subjects recruited on campus. Experimental results show that our system achieves a high novelty score and a competitive user-preference score. Ning Lin, Ping-Chia Tsai, Yu-An Chen, Homer H. Chen |
MMSP | 1 |
| 2005 | A Boundary Element-Based Approach to Analysis of LV Deformation
Ning Lin, Albert J. Sinusas, James S. Duncan |
MICCAI | 2 |
| 2003 | Analysis of Left Ventricular Motion Using a General Robust Point Matching Algorithm
Ning Lin, Xenophon Papademetris, Albert J. Sinusas, James S. Duncan |
MICCAI (1) | 1 |
| 2003 | Combinative multi-scale level set framework for echocardiographic image segmentation
Ning Lin, Weichuan Yu, James S. Duncan |
Medical Image Anal. | 1 |
| 2002 | Combinative Multi-scale Level Set Framework for Echocardiographic Image Segmentation
Ning Lin, Weichuan Yu, James S. Duncan |
MICCAI (1) | 1 |