Jian Meng

dblp:18/4220 · DBLP profile ↗
← Back
38ranked-venue papers
12as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 23 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 20 · 10 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SAVAF: Sparse Audio-Visual Rendering with Multihead Acoustic Field Attention Network
Ahmed Hasssan, Jian Meng, Jae-sun Seo
ICPR (4)2
2026 DCSHARP: 3D Gaussian Splatting with Direction Cosine Spherical Harmonics and Shape-Aware Pruning
abstract
3D Gaussian Splatting (3DGS) shows outstanding rendering quality for novel view synthesis. Despite its performance, the massive amount of Gaussian blobs leads to expensive run-time sorting and irregular memory access during rendering. Although 3DGS-based pruning algorithm has been widely explored, most of the current research has mainly focused on designing a proper pruning metric and the root cause behind the inevitable quality degradation remains underexplored for highly-sparse 3DGS. In particular, our investigation shows that the Spherical Harmonics (SH) of 3DGS is insufficient to capture high-frequency anisotropic reflections and specular highlights during rendering, especially with sparsified Gaussians. Motivated by that, this work proposes Direction Cosine Spherical Harmonics with Shape-Aware Pruning (DCSHARP). Specifically, the proposed Direction Cosine Spherical Harmonics (DCSH) replaces the vanilla spherical harmonics by facilitating the expressiveness of 3DGS on high-frequency and highly reflective scenes. Unlike recent works that rely on trainable masks or pseudo-rendering scores, the proposed Shape-aware Pruning method enables "pruning on-the-fly" while achieving high quality rendering. As a combined scheme, the proposed DCSHARP reduces the number of active Gaussians by up to 3.9× and improves rendering throughput by 1.9× with ZERO quality degradation compared to the vanilla 3DGS. Furthermore, the proposed DCSH scheme outperforms the vanilla 3DGS on all the mainstream benchmarks by simply replacing the vanilla SH with the DCSH. The source code of the proposed method will be open-sourced.
Ahmed Hasssan, Jian Meng, Yuanbo Xiangli, Jae-sun Seo
WACV2
2026 Unveiling the true potential of blockchain consensus: A comprehensive survey
Jiguo Yu, Baobao Chai, Qin Hu 0001, Tianqing He, Jianyuan Li, Jian Meng
J. Syst. Archit.7
2025 Closest Neighbors are Harmful for Lightweight Masked Auto-encoders
abstract
Learning the visual representation via masked auto-encoder (MAE) training has been proven to be a powerful technique. Transferring the pre-trained vision transformer (ViT) to downstream tasks leads to superior performance compared to conventional task-by-task supervised learning. Recent research works on MAE focus on large-sized vision transformers (>50 million parameters) with outstanding performance. However, improving the generality of the under-parametrized lightweight model has been widely ignored. In practice, downstream applications are commonly intended for resource-constrained platforms, where large-scale ViT cannot easily meet the resource budget. Current lightweight MAE training heavily relies on knowledge distillation with a pre-trained teacher, whereas the root cause behind the poor performance remains under-explored. Motivated by that, this paper first introduces the concept of "closest neighbor patch" to characterize the local semantics among the input tokens. Our discovery shows that the lightweight model failed to distinguish different local information, leading to aliased understanding and poor accuracy. Motivated by this finding, we propose NoR-MAE, a novel MAE training algorithm for lightweight vision transformers. NoR-MAE elegantly repels the semantic aliasing between patches and their closest neighboring patch (semantic centroid) with negligible training cost overhead. With the ViT-Tiny model, NoR-MAE achieves up to 7.22%/3.64% accuracy improvements on ImageNet-100/ImageNet-1K datasets, as well as up to 5.13% accuracy improvements in tested downstream tasks. https://github.com/SeoLabCornell/NoR-MAE
Jian Meng, Li Yang 0009, Deliang Fan, Jinwoo Shin, Jae-sun Seo
CVPR1
2025 Quant-NeRF: Efficient End-to-End Quantization of Neural Radiance Fields with Low-Precision 3D Gaussian Representation
abstract
Neural Radiance Field (NeRF) has been widely investigated for high-quality 3D object rendering based on captured 2D images. Previous research works have continuously improved the rendering quality with various sample representation and encoding strategies. However, a common bottleneck of NeRF is the extreme computational cost and the lack of compatibility with resource-constrained hardware. Despite the high fidelity of the rendered object, the extensive processing time of the pre-trained NeRF model largely degrades the feasibility of energy-efficient NeRF, especially for resource-constrained edge devices such as augmented/virtual reality (AR/VR) headsets. Most prior works focused on efficient hash table representation or simplified tensorial radiance fields with high-precision representation. However, the efficient, low precision, and hardware deployable NeRF with Gaussian-based modeling remains largely under-explored. Motivated by that, this paper proposes Quant-NeRF, a novel hardware-aware algorithm that performs 3D rendering with end-to-end low-precision representation and hardware deployable computation. Quant-NeRF achieves 60× acceleration compared to prior works on GPU, while maintaining high rendering quality as the full-precision baseline. The proposed algorithm achieves peak performance of 250 FPS.
Ahmed Hasssan, Anupreetham Anupreetham, Jian Meng, Jae-sun Seo
ICASSP3
2025 Hierarchical and Multi-scale Attention Network for Retinal Artery/Vein Segmentation and Diameter Estimation
Jian Meng, Zhenchao Cui
ICONIP (5)1
2025 Low-Precision Normalization Algorithm and Accelerator for Neural Network Training
abstract
As one of the most essential techniques in modern deep learning, normalization layer largely improves the convergence speed and performance of deep neural networks (DNN). However, calculating normalization statistics during training is costly, as the two-pass algorithm requires repeated accumulation on the same data. Although parallelized one-pass algorithms are used for variance calculation, low-precision floating-point arithmetic often leads to catastrophic cancellation and accuracy degradation due to uncentered batch statistics and rounding errors. In this paper, we propose Welford-Pairwise Normalization (WPN), a novel normalization algorithm with accelerator design for low-precision training. WPN resolves numerical instability while performing parallel computation. Implemented on the Xilinx Alveo U280 FPGA, WPN achieves up to 28× throughput improvement compared to standard normalization layer with less than 1% accuracy loss with BEiT3 vision transformer, ResNet, MobileNet, and VGG16 model.
Han-Sok Suh, Jian Meng, Jae-sun Seo
ISCAS2
2025 Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
abstract
Edge inference for large language models (LLM) offers secure, low-latency, and cost-effective inference solutions. We emphasize that an edge accelerator should achieve high area efficiency and minimize external memory access (EMA) during the memory-bound decode stage, while maintaining high energy efficiency during the compute-intensive prefill stage. This paper proposes an edge LLM inference accelerator featuring a hybrid systolic array (HSA) architecture that optimizes inference efficiency in both stages. To further reduce EMA, we adopt MXINT4 weight quantization and propose an optimized dataflow tailored for HSA, ensuring negligible dequantization overhead and achieving 100% hardware utilization with minimal accuracy loss under edge DRAM bandwidth constraints. For non-linear operations, we incorporate optimized root mean square normalization (RMSNorm) and rotary position embedding (RoPE) units, reducing their latency, area, and memory access overhead while enabling end-to-end inference on our accelerator. Our solution achieves 247/117 (token/s/mm2) while running a 1.3B LLM on long-input/long-output scenarios, providing >2.45×/13.5× improvement over existing approaches, while maintaining superior energy efficiency in token generation.
Chun-Ting Chen, Jian Meng, Mohamed S. Abdelfattah, Jae-sun Seo
ISLPED3
2024 POCA: Post-training Quantization with Temporal Alignment for Codec Avatars
Jian Meng, Yuecheng Li, Chenghui Li, Syed Shakib Sarwar, Dilin Wang, Jae-sun Seo
ECCV (40)1
2024 Spiking Neural Network with Learnable Threshold for Event-based Classification and Object Detection
abstract
Spiking neural networks (SNNs) have received increasing attention due to their high biological plausibility and energy efficiency. The binary spike-based information propagation enables efficient sparse computation for event-based computer vision applications. However, most prior works use the heuristically selected fixed threshold for spiking neurons, which limits the dynamics of SNNs toward further optimizing the performance. In the meantime, the optimization space of the existing trainable spike neurons is often limited by various constraints. Motivated by this, this paper investigates the plausibility of freely optimizing the threshold during direct SNN training. Specifically, we propose LT-SNN, a novel SNN training algorithm with a self-adaptive learnable potential threshold to improve SNN performance. LT-SNN optimizes the layer-wise firing threshold throughout SNN training without any high-precision spike representation or learning constraints. Extensive experiments are performed across event-based and static computer vision datasets, including both image classification and object detection tasks. Equipped with high adaptiveness that fully captures the dynamics of SNNs, LT-SNN outperforms the recent state-of-the-art works. Furthermore, LT-SNN is compatible with SNN models based on both convolutional neural networks (CNN) and vision transformers (ViT).
Ahmed Hasssan, Jian Meng, Jae-sun Seo
IJCNN2
2024 A 28nm Scalable and Flexible Accelerator for Sparse Transformer Models
abstract
Transformer-based model has been widely utilized in deep learning. The accuracy-driven applications broadly expand the model size, whereas the current hardware accelerator designs failed to adaptively alternate the scalability to match the corresponding computation intensity of different model sizes. Meanwhile, supporting the transformer models with different sizes requires flexibility for various matrix multiplication under different dimensions. On the higher level, the complex computation flow within transformer models urges a flexible data management design for accelerators. Furthermore, the massive model size enables the possibility of utilizing sparsity and eliminating the redundancy of the model. However, exploring the fine-grained sparsity on hardware remains challenging and under-explored for transformer accelerators. Finally, the non-linear functions and modules of the transformer model require a dedicated hardware design to balance the trade-off between accuracy and hardware cost. Motivated by that, we propose a novel hardware accelerator designed for transformer-based models. In particular, we propose the row-wise matrix multiplication processing elements (RMMPE) and the post-PE processors (PPE). RMMPE computes matrix multiplication in row-wise products with high data reuse. Furthermore, RMMPE efficiently handles the unstructured sparse matrix multiplication with various dimensionality, elevating the scalability and flexibility for different transformer models. PPE computes complex functions in linear approximation. The proposed accelerator achieves 17.1 TOPS peak throughput and 19.5 TOPS/W peak energy efficiency, outperforming the recent SoTA transformer accelerators.
Yuan Liao 0004, Jian Meng, Jae-sun Seo
ISLPED2
2024 BBS: Bi-Directional Bit-Level Sparsity for Deep Learning Acceleration
abstract
Bit-level sparsity methods skip ineffectual zero-bit operations and are typically applicable within bit-serial deep learning accelerators. This type of sparsity at the bit-level is especially interesting because it is both orthogonal and compatible with other deep neural network (DNN) efficiency methods such as quantization and pruning. Furthermore, it comes at little or no accuracy degradation and can be performed completely post-training. However, current bit-sparsity approaches lack practicality because of (1) load imbalance from the random distribution of zero bits, (2) unoptimized external memory access because all bits are fetched from off-chip memory, and (3) high hardware implementation overhead, including large multiplexers and shifters to support sparsity at the bit level. In this work, we improve the practicality and efficiency of bit-level sparsity through a novel algorithmic bit-pruning, averaging, and compression method, and a co-designed efficient bit-serial hardware accelerator. On the algorithmic side, we introduce bi-directional bit sparsity (BBS). The key insight of BBS is that we can leverage bit sparsity in a symmetrical way to prune either zero-bits or one-bits. This significantly improves the load balance of bit-serial computing and guarantees the level of sparsity to be more than 50%. On top of BBS, we further propose two bit-level binary pruning methods that require no retraining, and can be seamlessly applied to quantized DNNs. Combining binary pruning with a new tensor encoding scheme, BBS can both skip computation and reduce the memory footprint associated with bi-directional sparse bit columns. On the hardware side, we demonstrate the potential of BBS through BitVert, a bit-serial architecture with an efficient PE design to accelerate DNNs with low overhead, exploiting our proposed binary pruning. Evaluation on seven representative DNN models shows that our approach achieves: (1) on average 1.66× reduction in model size with negligible accuracy loss of < 0.5%; (2) up to 3.03× speedup and 2.44× energy saving compared to prior DNN accelerators.
Yuzong Chen 0001, Jian Meng, Jae-sun Seo, Mohamed S. Abdelfattah
MICRO2
2024 XGrad: Boosting Gradient-Based Optimizers With Weight Prediction
abstract
In this paper, we propose a general deep learning training framework XGrad which introduces weight prediction into the popular gradient-based optimizers to boost their convergence and generalization when training the deep neural network (DNN) models. In particular, ahead of each mini-batch training, the future weights are predicted according to the update rule of the used optimizer and are then applied to both the forward pass and backward propagation. In this way, during the whole training period, the optimizer always utilizes the gradients w.r.t. the future weights to update the DNN parameters, making the gradient-based optimizer achieve better convergence and generalization compared to the original optimizer without weight prediction. XGrad is rather straightforward to implement yet pretty effective in boosting the convergence of gradient-based optimizers and the accuracy of DNN models. Empirical results concerning five popular optimizers including SGD with momentum, Adam, AdamW, AdaBelief, and AdaM3 demonstrate the effectiveness of our proposal. The experimental results validate that XGrad can attain higher model accuracy than the baseline optimizers when training the DNN models.
Lei Guan 0001, Dongsheng Li 0001, Yanqi Shi, Jian Meng
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Lexical Entrainment in Bilingual Language Use
Yongjia Song, Kathlyn Canales, Yuting Gu, Jiachen Jin, Jian Meng, Judith F. Kroll, Gregory Scontras
CogSci5
2023 PRIVE: Efficient RRAM Programming with Chip Verification for RRAM-based In-Memory Computing Acceleration
abstract
As deep neural networks (DNNs) have been success-fully developed in many applications with continuously increasing complexity, the number of weights in DNNs surges, leading to consistent demands for denser memories than SRAMs. RRAM-based in-memory computing (IMC) achieves high density and energy-efficiency for DNN inference, but RRAM programming remains to be a bottleneck due to high write latency and energy consumption. In this work, we present the Progressive-wRite In-memory program-VErify (PRIVE) scheme, which we verify with an RRAM testchip for IMC-based hardware acceleration for DNNs. We optimize the progressive write operations on different bit positions of RRAM weights to enable error compensation and reduce programming latency/energy, while achieving high DNN accuracy. For 5-bit precision DNNs, PRIVE reduces the RRAM programming energy by 1.82×, while maintaining high accuracy of 91.91% (VGG-7) and 71.47% (ResNet-18) on CIFAR-10 and CIFAR-100 datasets, respectively.
Wangxin He, Jian Meng, Sujan K. Gonugondla, Shimeng Yu, Naresh R. Shanbhag, Jae-sun Seo
DATE2
2023 Slimmed Asymmetrical Contrastive Learning and Cross Distillation for Lightweight Model Training
abstract
Contrastive learning (CL) has been widely investigated with various learning mechanisms and achieves strong capability in learning representations of data in a self-supervised manner using unlabeled data. A common fashion of contrastive learning on this line is employing mega-sized encoders to achieve comparable performance as the supervised learning counterpart. Despite the success of the labelless training, current contrastive learning algorithms *failed* to achieve good performance with lightweight (compact) models, e.g., MobileNet, while the requirements of the heavy encoders impede the energy-efficient computation, especially for resource-constrained AI applications. Motivated by this, we propose a new self-supervised CL scheme, named SACL-XD, consisting of two technical components, **S**limmed **A**symmetrical **C**ontrastive **L**earning (SACL) and **Cross**-**D**istillation (XD), which collectively enable efficient CL with compact models. While relevant prior works employed a strong pre-trained model as the teacher of unsupervised knowledge distillation to a lightweight encoder, our proposed method trains CL models from scratch and outperforms them even without such an expensive requirement. Compared to the SoTA lightweight CL training (distillation) algorithms, SACL-XD achieves 1.79% ImageNet-1K accuracy improvement on MobileNet-V3 with 64$\times$ training FLOPs reduction.
Jian Meng, Li Yang 0009, Kyungmin Lee, Jinwoo Shin, Deliang Fan, Jae-sun Seo
NeurIPS1
2023 Algorithm-hardware Co-optimization for Energy-efficient Drone Detection on Resource-constrained FPGA
abstract
Convolutional neural network (CNN)-based object detection has achieved very high accuracy; e.g., single-shot multi-box detectors (SSDs) can efficiently detect and localize various objects in an input image. However, they require a high amount of computation and memory storage, which makes it difficult to perform efficient inference on resource-constrained hardware devices such as drones or unmanned aerial vehicles (UAVs). Drone/UAV detection is an important task for applications including surveillance, defense, and multi-drone self-localization and formation control. In this article, we designed and co-optimized an algorithm and hardware for energy-efficient drone detection on resource-constrained FPGA devices. We trained an SSD object detection algorithm with a custom drone dataset. For inference, we employed low-precision quantization and adapted the width of the SSD CNN model. To improve throughput, we use dual-data rate operations for DSPs to effectively double the throughput with limited DSP counts. For different SSD algorithm models, we analyze accuracy or mean average precision (mAP) and evaluate the corresponding FPGA hardware utilization, DRAM communication, and throughput optimization. We evaluated the FPGA hardware for a custom drone dataset, Pascal VOC, and COCO2017. Our proposed design achieves a high mAP of 88.42% on the multi-drone dataset, with a high energy efficiency of 79 GOPS/W and throughput of 158 GOPS using the Xilinx Zynq ZU3EG FPGA device on the Open Vision Computer version 3 (OVC3) platform. Our design achieves 1.1 to 8.7× higher energy efficiency than prior works that used the same Pascal VOC dataset, using the same FPGA device, but at a low-power consumption of 2.54 W. For the COCO dataset, our MobileNet-V1 implementation achieved an mAP of 16.8, and 4.9 FPS/W for energy-efficiency, which is ∼ 1.9× higher than prior FPGA works or other commercial hardware platforms.
Han-Sok Suh, Jian Meng, Ty Nguyen, Vijay Kumar 0001, Yu Cao 0001, Jae-sun Seo
ACM Trans. Reconfigurable Technol. Syst.2
2022 XBM: A Crossbar Column-wise Binary Mask Learning Method for Efficient Multiple Task Adaption
abstract
Recently, utilizing ReRAM crossbar array to accelerate DNN inference on single task has been widely studied. However, using the crossbar array for multiple task adaption has not been well explored. In this paper, for the first time, we propose XBM, a novel crossbar column-wise binary mask learning method for multiple task adaption in ReRAM crossbar DNN accelerator. XBM leverages the mask-based learning algorithm's benefit to avoid catastrophic forgetting to learn a task-specific mask for each new task. With our hardware-aware design innovation, the required masking operation to adapt for a new task could be easily implemented in existing crossbar based convolution engine with minimal hardware/ memory overhead and, more importantly, no need of power hungry cell re-programming, unlike prior works. The extensive experimental results show that compared with state-of-the-art multiple task adaption methods, XBM keeps the similar accuracy on new tasks while only requires 1.4% mask memory size compared with popular piggyback. Moreover, the elimination of cell re-programming or tuning saves up to 40% energy during new task adaption.
Fan Zhang 0069, Li Yang 0009, Jian Meng, Yu Cao 0001, Jae-sun Seo, Deliang Fan
ASP-DAC3
2022 Contrastive Dual Gating: Learning Sparse Features With Contrastive Learning
abstract
Contrastive learning (or its variants) has recently become a promising direction in the self-supervised learning domain, achieving similar performance as supervised learning with minimum fine-tuning. Despite the labeling efficiency, wide and large networks are required to achieve high accuracy, which incurs a high amount of computation and hinders the pragmatic merit of self-supervised learning. To effectively reduce the computation of insignificant features or channels, recent dynamic pruning algorithms for supervised learning employed auxiliary salience predictors. However, we found that such salience predictors cannot be easily trained when they are naïvely applied to contrastive learning from scratch. To address this issue, we propose contrastive dual gating (CDG), a novel dynamic pruning algorithm that skips the uninformative features during contrastive learning without hurting the trainability of the networks. We demonstrate the superiority of CDG with ResNet models for CIFAR-10, CIFAR-100, and ImageNet-100 datasets. Compared to our implementations of state-of-the-art dynamic pruning algorithms for self-supervised learning, CDG achieves up to 15% accuracy improvement for CIFAR-10 dataset with higher computation reduction.
Jian Meng, Li Yang 0009, Jinwoo Shin, Deliang Fan, Jae-sun Seo
CVPR1
2022 XMA: a crossbar-aware multi-task adaption framework via shift-based mask learning method
abstract
ReRAM crossbar array as a high-parallel fast and energy-efficient structure attracts much attention, especially on the acceleration of Deep Neural Network (DNN) inference on one specific task. However, due to the high energy consumption of weight re-programming and the ReRAM cells' low endurance problem, adapting the crossbar array for multiple tasks has not been well explored. In this paper, we propose XMA, a novel crossbar-aware shift-based mask learning method for multiple task adaption in the ReRAM crossbar DNN accelerator for the first time. XMA leverages the popular mask-based learning algorithm's benefit to mitigate catastrophic forgetting and learn a task-specific, crossbar column-wise, and shift-based multi-level mask, rather than the most commonly used element-wise binary mask, for each new task based on a frozen backbone model. With our crossbar-aware design innovation, the required masking operation to adapt for a new task could be implemented in an existing crossbar-based convolution engine with minimal hardware/memory overhead and, more importantly, no need for power-hungry cell re-programming, unlike prior works. The extensive experimental results show that, compared with state-of-the-art multiple task adaption Piggyback method [1], XMA achieves 3.19% higher accuracy on average, while saving 96.6% memory overhead. Moreover, by eliminating cell re-programming, XMA achieves ~4.3x higher energy efficiency than Piggyback.
Fan Zhang 0069, Li Yang 0009, Jian Meng, Jae-sun Seo, Yu Cao 0001, Deliang Fan
DAC3
2022 XST: A Crossbar Column-wise Sparse Training for Efficient Continual Learning
abstract
Leveraging the ReRAM crossbar-based In-Memory-Computing (IMC) to accelerate single task DNN inference has been widely studied. However, using the ReRAM crossbar for continual learning has not been explored yet. In this work, we propose XST, a novel crossbar column-wise sparse training framework for continual learning. XST significantly reduces the training cost and saves inference energy. More importantly, it is friendly to existing crossbar-based convolution engine with almost no hardware overhead. Compared with the state-of-the-art CPG method, the experiments show that XST's accuracy achieves 4.95 % higher accuracy. Furthermore, XST demonstrates ~5.59 × training speedup and 1.5 × inference energy-saving.
Fan Zhang 0069, Li Yang 0009, Jian Meng, Jae-sun Seo, Yu Cao 0001, Deliang Fan
DATE3
2022 Get More at Once: Alternating Sparse Training with Gradient Correction
abstract
Recently, a new trend of exploring training sparsity has emerged, which remove parameters during training, leading to both training and inference efficiency improvement. This line of works primarily aims to obtain a single sparse model under a pre-defined large sparsity ratio. It leads to a static/fixed sparse inference model that is not capable of adjusting or re-configuring its computation complexity (i.e., inference structure, latency) after training for real-world varying and dynamic hardware resource availability. To enable such run-time or post-training network morphing, the concept of dynamic inference' ortraining-once-for-all' has been proposed to train a single network consisting of multiple sub-nets once, but each sub-net could perform the same inference function with different computing complexity. However, the traditional dynamic inference training method requires a joint training scheme with multi-objective optimization, which suffers from very large training overhead. In this work, for the first time, we propose a novel alternating sparse training (AST) scheme to train multiple sparse sub-nets for dynamic inference without extra training cost compared to the case of training a single sparse model from scratch. Furthermore, to mitigate the interference of weight update among sub-nets, we propose gradient correction within the inner-group iterations to reduce their weight update interference. We validate the proposed AST on multiple datasets against state-of-the-art sparse training method, which shows that AST achieves similar or better accuracy, but only needs to train once to get multiple sparse sub-nets with different sparsity ratios. More importantly, compared with the traditional joint training based dynamic inference training methodology, the large training overhead is completely eliminated without affecting the accuracy of each sub-net.
Li Yang 0009, Jian Meng, Jae-sun Seo, Deliang Fan
NeurIPS2
2022 Reversible data hiding with enhancing contrast and preserving brightness in medical image
Yang Yang 0059, Jian Meng, Weiming Zhang 0001
J. Inf. Secur. Appl.3
2022 Hybrid RRAM/SRAM in-Memory Computing for Robust DNN Acceleration
abstract
RRAM-based in-memory computing (IMC) effectively accelerates deep neural networks (DNNs) and other machine learning algorithms. On the other hand, in the presence of RRAM device variations and lower precision, the mapping of DNNs to RRAM-based IMC suffers from severe accuracy loss. In this work, we propose a novel hybrid IMC architecture that integrates an RRAM-based IMC macro with a digital SRAM macro using a programmable shifter to compensate for the RRAM variations and recover the accuracy. The digital SRAM macro consists of a small SRAM memory array and an array of multiply-and-accumulate (MAC) units. The nonideal output from the RRAM macro, due to device and circuit nonidealities, is compensated by adding the precise output from the SRAM macro. In addition, the programmable shifter allows for different scales of compensation by shifting the SRAM macro output relative to the RRAM macro output. On the algorithm side, we develop a framework for the training of DNNs to support the hybrid IMC architecture through ensemble learning. The proposed framework performs quantization (weights and activations), pruning, RRAM IMC-aware training, and employs ensemble learning through different compensation scales by utilizing the programmable shifter. Finally, we design a silicon prototype of the proposed hybrid IMC architecture in the 65-nm SUNY process to demonstrate its efficacy. Experimental evaluation of the hybrid IMC architecture shows that the SRAM compensation allows for a realistic IMC architecture with multilevel RRAM cells (MLCs) even though they suffer from high variations. The hybrid IMC architecture achieves up to 21.9%, 12.65%, and 6.52% improvement in post-mapping accuracy over state-of-the-art techniques, at minimal overhead, for ResNet-20 on CIFAR-10, VGG-16 on CIFAR-10, and ResNet-18 on ImageNet, respectively.
Zhenyu Wang 0016, Injune Yeo, Li Yang 0009, Jian Meng, Maximilian Liehr, Rajiv V. Joshi, Nathaniel C. Cady, Deliang Fan, Jae-sun Seo, Yu Cao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 Modeling and Optimization of SRAM-based In-Memory Computing Hardware Design
abstract
In-memory computing (IMC) has been demonstrated as a promising technique to significantly improve energy-efficiency for deep neural network (DNN) hardware accelerators. However, designing one involves setting many design variables such as the number of parallel rows to assert, analog-to-digital converter (ADC) at the periphery of memory sub-array, activation/weight precisions of DNNs, etc., which affect energy-efficiency, DNN accuracy, and area. While individual IMC designs have been presented in the literature, they have not investigated this multi-dimensional design optimization. In this paper, to fill this knowledge gap, we present a SRAM-based IMC hardware modeling and optimization framework. A unified systematic study closely models IMC hardware, and investigates how a number of design variables and nonidealities (e.g. device mismatch and ADC quantization) affect the DNN accuracy of IMC design. To maintain high DNN accuracy for the IMC SRAM hardware, it is shown that the number of activated rows, ADC resolution, ADC quantization range, and different sources of variability/noise need to be carefully selected and co-optimized with an underlying DNN algorithm to implement.
Jyotishman Saikia, Shihui Yin, Sai Kiran Cherupally, Bo Zhang 0105, Jian Meng, Mingoo Seok, Jae-sun Seo
DATE5
2021 FixyFPGA: Efficient FPGA Accelerator for Deep Neural Networks with High Element-Wise Sparsity and without External Memory Access
abstract
Convolutional neural networks (CNNs) have become very popular in real-time computer vision systems. CNNs involve a large amount of computation and storage and typically demand a highly efficient computing platform. Researchers have explored a diverse range of software and hardware optimizations to accelerate CNN inference in recent years. The high power consumption of GPUs and the lack of flexibility with ASIC has promoted interest in FPGAs as a promising platform to efficiently accelerate these CNN inference tasks. Various FPGA-based CNN accelerators have been proposed to low precision weights and high-sparsity in various forms. However, most of the previous work requires off-chip DDR memory to store the parameters and expensive DSP blocks to perform the computation. In this work, we propose the FixyFPGA, a fully on-chip CNN inference accelerator that naturally supports high-sparsity and low-precision computation. In our design, the weights of the trained CNN network are hard-coded into hardware and used as fixed operand for the multiplication. Convolution is performed by streaming the input images to the compute engine in a fully-paralleled, fully-pipelined manner. We analyzed the performance of the proposed scheme with both image classification tasks and object detection tasks based on the low precision, sparse compact CNN models. Compared to prior works, our design achieved 2.34× higher GOPS on ImageNet classification and 3.82× higher frames per second on Pascal VOC object detection.
Jian Meng, Shreyas K. Venkataramanaiah, Chuteng Zhou, Patrick Hansen, Paul N. Whatmough, Jae-sun Seo
FPL1
2021 Algorithm-Hardware Co-Optimization for Energy-Efficient Drone Detection on Resource-Constrained FPGA
abstract
Convolutional neural network (CNN) based object detection has achieved very high accuracy, e.g. single-shot multi-box detectors (SSD) can efficiently detect and localize various objects in an input image. However, they require a high amount of computation and memory storage, which makes it difficult to perform efficient inference on resource-constrained hardware devices such as drones or unmanned aerial vehicles (UAVs). Drone/UAV detection is an important task for applications including surveillance, defense, and multi-drone self-localization and formation control. In this paper, we designed and co-optimized algorithm and hardware for energy-efficient drone detection on resource-constrained FPGA devices. We trained SSD object detection algorithm with a custom drone dataset. For inference, we employed low-precision quantization and adapted the width of the SSD CNN model. To improve throughput, we use dual-data rate operations for DSPs to effectively double the throughput with limited DSP counts. For different SSD algorithm models, we analyze accuracy or mean average precision (mAP) and evaluate the corresponding FPGA hardware utilization, DRAM communication, throughput optimization. Our proposed design achieves a high mAP of 88.42% on the multi-drone dataset, with a high energy-efficiency of 79 GOPS/W and throughput of 158 GOPS using Xilinx Zynq ZU3EG FPGA device on the Open Vision Computer version 3 (OVC3) platform. Our design achieves 2.7X higher energy efficiency than prior works using the same FPGA device, at a low-power consumption of 1.98 W.
Han-Sok Suh, Jian Meng, Ty Nguyen, Shreyas K. Venkataramanaiah, Vijay Kumar 0001, Yu Cao 0001, Jae-sun Seo
FPT2
2020 Compressing LSTM Networks with Hierarchical Coarse-Grain Sparsity
Deepak Kadetotad, Jian Meng, Visar Berisha, Chaitali Chakrabarti, Jae-sun Seo
INTERSPEECH2
2007 Assembly Problem of Overconstrained and Clearance-free Parallel Manipulators
abstract
To avoid deteriorating the mechanism's performance, joint clearance can be eliminated by preloading the pairing elements of the joint. However, this paper proves rigorously that in the real world, the unavoidable assembly and manufacturing errors will cause overconstrained parallel manipulators to lose degree of freedoms, or even unable to be assembled if they are composed of purely clearance-free pairs (e.g., preloaded pairs). Introducing joint clearance is an essential and efficient way for the correct functioning and easy assembly of overconstrained parallel manipulators.
Jian Meng, Dongjun Zhang, Zexiang Li 0001
ICRA1
2007 Accuracy Analysis of General Parallel Manipulators with Joint Clearance
abstract
Due to the joint clearance, parallel manipulators always exhibit some position and orientation errors at the mobile platform. This paper aims to provide a systematic framework for the error analysis problem of general parallel mechanisms influenced by the joint clearance. A novel and efficient method is proposed to evaluate the maximal pose errors of general spatial parallel manipulators with joint clearance.
Jian Meng, Dongjun Zhang, Tinghua Zhang, Zexiang Li 0001
ICRA1
2007 A Geometric Theory for Analysis and Synthesis of Sub-6 DoF Parallel Manipulators
abstract
Mechanism synthesis is mostly dependent on the designer's experience and intuition and is difficult to automate. This paper aims to develop a rigorous and precise geometric theory for analysis and synthesis of sub-6 DoF (or lower mobility) parallel manipulators. Using Lie subgroups and submanifolds of the special Euclidean group${\rm SE}(3)$, we first develop a unified framework for modelling commonly used primitive joints and task spaces. We provide a mathematically rigorous definition of the notion of motion type using conjugacy classes. Then, we introduce a new structure for subchains of parallel manipulators using the product of two subgroups of${\rm SE}(3)$and discuss its realization in terms of the primitive joints. We propose the notion of quotient manipulators that substantially enriches the topologies of serial manipulators. Finally, we present a general procedure for specifying the subchain structures given the desired motion type of a parallel manipulator. The parallel mechanism synthesis problem is thus solved using the realization techniques developed for serial manipulators. Generality of the theory is demonstrated by systematically generating a large class of feasible topologies for (parallel or serial) mechanisms with a desired motion type of either a Lie subgroup or a submanifold.
Jian Meng, Guanfeng Liu 0001, Zexiang Li 0001
IEEE Trans. Robotics1
2006 Finite Motion Validation for Parallel Manipulators: A Differential Geometry Approach
abstract
Type synthesis of low (3-5) degree of freedom (Dof) spatial parallel manipulators is well documented in literature. Recent approaches such as proposed in J.M. Herve and F. Sparacino (1991) - Z. Huang and Q.C. Li (2003) showed some systematic design capability, but did not develop an equally effective means to check for prescribed finite motion. In this paper, we studied the finite motion set of parallel manipulators from a general input-affine nonlinear system viewpoint. Differential geometry tools for controllability (reachability) analysis of nonlinear system on a differential manifold are utilized together with lie group theory. Our techniques are shown to be effective by applying to a systematic type synthesis method proposed in M. Jian, et al. (2005) and W. Yuanqing, et al. (2005)
Yuanqing Wu 0001, Han Ding 0001, Jian Meng, Zexiang Li 0001
IROS3
2005 A Geometric Theory for Synthesis and Analysis of Sub-6 DoF Parallel Manipulators
abstract
This paper presents a rigorous and precise geometric theory for the analysis and synthesis of sub-6 DoF parallel manipulators. We give a rigorous definition for the parallel manipulator synthesis problem, and introduce a general method for specifying the corresponding subchains which will result in the desired parallel manipulator. Following this, a procedure for solving the parallel manipulator synthesis problem is proposed when the set of desired end-effector motions is in the form of Lie subgroup or a regular submanifold of SE(3). Numerous examples are used to illustrate the generality and effectiveness of the proposed synthesis method.
Jian Meng, Guanfeng Liu 0001, Zexiang Li 0001
ICRA1
2005 A Geometric Theory for Synthesis and Analysis of Sub-6 DoF Serial Manipulator Subchains
abstract
Motivated by the work of Herve and his coworkers, this paper presents a rigorous and precise geometric theory for the synthesis and analysis of sub-6 DoF serial manipulator subchains. First, we review the basic properties of the Special Euclidean group SE(3), Lie subgroups and submanifolds of SE(3). With low dimensional subgroups and submanifolds providing models for the so called primitive generators, the high dimensional subgroups and regular submanifolds provide models for the set of desired end-effector motions. Two important classes of regular submanifolds of SE(3) are studied in detail. Then, starting from a given list of primitive generators, we give a rigorous definition of the synthesis problem for a serial manipulator subchain, and develop a general procedure for solving the synthesis problem when the set of desired end-effector motions is a Lie subgroup or a regular submanifold.
Jian Meng, Guanfeng Liu 0001, Zexiang Li 0001
ICRA1
2005 A general approach for accuracy analysis of parallel manipulators with joint clearance
abstract
Due to the joint clearance, parallel manipulators always exhibit some position and orientation errors at the mobile platform. This paper aims to present a novel and general approach for evaluating the maximal pose deviation of the mobile platform under the influence of joint clearance. First, it shows and proves that overconstrained parallel manipulators can not work without clearance. Then, an efficient method is proposed to evaluate the maximal pose errors for general spatial parallel manipulators with joint clearance. A numerical example shows the application and efficiency of the proposed approach.
Jian Meng, Zexiang Li 0001
IROS1
2005 Lie theoretical approach to synthesizing T(3) parallel kinematic manipulators
abstract
Various parallel kinematic manipulator (PKM) type design papers enumerate eligible links as the combination of revolute and prismatic joints and synthesize using local screw theory, but analysis and comparison on real capacity of different types has not been developed yet. This paper applies differential Lie group tools to developing a spectrum of so called regular link spatial translation (T(3)) PKM, which maximized workspace from a topological point of view.
Yuanqing Wu 0001, Han Ding 0001, Jian Meng, Zexiang Li 0001
IROS3
2003 Auto-calibration for a parallel manipulator with sensor redundancy
abstract
In this paper, we propose two algorithms for the auto-calibration of the home position or the joint angle offsets for a parallel manipulator by utilizing the extra sensor(s) information (sensor redundancy), sampling over the workspace, and optimizing a suitably chosen cost function, without resorting to any other external equipment. Meanwhile, a measure or estimate of the precision of the machine is also obtained. It is very useful and convenient if the machine needs frequent re-calibration. Simulations and experiments are also performed to show the effectiveness of the algorithms.
Yiu Kuen Yiu, Jian Meng, Zexiang Li 0001
ICRA2
2003 Kinematic synthesis of parallel manipulators: a Lie theoretic approach
abstract
This paper provided a unified geometric framework for kinematic analysis and synthesis of parallel manipulators. We gave a strict definition on motion types of a mechanism based on distributions on a Lie group. We derived conditions for parallel manipulators with Lie subgroup motions using the intersection of the permissible velocity spaces, or the direct sum of the constraint force spaces of each subchain, and the integration theory on a Lie group. Several practical examples were studied in detail to verify our approach.
Guanfeng Liu 0002, Jian Meng, Jijie Xu, Zexiang Li 0001
IROS2