Yongjune Kim 0001

dblp:124/3256 · DBLP profile ↗
← Back
41ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0003-0120-3750ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Design of Outage-Limit-Approaching Protograph LDPC Codes via Generalized Rootchecks
abstract
This paper presents a new protograph-based LDPC code design framework that simultaneously achieves full diversity over block-fading channels (BFCs) and near-capacity performance over additive white Gaussian noise channels. By leveraging a Boolean approximation-based analysis-Diversity Evolution-we derive structural constraints with generalized rootchecks that guarantee full diversity. Building on these constraints, we propose a diversity-aligned protograph template tailored for the two-block BFC (M=2) that ensures full diversity under iterative belief propagation decoding. Furthermore, a genetic algorithm guided by density evolution is employed to optimize the protograph edges within this family for improved coding gain. The resulting codes, termed DA-GRP-LDPC codes, simultaneously achieve full diversity and enhanced coding gain, reaching a 0.8 dB gap to the outage limit for the two-block BFC at a block length of 16,896. This demonstrates that the proposed framework effectively bridges the gap between diversity optimality in non-ergodic channels and high coding gain in ergodic channels.
Inki Kim, Hyuntae Ahn, Yongjune Kim 0001, Hee-Youl Kwak, Dae-Young Yun, Sang-Hyo Kim
ISIT3
2026 Optimized layerwise approximation for efficient private inference on fully homomorphic encryption
Joon-Woo Lee, Eunsang Lee, Young-Sik Kim, Yongwoo Lee 0002, Yongjune Kim 0001, Jong-Seon No
Neurocomputing6
2025 Regularized Deep Joint Source-Channel Coding for Robust Task-Oriented Semantic Communications
abstract
Semantic communications based on deep joint source-channel coding (JSCC) aim to improve communication efficiency by transmitting only task-relevant information. How-ever, ensuring robustness to the stochasticity of communication channels remains a key challenge in learning-based JSCC. In this paper, we propose a novel regularization technique for learning-based JSCC to enhance robustness against channel noise. The proposed method utilizes the Kullback-Leibler (KL) divergence as a regularizer term in the training loss, measuring the discrepancy between two posterior distributions: one under noisy channel conditions (noisy posterior) and one for a noise-free system (noise-free posterior). We further show that the expectation of the KL divergence given the encoded representation can be analytically approximated using the Fisher information matrix and the covariance matrix of the channel noise. Notably, the proposed regularization is architecture-agnostic, making it broadly applicable to general semantic communication systems over noisy channels.
Taewoo Park, Eunhye Hong, Yo-Seb Jeon, Namyoon Lee, Yongjune Kim 0001
GLOBECOM5
2025 Vision Transformer-Aided Importance-Aware Quantization for Digital Semantic Communications
abstract
Semantic communications provide significant performance gains over traditional communications by transmitting task-relevant semantic features through wireless channels. However, most existing studies rely on end-to-end (E2E) training of neural-type encoders and decoders to ensure effective transmission of these semantic features. To enable semantic communications without relying on E2E training, this paper presents a vision transformer (ViT)-based semantic communication system with importance-aware quantization (IAQ) for wireless image transmission. The core idea of the presented system is to leverage the attention scores of a pretrained ViT model to quantify the importance levels of image patches. Then, our IAQ framework assigns different quantization bits to image patches based on their importance levels. This is achieved by formulating a weighted quantization error minimization problem, where the weight is set to be an increasing function of the attention score. Then, an optimal incremental bit-allocation method and a low-complexity water-filling method are devised to solve the formulated problem. Simulations on multi-view image classification tasks show that our IAQ framework outperforms existing quantization methods.
Joohyuk Park, Yongjeong Oh, Yongjune Kim 0001, Yo-Seb Jeon
ICC3
2025 CrossMPT: Cross-attention Message-passing Transformer for Error Correcting Codes
abstract
Error correcting codes (ECCs) are indispensable for reliable transmission in communication systems. Recent advancements in deep learning have catalyzed the exploration of ECC decoders based on neural networks. Among these, transformer-based neural decoders have achieved state-of-the-art decoding performance. In this paper, we propose a novel Cross-Attention Message-Passing Transformer (CrossMPT), which shares key operational principles with conventional message-passing decoders. While conventional transformer-based decoders employ a self-attention mechanism without distinguishing between magnitude and syndrome embeddings, CrossMPT updates these two types of embeddings separately and iteratively via two masked cross-attention blocks. The mask matrices are determined by the code's parity-check matrix, which explicitly captures and removes irrelevant relationships between the magnitude and syndrome embeddings. Our experimental results show that CrossMPT significantly outperforms existing neural network-based decoders for various code classes. Notably, CrossMPT achieves this decoding performance improvement while significantly reducing memory usage, computational complexity, inference time, and training time.
Seong-Joon Park, Heeyoul Kwak, Sang-Hyo Kim, Yongjune Kim 0001, Jong-Seon No
ICLR4
2025 Quantizing for Noisy Flash Memory Channels
abstract
Flash memory-based processing-in-memory (flashbased PIM) offers high storage capacity and computational efficiency but faces significant reliability challenges due to noise in high-density multi-level cell (MLC) flash memories. Existing verify level optimization methods are designed for general storage scenarios and fail to address the unique requirements of flashbased PIM systems, where metrics such as mean squared error (MSE) and peak signal-to-noise ratio (PSNR) are critical. This paper introduces an integrated framework that jointly optimizes quantization and verify levels to minimize the MSE, considering both quantization and flash memory channel errors. We develop an iterative algorithm to solve the joint optimization problem. Experimental results on quantized images and SwinIR model parameters stored in flash memory show that the proposed method significantly improves the reliability of flash-based PIM systems.
Juyun Oh, Taewoo Park, Jiwoong Im, Yuval Cassuto, Yongjune Kim 0001
ISIT5
2025 Vision Transformer-Based Semantic Communications With Importance-Aware Quantization
abstract
Semantic communications provide significant performance gains over traditional communications by transmitting task-relevant semantic features through wireless channels. However, most existing studies rely on end-to-end (E2E) training of neural-type encoders and decoders to ensure effective transmission of these semantic features. To enable semantic communications without relying on E2E training, this paper presents a vision transformer (ViT)-based semantic communication system with importance-aware quantization (IAQ) for wireless image transmission. The core idea of the presented system is to leverage the attention scores of a pretrained ViT model to quantify the importance levels of image patches. Based on this idea, our IAQ framework assigns different quantization bits to image patches based on their importance levels. This is achieved by formulating a weighted quantization error minimization problem, where the weight is set to be an increasing function of the attention score. Then, an optimal incremental allocation method and a low-complexity water-filling method are devised to solve the formulated problem. Our framework is further extended for realistic digital communication systems by modifying the bit allocation problem and the corresponding allocation methods based on an equivalent binary symmetric channel (BSC) model. Simulations on single-view image classification, multi-view image classification, and single-object detection tasks demonstrate that our IAQ framework outperforms conventional image compression methods under both error-free and realistic communication scenarios.
Joohyuk Park, Yongjeong Oh, Yongjune Kim 0001, Yo-Seb Jeon
IEEE Internet Things J.3
2025 Boosted Neural Decoders: Achieving Extreme Reliability of LDPC Codes for 6G Networks
abstract
Ensuring extremely high reliability in channel coding is essential for 6G networks. The next-generation of ultra-reliable and low-latency communications (xURLLC) scenario within 6G networks requires frame error rate (FER) below 10-9. However, low-density parity-check (LDPC) codes, the standard in 5G new radio (NR), encounter a challenge known as the error floor phenomenon, which hinders to achieve such low frame error rates. To tackle this problem, we introduce an innovative solution: boosted neural min-sum (NMS) decoder. This decoder operates identically to conventional NMS decoders, but is trained by novel training methods including: i) boosting learning with uncorrected vectors, ii) block-wise training schedule to address the vanishing gradient issue, iii) dynamic weight sharing to minimize the number of trainable weights, iv) transfer learning to reduce the required sample count, and v) data augmentation to expedite the sampling process. Leveraging these training strategies, the boosted NMS decoder achieves the state-of-the art performance in reducing the error floor as well as superior waterfall performance. Remarkably, we fulfill the 6G xURLLC requirement for 5G LDPC codes without a severe error floor. Additionally, the boosted NMS decoder, once its weights are trained, can perform decoding without additional modules, making it highly practical for immediate application. The source code is available athttps://github.com/ghy1228/LDPC_Error_Floor.
Heeyoul Kwak, Daeyoung Yun, Yongjune Kim 0001, Sang-Hyo Kim, Jong-Seon No
IEEE J. Sel. Areas Commun.3
2025 Multiple-Masks Error Correction Code Transformer for Short Block Codes
abstract
With the broadening applications of deep learning, neural decoders have emerged as a key research focus, specifically aimed at improving the decoding performance of conventional decoding algorithms. In particular, error correction code transformer (ECCT), which utilizes the transformer architecture, has achieved state-of-the-art performance among neural network-based decoders. We present three technical contributions to significantly enhance the performance of ECCT. First, we propose a novel transformer architecture of ECCT, termed themultiple-masks ECCT (MM ECCT). We employ multiple masked self-attention blocks with different mask matrices in a parallel manner to learn diverse relationships among the codeword bits. Second, we discover that constructing mask matrices based on systematic parity check matrices (PCMs) can make the attention mapssparse, which not only enhances the decoding performance but also reduces computational complexity. Finally, we propose using complementary mask matrices derived from cyclic permutations of the systematic PCM. These complementary mask matrices are specifically designed to enhance the decoding of cyclic codes. Our extensive simulation results show that the proposed MM ECCT architecture with carefully designed mask matrices outperforms the original ECCT by a large margin, achieving state-of-the-art decoding performance among neural decoders. The source code is available at https://github.com/iil-postech/mm-ecct.
Seong-Joon Park, Heeyoul Kwak, Sang-Hyo Kim, Sunghwan Kim 0001, Yongjune Kim 0001, Jong-Seon No
IEEE J. Sel. Areas Commun.5
2025 Hybrid Ordered Statistics Decoding of Short-Length BCH Codes for URLLC Systems: Theoretical Analysis and Decoder Implementation
abstract
The ordered statistics decoding (OSD) algorithm has been gaining popularity for ultra-reliable and low-latency communication (URLLC) scenarios due to its near-maximum likelihood decoding performance, especially for short linear block codes. However, its substantial computational complexity hinders practical applications. In this paper, we introduce an advanced hybrid OSD algorithm that fully utilizes the hard-decision algebraic decoding results to selectively activate soft-decision OSD operations, significantly mitigating computational complexity. Through a rigorous analysis of error-correction characteristics, we derive a theoretical condition under which the hybrid OSD algorithm guarantees superior error-correction performance over the baseline OSD. To apply the proposed hybrid algorithm to the emerging URLLC systems, we also present a novel decoder architecture that efficiently integrates hard-and soft-decision operations. For (127, 64) BCH codes, the prototype decoder in a 28-nm process achieves an average processing latency of 773 ns at a target block error rate of 10-5, improving information throughput by 4.2× and energy-efficiency by 35× and offering coding gain compared to previous OSD hardware designs.
Jaehee Kim, Sangbu Yun, Dongyun Kam, Soonhyun Kwon, Yongjune Kim 0001, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 Attention-Aware Semantic Communications for Collaborative Inference
abstract
We propose a communication-efficient collaborative inference framework in the domain of edge inference, focusing on the efficient use of vision transformer (ViT) models. The partitioning strategy of conventional collaborative inference fails to reduce communication cost because of the inherent architecture of ViTs maintaining consistent layer dimensions across the entire transformer encoder. Therefore, instead of employing the partitioning strategy, our framework utilizes a lightweight ViT model on the edge device, with the server deploying a complicated ViT model. To enhance communication efficiency and achieve the classification accuracy of the server model, we propose two strategies: 1) attention-aware patch selection and 2) entropy-aware image transmission. Attention-aware patch selection leverages the attention scores generated by the edge device’s transformer encoder to identify and select the image patches critical for classification. This strategy enables the edge device to transmit only the essential patches to the server, significantly improving communication efficiency. Entropy-aware image transmission uses min-entropy as a metric to accurately determine whether to depend on the lightweight model on the edge device or to request the inference from the server model. In our framework, the lightweight ViT model on the edge device acts as a semantic encoder, efficiently identifying and selecting the crucial image information required for the classification task. Our experiments demonstrate that the proposed collaborative inference framework can reduce communication overhead by 68% with only a minimal loss in accuracy compared to the server model on the ImageNet dataset.
Jiwoong Im, Nayoung Kwon, Taewoo Park, Jiheon Woo, Jaeho Lee 0001, Yongjune Kim 0001
IEEE Internet Things J.6
2023 Rate-Distortion via Energy-Based Models
abstract
Rate-distortion theory provides a framework for understanding the limits of source coding. Energy-based models (EBMs), which have a broad range of applications in fields such as physics, statistics, and machine learning, can be used to estimate these limits. In this work, we demonstrate how EBMs can be used to estimate rate-distortion functions, and show that our empirical estimates agree with known closed-form expressions and bounds.
Qing Li 0002, Yongjune Kim 0001, Cyril Guyot
DCC2
2023 Boosting Learning for LDPC Codes to Improve the Error-Floor Performance
abstract
Low-density parity-check (LDPC) codes have been successfully commercialized in communication systems due to their strong error correction capabilities and simple decoding process. However, the error-floor phenomenon of LDPC codes, in which the error rate stops decreasing rapidly at a certain level, presents challenges for achieving extremely low error rates and deploying LDPC codes in scenarios demanding ultra-high reliability. In this work, we propose training methods for neural min-sum (NMS) decoders to eliminate the error-floor effect. First, by leveraging the boosting learning technique of ensemble networks, we divide the decoding network into two neural decoders and train the post decoder to be specialized for uncorrected words that the first decoder fails to correct. Secondly, to address the vanishing gradient issue in training, we introduce a block-wise training schedule that locally trains a block of weights while retraining the preceding block. Lastly, we show that assigning different weights to unsatisfied check nodes effectively lowers the error-floor with a minimal number of weights. By applying these training methods to standard LDPC codes, we achieve the best error-floor performance compared to other decoding methods. The proposed NMS decoder, optimized solely through novel training methods without additional modules, can be integrated into existing LDPC decoders without incurring extra hardware costs. The source code is available at https://github.com/ghy1228/LDPC_Error_Floor.
Heeyoul Kwak, Daeyoung Yun, Yongjune Kim 0001, Sang-Hyo Kim, Jong-Seon No
NeurIPS3
2023 Ensemble Classification With Noisy Real-Valued Base Functions
abstract
In data-intensive applications, it is advantageous to perform partial processing close to the data, and communicate intermediate results to a central processor, instead of the data itself. When the communication or computation medium is noisy, the resulting degradation in computation quality at the central processor must be mitigated. We study this problem for the setup of binary classification performed by an ensemble of base functions communicating real-valued confidence levels. We propose a noise-mitigation solution that optimizes the transmission gains and aggregation coefficients of the base functions. Toward that, we formulate a post-training gradient-based optimization algorithm that minimizes the error probability given the training dataset and the noise parameters. We further derive lower and upper bounds on the optimized error probability, and show empirical results that demonstrate the enhanced performance achieved by our approach on real data.
Yuval Ben-Hur, Asaf Goren, Da El Klang, Yongjune Kim 0001, Yuval Cassuto
IEEE J. Sel. Areas Commun.4
2023 Distributed Boosting Classification Over Noisy Communication Channels
abstract
We address the design of inference-oriented communication systems where multiple transmitters send partial inference values through noisy communication channels, and the receiver aggregates these channel outputs to obtain a reliable final inference. Since large data items are replaced by compact inference values, these systems lead to significant savings of communication resources. In particular, we present a principled framework to optimize communication-resource allocation for distributed boosting classifiers. Boosting classification algorithms make a final decision via a weighted vote from the outputs of multiple base classifiers. Since these base classifiers transmit their partial inference values over noisy channels, communication errors would degrade the final classification accuracy. We formulate communication resource allocation problems to maximize the final classification accuracy by taking into account the importance of base classifiers and the resource budget. To solve these problems rigorously, we formulate convex optimization problems to optimize: 1) transmit-power allocations and 2) transmit-rate allocations. This framework departs from classical communication-systems optimizations in seeking to maximize the classification accuracy rather than the reliability of the individual communicated bits. Results from numerical experiments demonstrate the benefits of our approach.
Yongjune Kim 0001, Junyoung Shin, Yuval Cassuto, Lav R. Varshney
IEEE J. Sel. Areas Commun.1
2023 Information and Energy Transmission With Wavelet-Reconstructed Harvesting Functions
abstract
In practical simultaneous information and energy transmission (SIET), the exact energy harvesting function is usually unavailable because an energy harvesting circuit is nonlinear and nonideal. In this work, we consider a SIET problem where the harvesting function is accessible only at experimentally-taken sample points and study how close we can design SIET to the optimal system with such sampled knowledge. Assuming that the harvesting function is of bounded variation that may have discontinuities, we separately consider two settings where samples are taken without and with additive noise. For these settings, we propose to design a SIET system as if a wavelet-reconstructed harvesting function is the true one and study its asymptotic performance loss of energy and information delivery from the true optimal one. Specifically, for noiseless samples, it is shown that designing SIET as if the wavelet-reconstructed harvesting function is the truth incurs asymptotically vanishing energy and information delivery loss with the number of samples. For noisy samples, we propose to reconstruct wavelet coefficients via soft-thresholding estimation. Then, we not only obtain similar asymptotic losses to the noiseless case but also show that the energy loss by wavelets is asymptotically optimal up to a logarithmic factor.
Yongjune Kim 0001
IEEE Trans. Commun.2
2023 A Context-Aware CEO Problem
abstract
In many sensor network applications, a fusion center often has additional valuable information, such as context data, which cannot be obtained directly from the sensors. Motivated by this, we study a generalized CEO problem where a CEO has access to context information. The main contribution of this work is twofold. Firstly, we characterize the asymptotically optimal error exponent per rate as the number of sensors and sum rate grow without bound. The proof extends the Berger-Tung coding scheme and the converse argument by Berger et al. (1996) taking into account context information. The resulting expression includes the minimum Chernoff divergence over context information. Secondly, assuming that the sizes of the source and context alphabets are respectively$|\mathcal {X}|$and$|\mathcal {S}|$, we prove that it is asymptotically optimal to partition all sensors into at most$\binom {|\mathcal {X}|}{2} |\mathcal {S}|$groups and have the sensors in each group adopt the same encoding scheme. Our problem subsumes the original CEO problem by Berger et al. (1996) as a special case if there is only one letter for context information; in this case, our result tightens its required number of groups from$\binom {|\mathcal {X}|}{2}+2$to$\binom {|\mathcal {X}|}{2}$. We also numerically demonstrate the effect of context information for a simple Gaussian scenario.
Sung Hoon Lim, Yongjune Kim 0001
IEEE Trans. Commun.3
2023 Generalized LRS Estimator for Min-Entropy Estimation
abstract
The min-entropy is a widely used metric to quantify the randomness of generated random numbers, which measures the difficulty of guessing the most likely output. It is difficult to accurately estimate the min-entropy of a non-independent and identically distributed (non-IID) source. Hence, NIST Special Publication (SP) 800-90B adopts ten different min-entropy estimators and then conservatively selects the minimum value among ten min-entropy estimates. Among these estimators, the longest repeated substring (LRS) estimator estimates the collision entropy instead of the min-entropy by counting the number of repeated substrings. Since the collision entropy is an upper bound on the min-entropy, the LRS estimator inherently providesoverestimatedoutputs. In this paper, we propose two techniques to estimate the min-entropy of a non-IID source accurately. The first technique resolves the overestimation problem by translating the collision entropy into the min-entropy. Next, we generalize the LRS estimator by adopting the general Rényi entropy instead of the collision entropy (i.e., Rényi entropy of order two). We show that adopting a higher order can reduce the variance of min-entropy estimates. By integrating these techniques, we propose a generalized LRS estimator that effectively resolves the overestimation problem and provides stable min-entropy estimates. Theoretical analysis and empirical results support that the proposed generalized LRS estimator improves the estimation accuracy significantly, which makes it an appealing alternative to the LRS estimator.
Jiheon Woo, Chanhee Yoo, Young-Sik Kim, Yuval Cassuto, Yongjune Kim 0001
IEEE Trans. Inf. Forensics Secur.5
2022 High-Precision Bootstrapping for Approximate Homomorphic Encryption by Error Variance Minimization
Yongwoo Lee 0002, Joon-Woo Lee, Young-Sik Kim, Yongjune Kim 0001, Jong-Seon No, HyungChul Kang
EUROCRYPT (1)4
2022 Low-Complexity Deep Convolutional Neural Networks on Fully Homomorphic Encryption Using Multiplexed Parallel Convolutions
abstract
Recently, the standard ResNet-20 network was successfully implemented on the fully homomorphic encryption scheme, residue number system variant Cheon-Kim-Kim-Song (RNS-CKKS) scheme using bootstrapping, but the implementation lacks practicality due to high latency and low security level. To improve the performance, we first minimize total bootstrapping runtime using multiplexed parallel convolution that collects sparse output data for multiple channels compactly. We also propose the imaginary-removing bootstrapping to prevent the deep neural networks from catastrophic divergence during approximate ReLU operations. In addition, we optimize level consumptions and use lighter and tighter parameters. Simulation results show that we have 4.67x lower inference latency and 134x less amortized runtime (runtime per image) for ResNet-20 compared to the state-of-the-art previous work, and we achieve standard 128-bit security. Furthermore, we successfully implement ResNet-110 with high accuracy on the RNS-CKKS scheme for the first time.
Eunsang Lee, Joon-Woo Lee, Young-Sik Kim, Yongjune Kim 0001, Jong-Seon No, Woosuk Choi
ICML5
2022 Mitigating Noise in Ensemble Classification with Real-Valued Base Functions
abstract
In data-intensive applications, it is advantageous to perform some partial processing close to the data, and communicate to a central processor the partial results instead of the data itself. When the communication medium is noisy, one must mitigate the resulting degradation in computation quality. We study this problem for the setup of binary classification performed by an ensemble of functions communicating real-valued confidence levels. We propose a noise-mitigation solution that works by optimizing the aggregation coefficients at the central processor. Toward that, we formulate a post-training gradient algorithm that minimizes the error probability given the dataset and the noise parameters. We further derive lower and upper bounds on the optimized error probability, and show empirical results that demonstrate the enhanced performance achieved by our scheme on real data.
Yuval Ben-Hur, Asaf Goren, Da El Klang, Yongjune Kim 0001, Yuval Cassuto
ISIT4
2022 Information and Energy Transmission with Wavelet-Reconstructed Harvesting Functions
abstract
In practical simultaneous information and energy transmission (SIET), the exact energy harvesting function is usually unavailable because a harvesting circuit is nonlinear and nonideal. In this work, we consider a SIET problem where the harvesting function is accessible only at sample points that are experimentally taken in the presence of noise. Assuming that the harvesting function is of bounded variation that may have discontinuities, we propose to design SIET based on the wavelet-reconstructed harvesting function. The main focus is its asymptotic performance of expected energy and information rate. Specifically, we propose to design a SIET system based on the wavelet-reconstructed harvesting function with soft-thresholding estimation. Then, the expected loss in energy transmission asymptotically vanishes as the number of samples grows, which turns out to be optimal up to a logarithmic factor. The expected loss in information transmission also vanishes if the target energy delivery is in the interior of the deliverable energy range.
Yongjune Kim 0001
ISIT2
2022 Generalized Longest Repeated Substring Min-Entropy Estimator
abstract
The min-entropy is a widely used metric to quantify the randomness of generated random numbers, which measures the difficulty of guessing the most likely output. It is difficult to accurately estimate the min-entropy of a non-independent and identically distributed (non-IID) source. Hence, NIST Special Publication (SP) 800-90B adopts ten different min-entropy estimators and then conservatively selects the minimum value among ten min-entropy estimates. Among these estimators, the longest repeated substring (LRS) estimator estimates the collision entropy instead of the min-entropy by counting the number of repeated substrings. Since the collision entropy is an upper bound on the min-entropy, the LRS estimator inherently provides overestimated outputs. In this paper, we propose two techniques to estimate the min-entropy of a non-IID source accurately. The first technique resolves the overestimation problem by translating the collision entropy into the min-entropy. Next, we generalize the LRS estimator by adopting the general Rényi entropy instead of the collision entropy (i.e., Rényi entropy of order two). We show that adopting a higher order can reduce the variance of min-entropy estimates. By integrating these techniques, we propose a generalized LRS estimator that effectively resolves the overestimation problem and provides stable min-entropy estimates. Theoretical analysis and empirical results support that the proposed generalized LRS estimator improves the estimation accuracy significantly, which makes it an appealing alternative to the current-standard LRS estimator.
Jiheon Woo, Chanhee Yoo, Young-Sik Kim, Yuval Cassuto, Yongjune Kim 0001
ISIT5
2022 Optimizing Write Fidelity of MRAMs by Alternating Water-Filling Algorithm
abstract
Magnetic random-access memory (MRAM) is a promising memory technology due to its high density, non-volatility, and high endurance. However, achieving high memory fidelity incurs high write-energy costs, which should be reduced for large-scale deployment of MRAMs. In this paper, we formulate abiconvexoptimization problem to optimize write fidelity given energy and latency constraints. The basic idea is to allocate non-uniform write pulses depending on the importance of each bit position. The fidelity measure we consider is mean squared error (MSE), for which we optimize write pulses via alternating convex search (ACS). We derive analytic solutions and propose analternating water-fillingalgorithm by casting the MRAM’s write operation as communication over parallel channels. Hence, the proposed alternating water-filling algorithm is computationally more efficient than the original ACS while their solutions are identical. Since the formulated biconvex problem is non-convex, both the original ACS and the proposed algorithm do not guarantee global optimality. However, the MSEs obtained by the proposed algorithm are comparable to the MSEs by complicated global nonlinear programming solvers. Furthermore, we prove that our algorithm can reduce the MSE exponentially with the number of bits per word. For an 8-bit accessed word, the proposed algorithm reduces the MSE by a factor of 21. We also evaluate MNIST dataset classification supposing that the model parameters of deep neural networks are stored in MRAMs. The numerical results show that the optimized write pulses can achieve 40% write-energy reduction for the same classification accuracy.
Yongjune Kim 0001, Yoocharn Jeon, Hyeokjin Choi, Cyril Guyot, Yuval Cassuto
IEEE Trans. Commun.1
2021 Boosting for Straggling and Flipping Classifiers
abstract
Boosting is a well-known method in machine learning for combining multiple weak classifiers into one strong classifier. When used in distributed setting, accuracy is hurt by classifiers that flip or straggle due to communication and/or computation unreliability. While unreliability in the form of noisy data is well-treated by the boosting literature, the unreliability of the classifier outputs has not been explicitly addressed. Protecting the classifier outputs with an error/erasure-correcting code requires reliable encoding of multiple classifier outputs, which is not feasible in common distributed settings. In this paper we address the problem of training boosted classifiers subject to straggling or flips at classification time. We propose two approaches: one based on minimizing the usual exponential loss but in expectation over the classifier errors, and one by defining and minimizing a new worst-case loss for a specified bound on the number of unreliable classifiers.
Yuval Cassuto, Yongjune Kim 0001
ISIT2
2021 On the Efficient Estimation of Min-Entropy
abstract
The min-entropy is a widely used metric to quantify the randomness of generated random numbers in cryptographic applications; it measures the difficulty of guessing the most likely output. An important min-entropy estimator is thecompression estimatorof NIST Special Publication (SP) 800-90B, which relies on Maurer’s universal test. In this paper, we propose two kinds of min-entropy estimators to improve computational complexity and estimation accuracy by leveraging two variations of Maurer’s test: Coron’s test (for Shannon entropy) and Kim’s test (for Rényi entropy). First, we propose a min-entropy estimator based on Coron’s test. It is computationally more efficient than the compression estimator while maintaining the estimation accuracy. The secondly proposed estimator relies on Kim’s test that computes the Rényi entropy. This estimator improves estimation accuracy as well as computational complexity. We analytically characterize the bias-variance tradeoff, which depends on the order of Rényi entropy. By taking into account this tradeoff, we observe that the order of two is a proper assignment and focus on the min-entropy estimation based on the collision entropy (i.e., Rényi entropy of order two). The min-entropy estimation from the collision entropy can be described by a closed-form solution, whereas both the compression estimator and the proposed estimator based on Coron’s test do not have closed-form solutions. By leveraging the closed-form solution, we also propose a lightweight estimator that processes data samples in an online manner. Numerical evaluations demonstrate that the first proposed estimator achieves the same accuracy as the compression estimator with much less computation. The proposed estimator based on the collision entropy can even improve the accuracy and reduce the computational complexity.
Yongjune Kim 0001, Cyril Guyot, Young-Sik Kim
IEEE Trans. Inf. Forensics Secur.1
2020 Optimizing the Write Fidelity of MRAMs
abstract
Magnetic random-access memory (MRAM) is a promising memory technology due to its high density, non-volatility, and high endurance. However, achieving high memory fidelity incurs significant write-energy costs, which should be reduced for the large-scale deployment of MRAMs. In this paper, we formulate an optimization problem to maximize the memory fidelity given energy constraints, and propose a biconvex optimization approach to solve it. The basic idea is to allocate non-uniform write pulses depending on the importance of each bit position. We consider the mean squared error (MSE) as a fidelity metric and propose an iterative water-filling algorithm to minimize the MSE. Although the iterative algorithm does not guarantee the global optimality, we can choose a proper starting point that decreases the MSE exponentially and guarantees fast convergence. For an 8-bit accessed word, the proposed algorithm reduces the MSE by a factor of 21.
Yongjune Kim 0001, Yoocharn Jeon, Cyril Guyot, Yuval Cassuto
ISIT1
2020 Compression by and for Deep Boltzmann Machines
abstract
We answer two questions in this work: what Deep Boltzmann Machines (DBMs) can do for compression and vise versa. We show that (1) DBMs can be applied to learn the rate distortion approaching posterior as in the Blahut-Arimoto (BA) algorithm, and to construct a lossy source compression scheme based on the Deep AutoEncoder; (2) compression can improve DBMs’ training performances via compression-based denoising algorithms. The implementation of the BA algorithm in the form of DBMs is the foundation of the two applications.
Qing Li 0002, Yang Chen 0063, Yongjune Kim 0001
IEEE Trans. Commun.3
2019 On the Optimal Refresh Power Allocation for Energy-Efficient Memories
abstract
Refresh is an important operation to prevent loss of data in dynamic random-access memory (DRAM). However, frequent refresh operations incur considerable power consumption and degrade system performance. Refresh power cost is especially significant in high-capacity memory devices and battery-powered edge/mobile applications. In this paper, we propose a principled approach to optimizing the refresh power allocation. Given a model for the bit error rate dependence on power, we formulate a convex optimization problem to minimize the word mean squared error for a refresh power constraint; hence we can guarantee the optimality of the obtained refresh power allocations. In addition, we provide an integer programming problem to optimize the discrete refresh interval assignments. For an 8-bit accessed word, numerical results show that the optimized nonuniform refresh intervals reduce the refresh power by 29% at a peak signal-to-noise ratio of 50dB compared to the uniform assignment.
Yongjune Kim 0001, Won Ho Choi, Cyril Guyot, Yuval Cassuto
GLOBECOM1
2019 Shannon-Inspired Statistical Computing for the Nanoscale Era
abstract
Modern day computing systems are based on the von Neumann architecture proposed in 1945 but face dual challenges of: 1) unique data-centric requirements of emerging applications and 2) increased nondeterminism of nanoscale technologies caused by process variations and failures. This paper presents a Shannon-inspired statistical model of computation (statistical computing) that addresses the statistical attributes of both emerging cognitive workloads and nanoscale fabrics within a common framework. Statistical computing is a principled approach to the design of non-von Neumann architectures. It emphasizes the use of information-based metrics; enables the determination of fundamental limits on energy, latency, and accuracy; guides the exploration of statistical design principles for low signal-to-noise ratio (SNR) circuit fabrics and architectures such as deep in-memory architecture (DIMA) and deep in-sensor architecture (DISA); and thereby provides a framework for the design of computing systems that approach the limits of energy efficiency, latency, and accuracy. From its early origins, Shannon-inspired statistical computing has grown into a concrete design framework validated extensively via both theory and laboratory prototypes in both CMOS and beyond. The framework continues to grow at both of these levels, yielding new ways of connecting systems through architectures, circuits, and devices, for the semiconductor roadmap to march into the nanoscale era.
Naresh R. Shanbhag, Naveen Verma, Yongjune Kim 0001, Ameya Patil 0001, Lav R. Varshney
Proc. IEEE3
2018 Energy-Efficient Deep In-memory Architecture for NAND Flash Memories
abstract
This paper proposes an energy-efficient deep in-memory architecture for NAND flash (DIMA-F) to perform machine learning and inference algorithms on NAND flash memory. Algorithms for data analytics, inference, and decision-making require processing of large data volumes and are hence limited by data access costs. DIMA-F achieves energy savings and throughput improvement for such algorithms by reading and processing data in the analog domain at the periphery of NAND flash memory. This paper also provides behavioral models of DIMA-F that can be used for analysis and large scale system simulations in presence of circuit non-idealities and variations. DIMA-F is studied in the context of linear support vector machines and k-nearest neighbor for face detection and recognition, respectively. An estimated 8×-to-23× reduction in energy and 9×-to-15× improvement in throughput resulting in EDP gains up to 345× over the conventional NAND flash architecture incorporating an external digital ASIC for computation.
Sujan K. Gonugondla, Mingu Kang, Yongjune Kim 0001, Mark Helm, Sean Eilert, Naresh R. Shanbhag
ISCAS3
2018 SRAM Bit-line Swings Optimization using Generalized Waterfilling
abstract
We propose an information-theoretic approach to optimize non-uniform bit-line swings for static random access memories (SRAMs). We formulate convex optimization problems whose objectives are to minimize energy (for low-power SRAMs), maximize speed (for high-speed SRAMs), and minimize energy-delay product for a given constraint on mean squared error of retrieved words. We show that these optimization problems can be interpreted as generalized water-filling including classical waterfilling, ground-flattening and water-filling, and sand-pouring and water-filling, respectively. Numerical results show that energy-optimal swing assignment reduces energy consumption by half at a peak signal-to-noise ratio of 30dB for an 8-bit accessed word.
Yongjune Kim 0001, Mingu Kang, Lav R. Varshney, Naresh R. Shanbhag
ISIT1
2018 Generalized Water-Filling for Source-Aware Energy-Efficient SRAMs
abstract
Conventional low-power static random access memories (SRAMs) reduce read energy by decreasing the bit-line voltage swings uniformly across the bit-line columns. This is because the read energy is proportional to the bit-line swings. On the other hand, bit-line swings are limited by the need to avoid decision errors especially in the most significant bits. We propose a principled approach to determine optimal non-uniform bit-line swings by formulating convex optimization problems. For a given constraint on mean squared error of retrieved words, we consider criteria to minimize energy (for low-power SRAMs), maximize speed (for high-speed SRAMs), and minimize energy-delay product. These optimization problems can be interpreted as classical water-filling, ground-flattening and water-filling, and sand-pouring and water-filling, respectively. By leveraging these interpretations, we also propose greedy algorithms to obtain optimized discrete swings. Numerical results show that energy-optimal swing assignment reduces energy consumption by half at a peak signal-to-noise ratio of 30 dB for an 8-bit accessed word. The energy savings increase to four times for a 16-bit accessed word.
Yongjune Kim 0001, Mingu Kang, Lav R. Varshney, Naresh R. Shanbhag
IEEE Trans. Commun.1
2017 Minimum precision requirements for the SVM-SGD learning algorithm
abstract
It is well-known that the precision of data, weight vector, and internal representations employed in learning systems directly impacts their energy, throughput, and latency. The precision requirements for the training algorithm are also important for systems that learn on-the-fly. In this paper, we present analytical lower bounds on the precision requirements for the commonly employed stochastic gradient descent (SGD) on-line learning algorithm in the specific context of a support vector machine (SVM). These bounds are obtained subject to desired system performance. These bounds are validated using the UCI breast cancer dataset. Additionally, the impact of these precisions on the energy consumption of a fixed-point SVM with on-line training is studied. Simulation results in 45 nm CMOS process show that operating at the minimum precision as dictated by our bounds improves energy consumption by a factor of 5.3× as compared to conventional precision assignments with no observable loss in accuracy.
Charbel Sakr, Ameya Patil 0001, Yongjune Kim 0001, Naresh R. Shanbhag
ICASSP4
2017 Analytical Guarantees on Numerical Precision of Deep Neural Networks
abstract
The acclaimed successes of neural networks often overshadow their tremendous complexity. We focus on numerical precision – a key parameter defining the complexity of neural networks. First, we present theoretical bounds on the accuracy in presence of limited precision. Interestingly, these bounds can be computed via the back-propagation algorithm. Hence, by combining our theoretical analysis and the back-propagation algorithm, we are able to readily determine the minimum precision needed to preserve accuracy without having to resort to time-consuming fixed-point simulations. We provide numerical evidence showing how our approach allows us to maintain high accuracy but with lower complexity than state-of-the-art binary networks.
Charbel Sakr, Yongjune Kim 0001, Naresh R. Shanbhag
ICML2
2017 PredictiveNet: An energy-efficient convolutional neural network via zero prediction
abstract
Convolutional neural networks (CNNs) have gained considerable interest due to their record-breaking performance in many recognition tasks. However, the computational complexity of CNNs precludes their deployments on power-constrained embedded platforms. In this paper, we propose predictive CNN (PredictiveNet), which predicts the sparse outputs of the non-linear layers thereby bypassing a majority of computations. PredictiveNet skips a large fraction of convolutions in CNNs at runtime without modifying the CNN structure or requiring additional branch networks. Analysis supported by simulations is provided to justify the proposed technique in terms of its capability to preserve the mean square error (MSE) of the nonlinear layer outputs. When applied to a CNN for handwritten digit recognition, simulation results show that PredictiveNet can reduce the computational cost by a factor of 2.9χ compared to a state-of-the-art CNN, while incurring marginal accuracy degradation.
Yingyan (Celine) Lin, Charbel Sakr, Yongjune Kim 0001, Naresh R. Shanbhag
ISCAS3
2016 Locally rewritable codes for resistive memories
abstract
We propose locally rewritable codes (LWC) for resistive memories inspired by locally repairable codes (LRC) for distributed storage systems. Small values of repair locality of LRC enable fast repair of a single failed node since the lost data in the failed node can be recovered by accessing only a small fraction of other nodes. By using rewriting locality, LWC can improve endurance and power consumption which are major challenges for resistive memories. We point out the duality between LRC and LWC, which indicates that existing construction methods of LRC can be applied to construct LWC.
Yongjune Kim 0001, Abhishek A. Sharma, Robert Mateescu, Seung-Hwan Song, Zvonimir Bandic, James A. Bain, B. V. K. Vijaya Kumar
ICC1
2016 Locally Rewritable Codes for Resistive Memories
abstract
Resistive memories, such as phase change memories and resistive random access memories, have attracted significant research interest because of their scalability, non-volatility, fast speed, and rewritability. However, their write endurance needs to be improved substantially for large-scale deployment of resistive memories. In addition, their write power consumption is much higher than the power consumption of read operation. Inspired by locally repairable codes (LRCs) recently introduced for distributed storage systems, we propose locally rewritable codes (LWCs) for resistive memories. We define a novel parameter ofrewriting locality, which can be connected torepair localityof LRC. As small values of repair locality of LRC enable fast repair in distributed storage systems, small values of rewriting locality of LWC are able to reduce the problems of write endurance and write power consumption. We show how a small value of rewriting locality can improve write endurance and power consumption by deriving the upper bounds on writing cost. Also, we point out the dual relation of LRC and LWC, which indicates that the existing construction methods of LRC can be applied to construct LWC. Finally, we investigate the construction of LWC with error correcting capability for random errors.
Yongjune Kim 0001, Abhishek A. Sharma, Robert Mateescu, Seung-Hwan Song, Zvonimir Bandic, James A. Bain, B. V. K. Vijaya Kumar
IEEE J. Sel. Areas Commun.1
2015 Coding scheme for 3D vertical flash memory
abstract
Recently introduced 3D vertical flash memory is expected to be a disruptive technology since it overcomes scaling challenges of conventional 2D planar flash memory by stacking up cells in the vertical direction. However, 3D vertical flash memory suffers from a new problem known as fast detrapping, which is a rapid charge loss problem. In this paper, we propose a scheme to compensate the effect of fast detrapping by intentional inter-cell interference (ICI). In order to properly control the intentional ICI, our scheme relies on a coding technique that incorporates the side information of fast detrapping during the encoding stage. This technique is closely connected to the well-known problem of coding in a memory with defective cells. Numerical results show that the proposed scheme can effectively address the problem of fast detrapping.
Yongjune Kim 0001, Robert Mateescu, Seung-Hwan Song, Zvonimir Bandic, B. V. K. Vijaya Kumar
ICC1
2013 Coding for memory with stuck-at defects
abstract
In this paper, we propose an encoding scheme for partitioned linear block codes (PLBC) which mask the stuck-at defects in memories. In addition, we derive an upper bound and the estimate of the probability that masking fails. Numerical results show that PLBC can efficiently mask the defects with the proposed encoding scheme. Also, we show that our upper bound is very tight by using numerical results.
Yongjune Kim 0001, B. V. K. Vijaya Kumar
ICC1
2013 Redundancy allocation of partitioned linear block codes
abstract
Most memories suffer from both permanent defects and intermittent random errors. The partitioned linear block codes (PLBC) were proposed by Heegard to efficiently mask stuck-at defects and correct random errors. The PLBC have two separate redundancy parts for defects and random errors. In this paper, we investigate the allocation of redundancy between these two parts. The optimal redundancy allocation will be investigated using simulations and the simulation results show that the PLBC can significantly reduce the probability of decoding failure in memory with defects. In addition, we will derive the upper bound on the probability of decoding failure of PLBC and estimate the optimal redundancy allocation using this upper bound. The estimated redundancy allocation matches the optimal redundancy allocation well.
Yongjune Kim 0001, B. V. K. Vijaya Kumar
ISIT1