Kee-Bong Song

dblp:33/4194 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Mobile-StereoHPE: Real-Time Mobile 3D Hand Pose Estimation from Stereo Gray Images
abstract
Mobile XR scenarios pose significant challenges for 3D bare-hand interaction, requiring accurate and low-latency 3D hand tracking within strict computational constraints. Monocular RGB-based methods struggle with estimating absolute 3D hand joints, while depth sensor or large-model approaches are unsuitable for mobile devices. We propose Mobile-StereoHPE, a lightweight and efficient framework for absolute 3D hand pose estimation using stereo gray images. Our approach introduces two novel components: a Feature Fusion Module (FFM) that aggregates stereo image features while minimizing computational overhead, and a Cross Feature Attention (CFA) module that enhances inter-view correspondence learning. These innovations empower our compact CNN backbone to achieve state-of-the-art accuracy on the MVHand benchmark with the smallest model size, a 33% improvement in 3D joint accuracy, and more than twice the speed of current large-model methods. Extensive experiments validate Mobile-StereoHPE as an ideal solution for next-generation XR applications.
Dongfang Zhao 0017, Menghe Zhang, Yangwen Liang, Shuangquan Wang, Kee-Bong Song
ICME5
2025 3D-AMTA: Occlusion-Aware Real-Time 3D Hand Pose Estimation with Auto Mask and Token-Specific Attention
abstract
Understanding hand motion from a single RGB image is challenging due to occlusions and high articulation. This paper presents 3D-AMTA, a transformer-based framework with Auto Mask and Token-specific Attention for occlusion-aware 3D hand pose estimation (HPE). We propose two novel architectural enhancements: auto mask for high-occlusion scenarios, and token-specific attention for fine-grained hand articulations. These modules seamlessly integrate into transformer-based architectures that enhance real-time performance in interactive systems. To enable efficient deployment on robotic and embedded platforms, we propose 3D-AMTA-Mobile, a lightweight variant optimized for on-device processing. It achieves 267 FPS on NVIDIA RTX 2080Ti-GPU while maintaining high accuracy, making it well-suited for resource-constrained robotic applications. Extensive evaluations on FreiHAND and HO3D demonstrate that our approach consistently outperforms state-of-the-art methods in terms of accuracy, efficiency, and inference speed. These advancements contribute to robust hand perception for interactive robotics and AR-based teleoperation.
Dongfang Zhao 0017, Menghe Zhang, Yangwen Liang, Shuangquan Wang, Kee-Bong Song
IROS5
2025 Leveraging Latent Diffusion in 3D Gaussian Splatting for Novel View Synthesis
Yangwen Liang, Shuangquan Wang, Kee-Bong Song
MMM (5)5
2024 Knowledge Distillation for Tiny Speech Enhancement with Latent Feature Augmentation
Behnam Gholami, Mostafa El-Khamy, Kee-Bong Song
INTERSPEECH3
2023 Domain invariant regularization by disentangling content and style Features for visual domain generalization
abstract
In this paper, taking the advantage of multiple source domains, we propose a novel approach for visual Domain Generalization (DG). The three key ideas underlying our formulation are (1) leveraging disentangled representations of the images to define different factors of variations, (2) generating perturbed images by changing such factors composing the representations of the images, (3) enforcing the learner (classifier) to be invariant to such changes in the images. We demonstrate the effectiveness of our approach on several widely used datasets for the domain generalization problem, on all of which we achieve competitive results with state-of-the-art models.
Behnam Gholami, Mostafa El-Khamy, Kee-Bong Song
ICIP3
2023 Latent Feature Disentanglement for Visual Domain Generalization
abstract
Despite remarkable success in a variety of computer vision applications, it is well-known that deep learning can fail catastrophically when presented with out-of-distribution data, where there are usually style differences between the training and test images. Toward addressing this challenge, we consider the domain generalization problem, wherein predictors are trained using data drawn from a family of related training (source) domains and then evaluated on a distinct and unseen test domain. Naively training a model on the aggregate set of data (pooled from all source domains) has been shown to perform suboptimally, since the information learned by that model might be domain-specific and generalizes imperfectly to test domains. Data augmentation has been shown to be an effective approach to overcome this problem. However, its application has been limited to enforcing invariance to simple transformations like rotation, brightness change, etc. Such perturbations do not necessarily cover plausible real-world variations that preserve the semantics of the input (such as a change in the image style). In this paper, taking the advantage of multiple source domains, we propose a novel approach to express and formalize robustness to these kind of real-world image perturbations. The three key ideas underlying our formulation are (1) leveraging disentangled representations of the images to define different factors of variations, (2) generating perturbed images by changing such factors composing the representations of the images, (3) enforcing the learner (classifier) to be invariant to such changes in the images. We use image-to-image translation models to demonstrate the efficacy of this approach. Based on this, we propose a domain-invariant regularization (DIR) loss function that enforces invariant prediction of targets (class labels) across domains which yields improved generalization performance. We demonstrate the effectiveness of our approach on several widely used datasets for the domain generalization problem, on all of which our results are competitive with the state-of-the-art.
Behnam Gholami, Mostafa El-Khamy, Kee-Bong Song
IEEE Trans. Image Process.3
2022 DeepGBASS: Deep Guided Boundary-Aware Semantic Segmentation
abstract
Image semantic segmentation is ubiquitously used in scene understanding applications, such as AI Camera, which require high accuracy and efficiency. Deep learning has significantly advanced the state-of-the-art in semantic segmentation. However, many of recent semantic segmentation works only consider class accuracy and ignore the accuracies at the boundaries between semantic classes. To improve the semantic boundary accuracy, we propose low complexity Deep Guided Decoder (DGD) networks, trained with a novel Semantic Boundary-Aware Learning (SBAL) strategy. Our ablation studies on Cityscapes and the ADE20K-32 confirm the effectiveness of our approach with network of different complexities. We show that our DeepGBASS approach significantly improves the mIoU by up to 11% relative gain and the mean boundary F1-score (mBF) by up to 39.4% when training MobileNetEdgeTPU DeepLab on ADE20K-32 dataset.
Qingfeng Liu, Hai Su, Mostafa El-Khamy, Kee-Bong Song
ICASSP4
2021 Learning-aided joint time-frequency channel estimation for 5G new radio
abstract
In this paper, we propose a learning-aided signal processing solution for channel estimation in 5G new radio (NR). Channel estimation is an important algorithm for baseband modem design. In 5G NR, estimating the channel is challenging due to two reasons. First, the pilot signals are transmitted over a small fraction of the available time-frequency resources. Second, the real time nature of physical layer processing introduces a strict limitation on the computational complexity of channel estimation. To this end, we propose a channel estimation technique that integrates a small one hidden layer neural network between two linear minimum mean squared error (LMMSE) interpolation blocks. While the neural network leverages the advantages of offline data-driven learning, the LMMSE blocks exploit the second order online channel statistics along time and frequency dimensions. The training procedure tunes the weights of the neural network by back-propagating through the time domain LMMSE interpolation block. We derive bounds on the training loss with the proposed method and show that our approach can improve the channel estimate.
Nitin Jonathan Myers, Hyukjoon Kwon, Yacong Ding, Kee-Bong Song
GLOBECOM4
2021 Line-of-Sight Communications with Antenna Misalignments
abstract
Line-of-sight (LOS) communications are becoming one of the promising use cases in the near future’s wireless communications thanks to the potential use of very high carrier frequency, for example tera-hertz band. However, most traditional communication techniques have been developed assuming abundant reflected multipaths, and there is relatively not enough in-depth study regarding LOS communications. Therefore, this paper considers LOS communications from the aspect of antenna misalignments assuming multiple-in multiple-out (MIMO) is available. We propose a communication system that utilizes antenna subarrays for estimating the antenna misalignments and other necessary parameters to achieve the best performance in LOS channels. This information is fed back to the transmitter for deriving the optimal precoder. We exploit multiple techniques for enhancing the estimation performance for stable system performance. The results indicate that the performance with ideal information can be closely achieved with the proposed techniques. The proposed techniques are also shown to be flexibly applicable to any shape of Tx antenna arrays.
Jangwook Moon, Hongbing Cheng, Kee-Bong Song
ICC3
2021 A Maximum Likelihood Detection Method for NR Sidelink SSS Searcher
abstract
One of the major challenges in vehicle to everything (V2X) system is robust and low-cost synchronization for ultra-reliable low latency communication under various fading scenarios. To address this issue, this paper presents a maximum likelihood (ML) based algorithm for the sidelink secondary synchronization signal (S-SSS) detection, assuming the sidelink primary synchronization signal (S-PSS) detection has been successfully achieved at an earlier stage. Based on the specific signal structure, the proposed method exploits the channel selectivity in frequency domain (FD) and channel correlation in time domain (TD) to obtain an ML solution. In order to avoid the need to estimate the Doppler frequency and simplify the algorithm, we provide the ML solutions based on the assumptions of infinite or zero Doppler. Furthermore, we propose a practical method with limited TD channel taps that can be efficiently implemented in real systems. Simulation results show that the proposed method significantly improves the performance over the conventional detection methods under different Doppler scenarios.
Sili Lu, Hongbing Cheng, Kee-Bong Song
VTC Spring3
2021 Channel Recovery Using History Information for Hybrid Beamforming Systems
abstract
Millimeter-wave channels are angular sparse. Under practical channels, angle of arrivals (AoAs) tend to change slowly over time, while the path gain corresponding to each AoA may vary relatively fast. To exploit the property of slow AoAs variation, this paper proposes two approaches that consider measurements of current and multiple previous beam sweeping periods to improve channel recovery quality at the user equipment (UE). Given slow AoAs variation, the first approach suggests that UE update the beam sweeping codebook based on estimated AoAs to improve signal to noise ratio for the following beam sweeping. When path gains also vary slowly, the second approach suggests that UE estimate channel AoAs using both current and history measurements, and estimate path gains only using current measurement. The first approach is more robust to channel variation since it only assumes slow AoAs variation. When channel variation is small, the second approach exhibits advantage over the first one, and both outperform closed form simultaneous orthogonal matching pursuit (SOMP) proposed in [1], which only uses current measurement for channel recovery. When channel variation is relatively large, codebook update still brings gains over closed form SOMP.
Yanru Tang, Hongbing Cheng, Kee-Bong Song
VTC Spring3
2020 MIMO Detector Selection for Multiple High-Order Modulations with Unified Neural Network
abstract
We propose a unified multi-layer perceptron (MLP) network to select an appropriate multiple-input, multiple-output (MIMO) detector for high-order quadrature amplitude modulations (QAMs) such as 64-, 256-, and 1024-QAM. The network is trained to select a low-complexity detector dynamically for 64-, 256-, and 1024-QAM from a set of candidate detectors, while simultaneously maintaining a block error rate (BLER) close to the highest complexity candidate detector. We train the network on a combined data set collected under various environments such as multiple modulation orders, channel profiles, and signal-to-noise ratios. For offline training, a selective back-propagation method is proposed wherein samples of specific$M$-QAM are used to update modulation specific weights between the hidden layer and the output layer of the MLP network. The common weights between the input layer and the hidden layer are updated using all the samples in the data set. Thus, a single unified network can be deployed for multiple modulation orders together. The performance is evaluated with different channel profiles including channels for which the network was not trained. Simulation results show that the proposed algorithm maintains the BLER close to that of the most complex candidate detector even under un-trained channel profiles. The proposed algorithm also reduces the computational complexity of the MIMO detection block up to 10×.
Shailesh Chaudhari, HyukJoon Kwon, Kee-Bong Song
GLOBECOM3
2020 Antenna Location Design for Line-of-Sight Communications
abstract
The antenna location design problem is addressed in line-of-sight channel with superimposed-concentric uniform circular arrays (SC-UCA). The analysis is done by first, converting the capacity maximization to the channel correlation minimization problem, second, calculating channel correlation terms explicitly that need to be minimized, and third, calculating necessary conditions to achieve minimum correlations. The analysis is performed for 4x4 channel matrix, and a generic case with relative antenna rotation is handled and the closed-form solutions are obtained. As another special case, the modified antenna configuration with inner-circle rotation is also considered. With this model, it is shown that the existing results for uniform linear array (ULA) and uniform circular array (UCA) can be obtained. Moreover, it is also shown that it is possible to provide maximum multiple-in multiple-out (MIMO) gain with either non-uniform linear or circular arrays. With detailed analysis and simulations, the flexible design of antenna locations with multiple concentric circular arrays are presented.
Jangwook Moon, Hongbing Cheng, Kee-Bong Song
VTC Fall3
2020 Low Complexity Channel Estimation for Hybrid Beamforming Systems
abstract
This paper considers the problem of channel estimation in hybrid beamforming systems at the user equipment (UE). After the beam sweeping process, UE estimates the channel such that the best beamforming vector can be derived accordingly to improve analog beamforming gain. To exploit the angular sparsity of millimeter-wave (mmWave) channels, a compressed sensing (CS) based channel estimation algorithm, termed as closed form simultaneous orthogonal matching pursuit (SOMP), is proposed. The proposed algorithm works for a class of beamforming codebooks and scenarios with very small number of beam sweeping symbols, with reduced complexity compared with the existing SOMP algorithm [1]. Simulation results show that the proposed algorithm achieves similar performance as that of the SOMP algorithm, and has advantages over the non-CS based channel estimation algorithm.
Yanru Tang, Hongbing Cheng, Kee-Bong Song
VTC Spring3
2020 Learning based Dynamic Codebook Selection for Analog Beamforming
abstract
In this paper, a dynamic codebook selection scheme is proposed for beam sweeping at millimeter wave receiver. The codebook selection problem is formulated as a Partially Observed Markov Decision Process (POMDP) and Q-learning method is used to learn a policy to dynamically select an appropriate codebook from a pre-defined codebook set in each beam sweeping period. A hierarchical structure codebook set design method based on Lloyd's method is proposed to generate a proper codebook set with good coverage and high average beamforming gain. The proposed scheme offers more than 1 dB gain in terms of block error rate (BLER) performance compared with the existing methods which use a fixed codebook in all beam sweeping periods.
Qi Zhan, Hongbing Cheng, Kee-Bong Song
VTC Fall3
2019 Offset min-sum Optimization for General Decoding Scheduling: A Deep Learning Approach
abstract
Deep learning has shown an unprecedented success in many fields such as computer vision and speech recognition, providing solutions to intractable problems. Deep learning has also provided solutions to several intractable problems in communications systems. This paper uses a deep learning approach to optimize the analytically intractable offset value in the offset-min-sum (OMS) algorithm. OMS algorithm is a very attractive low complexity algorithm that is used in the belief propagation decoding of linear codes. The contributions of this paper are: First, providing a low complexity offset optimization framework based on gradient descent and back propagation on the original Tanner graph. Our proposed algorithm has comparable complexity and similar operation as the forward belief propagation decoding algorithm and hence, can be trained much more efficiently. Second, the framework can be easily extended to any decoding scheduling such as flooding or layered scheduling. Training results show that the proposed framework can find the optimal offset value under different decoding scheduling with the same complexity of the belief propagation algorithm.
Ahmed Abotabl, Jung Hyun Bae, Kee-Bong Song
VTC Fall3
2006 Adaptive modulation and coding (AMC) for bit-interleaved coded OFDM (BIC-OFDM)
abstract
This paper proposes adaptive modulation and coding (AMC) as a method for the bit-interleaved coded OFDM (BIC-OFDM) packet transmission. Following a pair-wise error probability (PEP) analysis of AMC-BIC-OFDM in a slowly fading frequency selective channel, AMC scheme maximizes the total rate by optimally selecting the code and efficiently allocating rate and power over the frequency band. The proposed method improves upon the performance of uniform rate and power allocation scheme by 6.5 to 10 dB
Kee-Bong Song, Amal Ekbal, Seong Taek Chung, John M. Cioffi
IEEE Trans. Wirel. Commun.1
2005 Single user random beamforming in Gaussian MIMO broadcast channels
abstract
This paper proposes a new transmitting scheme for a multi-antenna (MIMO) Gaussian broadcast channel. The transmitter exploits multiuser diversity using a small amount of feedback information from the receivers. The feedback information consists of the maximum achievable rate for each user and the necessary power distribution profile across the spatial dimensions that the transmitter should use to achieve that rate. The receivers use the iterative water-filling algorithm to maximize the achievable rate. The new scheme serves a single user at each transmission, hence the name single user random beamforming (SUBF) scheme. The proposed scheme maximizes the average sum rate under a TDMA environment (i.e., supporting a single user at a time) with the partial channel state information (CSI) feedback constraint.
Kee-Bong Song, Ravi Narasimhan, John M. Cioffi
ICC2
2004 QoS-constrained physical layer optimization for correlated flat-fading wireless channels
abstract
In next generation data networks, joint optimization of physical layer parameters and scheduling layer parameters will be necessary to meet the demands of quality of service (QoS)-constrained traffic like video traffic. Until now, such optimization problems considered only simple and non-realistic Markov chain models to represent physical layer channel dynamics. In this paper, we consider the incorporation of realistic channel models in a Markov decision process (MDP) formulation for the QoS-constrained optimization of the physical layer. The channel random process is transformed into an extended Markovian setup through hidden Markov models (HMM) and the solution of the resulting optimization problem is shown to be a partially observable Markov decision process (POMDP). The optimal scheduling agent bases its decision on a parameter that summarizes complete system history until the decision instant. This parameter can he computed at each decision instant based only on newly available information, avoiding the need to record all the past states.
Amal Ekbal, Kee-Bong Song, John M. Cioffi
ICC2
2004 Adaptive modulation and coding (AMC) for bit-interleaved coded OFDM (BIC-OFDM)
abstract
Adaptive modulation and coding (AMC) is receiving increasing attention as an effective method to enhance the performance of wireless systems. This paper proposes AMC as a method for the bit-interleaved coded OFDM (BIC-OFDM) packet transmission. We provide a pair-wise error probability (PEP) analysis of AMC-BIC-OFDM in a slowly fading frequency selective channel. We maximize the total rate by optimally selecting the code and efficiently allocating rate and power over the frequency band. The proposed method improves the performance of uniform rate and power allocation scheme by 8 to 19 dB.
Kee-Bong Song, Amal Ekbal, Seong Taek Chung, John M. Cioffi
ICC1
2004 Rate-compatible punctured convolutionally (RCPC) space-frequency bit-interleaved coded modulation (SF-BICM)
abstract
This paper proposes a low-complexity space-frequency bit-interleaved coded modulation (SF-BICM) transceiver with successive interference cancelling (SIC) processing for multiinput multioutput (MIMO) orthogonal frequency division multiplexing (OFDM) system. Two schemes are presented in order to improve the detection performance of SF-BICM perturbated by the error propagation at SIC receiver. First, the bit stream to each transmit antenna is distributed via the rate-compatible puncturing (RCP) such that the SIC process is "code-assisted". Second, a simple heuristic soft bit-metric computation is proposed to compensate for the incorrect soft information caused by the error propagation. Simulation results in the indoor wireless local area network (WLAN) channels show that the proposed RCPC SF-BICM with the new soft bit-metric has significant power gains up to 6dB over the conventional SIC receiver.
Kee-Bong Song, Chan-Soo Hwang, John M. Cioffi
ICC1
2003 Outage capacity and cutoff rate of bit-interleaved coded OFDM under quasistatic frequency selective fading
abstract
IEEE 802.11a wireless local area network (WLAN) system which employs bit-interleaved coded orthogonal frequency division multiplexing (OFDM) is one of the most important applications of bit-interleaved coded modulation (BICM). The capacity and cutoff rate of BICM have been well studied in flat-fading channels assuming infinite-length ideal interleaving. This assumption is practically justifiable only in fast-fading channels. WLAN channel environment is, however, highly frequency selective while slowly varying in time. In this work, we obtain outage capacity and outage cutoff rate expressions for BICM in such quasistatic frequency selective fades, under a uniform ideal interleaving assumption. These results are also extended to the general case of space-frequency BICM (SF-BICM) which is the multiinput multioutput (MIMO) extension of BICM.
Amal Ekbal, Kee-Bong Song, John M. Cioffi
GLOBECOM2
2003 A low complexity space-frequency BICM MIMO-OFDM system for next-generation WLAN
abstract
In this paper, we propose a low complexity space-frequency bit interleaved coded (SF-BICM) multiple-input multiple-output (MIMO) OFDM system as a promising candidate for next generation wireless LANs, expected to operate in excess of 100 Mb/s. A MIMO overlay on a 64-QAM OFDM system can increase computational complexity exponentially if maximum likelihood performance is desired. In this paper, we present two schemes for reducing implementation complexity without compromising PER performance. First, we show that multiplexing OFDM symbols across multiple antennas achieves the best performance, and also leads to the most cost effective implementation. Second, we propose a novel ad-hoc scheme for MIMO soft-symbol demapping that uses hard-decision feedback to perform interference cancellation and maximal ratio combining. Our proposed low-complexity architecture offers a throughput of 108 Mb/s at the comparable SNR of an IEEE 802.11a/g-compliant 54 Mb/s transmission with 1% PER.
Kee-Bong Song, Syed Aon Mujtaba
GLOBECOM1