Youngchul Sung

dblp:17/6798 · DBLP profile ↗
← Back
68ranked-venue papers
12as first author
25since 2021 · last 2025
0000-0003-4536-6690ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 23 since 2021Computer networks · 19 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorTheory of computation · 2 · 2 first-author
YearPublicationVenuePosition
2025 ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
abstract
Recent advances in LLM agents have largely built on reasoning backbones like ReAct (Yao et al., 2023), which interleaves thought and action in complex environments.However, Re-Act often produces ungrounded or incoherent reasoning steps, leading to misalignment between the agent's actual state and goal.Our analysis finds that this stems from ReAct's inability to maintain consistent internal beliefs and goal alignment, causing compounding errors and hallucinations.To address this, we introduce ReflAct, a novel backbone that shifts reasoning from merely planning next actions to continuously reflecting on the agent's state relative to its goal.By explicitly grounding decisions in states and enforcing ongoing goal alignment, ReflAct dramatically improves strategic reliability.This design delivers substantial empirical gains: ReflAct surpasses Re-Act by 27.7% on average, achieving a 93.3% success rate in ALFWorld.Notably, ReflAct even outperforms ReAct with added enhancement modules (e.g., Reflexion, WKM), showing that strengthening the core reasoning backbone is key to reliable agent performance.Observation: You pick up the apple 1 from the countertop 1. Thought: Now I take an apple 1.Next, I need to go to a microwave 1 and heat it.Action: go to microwave 1 Observation: The microwave 1 is closed.Thought: The microwave 1 is closed, so I need to open it first.Action: open microwave 1 Observation: You open the microwave 1.The microwave 1 is open.In it, you see nothing.
Jeonghye Kim, Sojeong Rhee, Minbeom Kim, Sangmook Lee, Youngchul Sung, Kyomin Jung
EMNLP6
2025 Reward Dimension Reduction for Scalable Multi-Objective Reinforcement Learning
abstract
In this paper, we introduce a simple yet effective reward dimension reduction method to tackle the scalability challenges of multi-objective reinforcement learning algorithms. While most existing approaches focus on optimizing two to four objectives, their abilities to scale to environments with more objectives remain uncertain. Our method uses a dimension reduction approach to enhance learning efficiency and policy performance in multi-objective settings. While most traditional dimension reduction methods are designed for static datasets, our approach is tailored for online learning and preserves Pareto-optimality after transformation. We propose a new training and evaluation framework for reward dimension reduction in multi-objective reinforcement learning and demonstrate the superiority of our method in environments including one with sixteen objectives, significantly outperforming existing online dimension reduction methods.
Giseung Park, Youngchul Sung
ICLR2
2025 ARS: Adaptive Reward Scaling for Multi-Task Reinforcement Learning
abstract
Multi-task reinforcement learning (RL) encounters significant challenges due to varying task complexities and their reward distributions from the environment. To address these issues, in this paper, we propose Adaptive Reward Scaling (ARS), a novel framework that dynamically adjusts reward magnitudes and leverages a periodic network reset mechanism. ARS introduces a history-based reward scaling strategy that ensures balanced reward distributions across tasks, enabling stable and efficient training. The reset mechanism complements this approach by mitigating overfitting and ensuring robust convergence. Empirical evaluations on the Meta-World benchmark demonstrate that ARS significantly outperforms baseline methods, achieving superior performance on challenging tasks while maintaining overall learning efficiency. These results validate ARS’s effectiveness in tackling diverse multi-task RL problems, paving the way for scalable solutions in complex real-world applications.
Myungsik Cho, Jongeui Park, Jeonghye Kim, Youngchul Sung
ICML4
2025 Penalizing Infeasible Actions and Reward Scaling in Reinforcement Learning with Offline Data
abstract
Reinforcement learning with offline data suffers from Q-value extrapolation errors. To address this issue, we first demonstrate that linear extrapolation of the Q-function beyond the data range is particularly problematic. To mitigate this, we propose guiding the gradual decrease of Q-values outside the data range, which is achieved through reward scaling with layer normalization (RS-LN) and a penalization mechanism for infeasible actions (PA). By combining RS-LN and PA, we develop a new algorithm called PARS. We evaluate PARS across a range of tasks, demonstrating superior performance compared to state-of-the-art algorithms in both offline training and online fine-tuning on the D4RL benchmark, with notable success in the challenging AntMaze Ultra task.
Jeonghye Kim, Yongjae Shin, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngchul Sung, Kanghoon Lee, Woohyung Lim
ICML6
2025 Online Pre-Training for Offline-to-Online Reinforcement Learning
abstract
Offline-to-online reinforcement learning (RL) aims to integrate the complementary strengths of offline and online RL by pre-training an agent offline and subsequently fine-tuning it through online interactions. However, recent studies reveal that offline pre-trained agents often underperform during online fine-tuning due to inaccurate value estimation caused by distribution shift, with random initialization proving more effective in certain cases. In this work, we propose a novel method, Online Pre-Training for Offline-to-Online RL (OPT), explicitly designed to address the issue of inaccurate value estimation in offline pre-trained agents. OPT introduces a new learning phase, Online Pre-Training, which allows the training of a new value function tailored specifically for effective online fine-tuning. Implementation of OPT on TD3 and SPOT demonstrates an average 30% improvement in performance across a wide range of D4RL environments, including MuJoCo, Antmaze, and Adroit.
Yongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong, Deunsol Yoon, Youngsoo Jang, Geon-Hyeong Kim, Jongseong Chae, Youngchul Sung, Kanghoon Lee, Woohyung Lim
ICML9
2025 Multi-Objective Reinforcement Learning with Max-Min Criterion: A Game-Theoretic Approach
abstract
In this paper, we propose a provably convergent and practical framework for multi-objective reinforcement learning with max-min criterion. From a game-theoretic perspective, we reformulate max-min multi-objective reinforcement learning as a two-player zero-sum regularized continuous game and introduce an efficient algorithm based on mirror descent. Our approach simplifies the policy update while ensuring global last-iterate convergence. We provide a comprehensive theoretical analysis on our algorithm, including iteration complexity under both exact and approximate policy evaluations, as well as sample complexity bounds. To further enhance performance, we modify the proposed algorithm with adaptive regularization. Our experiments demonstrate the convergence behavior of the proposed algorithm in tabular settings, and our implementation for deep reinforcement learning significantly outperforms previous baselines in many MORL environments.
Woohyeon Byeon, Giseung Park, Jongseong Chae, Amir Leshem, Youngchul Sung
NeurIPS5
2025 Adaptive multi-model fusion learning for sparse-reward reinforcement learning
Giseung Park, Whiyoung Jung, Seungyul Han, Sungho Choi, Youngchul Sung
Neurocomputing5
2024 Decision ConvFormer: Local Filtering in MetaFormer is Sufficient for Decision Making
abstract
The recent success of Transformer in natural language processing has sparked its use in various domains. In offline reinforcement learning (RL), Decision Transformer (DT) is emerging as a promising model based on Transformer. However, we discovered that the attention module of DT is not appropriate to capture the inherent local dependence pattern in trajectories of RL modeled as a Markov decision process. To overcome the limitations of DT, we propose a novel action sequence predictor, named Decision ConvFormer (DC), based on the architecture of MetaFormer, which is a general structure to process multiple entities in parallel and understand the interrelationship among the multiple entities. DC employs local convolution filtering as the token mixer and can effectively capture the inherent local associations of the RL dataset. In extensive experiments, DC achieved state-of-the-art performance across various standard RL benchmarks while requiring fewer resources. Furthermore, we show that DC better understands the underlying meaning in data and exhibits enhanced generalization capability.
Jeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul Sung
ICLR4
2024 Hard Tasks First: Multi-Task Reinforcement Learning Through Task Scheduling
abstract
Multi-task reinforcement learning (RL) faces the significant challenge of varying task difficulties, often leading to negative transfer when simpler tasks overshadow the learning of more complex ones. To overcome this challenge, we propose a novel algorithm, Scheduled Multi-Task Training (SMT), that strategically prioritizes more challenging tasks, thereby enhancing overall learning efficiency. SMT introduces a dynamic task prioritization strategy, underpinned by an effective metric for assessing task difficulty. This metric ensures an efficient and targeted allocation of training resources, significantly improving learning outcomes. Additionally, SMT incorporates a reset mechanism that periodically reinitializes key network parameters to mitigate the simplicity bias, further enhancing the adaptability and robustness of the learning process across diverse tasks. The efficacy of SMT's scheduling method is validated by significantly improving performance on challenging Meta-World benchmarks.
Myungsik Cho, Jongeui Park, Suyoung Lee, Youngchul Sung
ICML4
2024 The Max-Min Formulation of Multi-Objective Reinforcement Learning: From Theory to a Model-Free Algorithm
abstract
In this paper, we consider multi-objective reinforcement learning, which arises in many real-world problems with multiple optimization goals. We approach the problem with a max-min framework focusing on fairness among the multiple goals and develop a relevant theory and a practical model-free algorithm under the max-min framework. The developed theory provides a theoretical advance in multi-objective reinforcement learning, and the proposed algorithm demonstrates a notable performance improvement over existing baseline methods.
Giseung Park, Woohyeon Byeon, Elad Havakuk, Amir Leshem, Youngchul Sung
ICML6
2024 Adaptive Q-Aid for Conditional Supervised Learning in Offline Reinforcement Learning
abstract
Offline reinforcement learning (RL) has progressed with return-conditioned supervised learning (RCSL), but its lack of stitching ability remains a limitation. We introduce $Q$-Aided Conditional Supervised Learning (QCS), which effectively combines the stability of RCSL with the stitching capability of $Q$-functions. By analyzing $Q$-function over-generalization, which impairs stable stitching, QCS adaptively integrates $Q$-aid into RCSL's loss function based on trajectory return. Empirical results show that QCS significantly outperforms RCSL and value-based methods, consistently achieving or exceeding the highest trajectory returns across diverse offline RL benchmarks. QCS represents a breakthrough in offline RL, pushing the limits of what can be achieved and fostering further innovations.
Jeonghye Kim, Suyoung Lee, Woojun Kim, Youngchul Sung
NeurIPS4
2023 LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option Framework
abstract
In this paper, a unified framework for exploration in reinforcement learning (RL) is proposed based on an option-critic architecture. The proposed framework learns to integrate a set of diverse exploration strategies so that the agent can adaptively select the most effective exploration strategy to realize an effective exploration-exploitation trade-off for each given task. The effectiveness of the proposed exploration framework is demonstrated by various experiments in the MiniGrid and Atari environments.
Woojun Kim, Jeonghye Kim, Youngchul Sung
ICML3
2023 An Adaptive Entropy-Regularization Framework for Multi-Agent Reinforcement Learning
abstract
In this paper, we propose an adaptive entropy-regularization framework (ADER) for multi-agent reinforcement learning (RL) to learn the adequate amount of exploration of each agent for entropy-based exploration. In order to derive a metric for the proper level of exploration entropy for each agent, we disentangle the soft value function into two types: one for pure return and the other for entropy. By applying multi-agent value factorization to the disentangled value function of pure return, we obtain a metric to determine the relevant level of exploration entropy for each agent, given by the partial derivative of the pure-return value function with respect to (w.r.t.) the policy entropy of each agent. Based on this metric, we propose the ADER algorithm based on maximum entropy RL, which controls the necessary level of exploration across agents over time by learning the proper target entropy for each agent. Experimental results show that the proposed scheme significantly outperforms current state-of-the-art multi-agent RL algorithms.
Woojun Kim, Youngchul Sung
ICML2
2023 Domain Adaptive Imitation Learning with Visual Observation
abstract
In this paper, we consider domain-adaptive imitation learning with visual observation, where an agent in a target domain learns to perform a task by observing expert demonstrations in a source domain. Domain adaptive imitation learning arises in practical scenarios where a robot, receiving visual sensory data, needs to mimic movements by visually observing other robots from different angles or observing robots of different shapes. To overcome the domain shift in cross-domain imitation learning with visual observation, we propose a novel framework for extracting domain-independent behavioral features from input observations that can be used to train the learner, based on dual feature extraction and image reconstruction. Empirical results demonstrate that our approach outperforms previous algorithms for imitation learning from visual observation with domain shift.
Sungho Choi, Seungyul Han, Woojun Kim, Jongseong Chae, Whiyoung Jung, Youngchul Sung
NeurIPS6
2023 Sample-Efficient and Safe Deep Reinforcement Learning via Reset Deep Ensemble Agents
abstract
Deep reinforcement learning (RL) has achieved remarkable success in solving complex tasks through its integration with deep neural networks (DNNs) as function approximators. However, the reliance on DNNs has introduced a new challenge called primacy bias, whereby these function approximators tend to prioritize early experiences, leading to overfitting. To alleviate this bias, a reset method has been proposed, which involves periodic resets of a portion or the entirety of a deep RL agent while preserving the replay buffer. However, the use of this method can result in performance collapses after executing the reset, raising concerns from the perspective of safe RL and regret minimization. In this paper, we propose a novel reset-based method that leverages deep ensemble learning to address the limitations of the vanilla reset method and enhance sample efficiency. The effectiveness of the proposed method is validated through various experiments including those in the domain of safe RL. Numerical results demonstrate its potential for real-world applications requiring high sample efficiency and safety considerations.
Woojun Kim, Yongjae Shin, Jongeui Park, Youngchul Sung
NeurIPS4
2023 Parameterizing Non-Parametric Meta-Reinforcement Learning Tasks via Subtask Decomposition
abstract
Meta-reinforcement learning (meta-RL) techniques have demonstrated remarkable success in generalizing deep reinforcement learning across a range of tasks. Nevertheless, these methods often struggle to generalize beyond tasks with parametric variations. To overcome this challenge, we propose Subtask Decomposition and Virtual Training (SDVT), a novel meta-RL approach that decomposes each non-parametric task into a collection of elementary subtasks and parameterizes the task based on its decomposition. We employ a Gaussian mixture VAE to meta-learn the decomposition process, enabling the agent to reuse policies acquired from common subtasks. Additionally, we propose a virtual training procedure, specifically designed for non-parametric task variability, which generates hypothetical subtask compositions, thereby enhancing generalization to previously unseen subtask compositions. Our method significantly improves performance on the Meta-World ML-10 and ML-45 benchmarks, surpassing current state-of-the-art techniques.
Suyoung Lee, Myungsik Cho, Youngchul Sung
NeurIPS3
2023 Joint Direct and Indirect Channel Estimation for RIS-Assisted Millimeter-Wave Systems Based on Array Signal Processing
abstract
Reconfigurable intelligent surface (RIS)-assisted millimeter wave (mmWave) communication is a promising technology for enlarging the coverage area of millimeter wave systems. Unfortunately, realizing the full potential of these systems requires addressing numerous challenges in channel estimation. In this paper, channel estimation for RIS-assisted mmWave communications is considered. Under the assumption that the array manifolds of the base station antennas and the RIS reflecting elements are given by uniform arrays, an efficient two-stage channel estimation method based on array signal processing techniques is proposed. In the proposed algorithm, the direct and indirect channels are jointly estimated by space-time processing that exploits the sparsity in RIS-assisted mmWave channels and the features associated with uniform arrays. Then, several practical issues, including detection of the number of channel paths, imperfect RIS hardware, and complexity, are addressed. Extensions to the cases of uniform planar array-based RIS, wideband communication, and multiple users are also discussed. Numerical results validate the effectiveness of the proposed method.
Song Noh, Kyungsik Seo, Youngchul Sung, David J. Love, Junse Lee, Heejung Yu
IEEE Trans. Wirel. Commun.3
2022 Blockwise Sequential Model Learning for Partially Observable Reinforcement Learning
abstract
This paper proposes a new sequential model learning architecture to solve partially observable Markov decision problems. Rather than compressing sequential information at every timestep as in conventional recurrent neural network-based methods, the proposed architecture generates a latent variable in each data block with a length of multiple timesteps and passes the most relevant information to the next block for policy optimization. The proposed blockwise sequential model is implemented based on self-attention, making the model capable of detailed sequential learning in partial observable settings. The proposed model builds an additional learning network to efficiently implement gradient estimation by using self-normalized importance sampling, which does not require the complex blockwise input data reconstruction in the model learning. Numerical results show that the proposed method significantly outperforms previous methods in various partially observable environments.
Giseung Park, Sungho Choi, Youngchul Sung
AAAI3
2022 Robust Imitation Learning against Variations in Environment Dynamics
abstract
In this paper, we propose a robust imitation learning (IL) framework that improves the robustness of IL when environment dynamics are perturbed. The existing IL framework trained in a single environment can catastrophically fail with perturbations in environment dynamics because it does not capture the situation that underlying environment dynamics can be changed. Our framework effectively deals with environments with varying dynamics by imitating multiple experts in sampled environment dynamics to enhance the robustness in general variations in environment dynamics. In order to robustly imitate the multiple sample experts, we minimize the risk with respect to the Jensen-Shannon divergence between the agent’s policy and each of the sample experts. Numerical results show that our algorithm significantly improves robustness against dynamics perturbations compared to conventional IL baselines.
Jongseong Chae, Seungyul Han, Whiyoung Jung, Myungsik Cho, Sungho Choi, Youngchul Sung
ICML6
2022 MASER: Multi-Agent Reinforcement Learning with Subgoals Generated from Experience Replay Buffer
abstract
In this paper, we consider cooperative multi-agent reinforcement learning (MARL) with sparse reward. To tackle this problem, we propose a novel method named MASER: MARL with subgoals generated from experience replay buffer. Under the widely-used assumption of centralized training with decentralized execution and consistent Q-value decomposition for MARL, MASER automatically generates proper subgoals for multiple agents from the experience replay buffer by considering both individual Q-value and total Q-value. Then, MASER designs individual intrinsic reward for each agent based on actionable representation relevant to Q-learning so that the agents reach their subgoals while maximizing the joint action value. Numerical results show that MASER significantly outperforms StarCraft II micromanagement benchmark compared to other state-of-the-art MARL algorithms.
Jeewon Jeon, Woojun Kim, Whiyoung Jung, Youngchul Sung
ICML4
2022 Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage Probability
abstract
Constrained reinforcement learning (RL) is an area of RL whose objective is to find an optimal policy that maximizes expected cumulative return while satisfying a given constraint. Most of the previous constrained RL works consider expected cumulative sum cost as the constraint. However, optimization with this constraint cannot guarantee a target probability of outage event that the cumulative sum cost exceeds a given threshold. This paper proposes a framework, named Quantile Constrained RL (QCRL), to constrain the quantile of the distribution of the cumulative sum cost that is a necessary and sufficient condition to satisfy the outage constraint. This is the first work that tackles the issue of applying the policy gradient theorem to the quantile and provides theoretical results for approximating the gradient of the quantile. Based on the derived theoretical results and the technique of the Lagrange multiplier, we construct a constrained RL algorithm named Quantile Constrained Policy Optimization (QCPO). We use distributional RL with the Large Deviation Principle (LDP) to estimate quantiles and tail probability of the cumulative sum cost for the implementation of QCPO. The implemented algorithm satisfies the outage probability constraint after the training period.
Whiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul Sung
NeurIPS4
2022 Training Signal Design for Sparse Channel Estimation in Intelligent Reflecting Surface-Assisted Millimeter-Wave Communication
abstract
In this paper, the problem of training signal design for intelligent reflecting surface (IRS)-assisted millimeter-wave (mmWave) communication under a sparse channel model is considered. The problem is approached based on the Cramér-Rao lower bound (CRB) on the mean-square error (MSE) of channel estimation. By exploiting the sparse structure of mmWave channels, the CRB for the channel parameter composed of path gains and path angles is derived in closed form under Bayesian and hybrid parameter assumptions. Based on the derivation and analysis, an IRS reflection pattern design method is proposed by minimizing the CRB as a function of design variables under constant modulus constraint on reflection coefficients. Extensions of the proposed design to a multi-antenna transceiver, a uniform planar array (UPA)-based IRS, and multi-user case are discussed. Numerical results validate the effectiveness of the proposed design method for sparse mmWave channel estimation.
Song Noh, Heejung Yu, Youngchul Sung
IEEE Trans. Wirel. Commun.3
2021 Communication in Multi-Agent Reinforcement Learning: Intention Sharing
Woojun Kim, Jongeui Park, Youngchul Sung
ICLR3
2021 Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient Exploration
abstract
In this paper, sample-aware policy entropy regularization is proposed to enhance the conventional policy entropy regularization for better exploration. Exploiting the sample distribution obtainable from the replay buffer, the proposed sample-aware entropy regularization maximizes the entropy of the weighted sum of the policy action distribution and the sample action distribution from the replay buffer for sample-efficient exploration. A practical algorithm named diversity actor-critic (DAC) is developed by applying policy iteration to the objective function with the proposed sample-aware entropy regularization. Numerical results show that DAC significantly outperforms existing recent algorithms for reinforcement learning.
Seungyul Han, Youngchul Sung
ICML2
2021 A Max-Min Entropy Framework for Reinforcement Learning
abstract
In this paper, we propose a max-min entropy framework for reinforcement learning (RL) to overcome the limitation of the soft actor-critic (SAC) algorithm implementing the maximum entropy RL in model-free sample-based learning. Whereas the maximum entropy RL guides learning for policies to reach states with high entropy in the future, the proposed max-min entropy framework aims to learn to visit states with low entropy and maximize the entropy of these low-entropy states to promote better exploration. For general Markov decision processes (MDPs), an efficient algorithm is constructed under the proposed max-min entropy framework based on disentanglement of exploration and exploitation. Numerical results show that the proposed algorithm yields drastic performance improvement over the current state-of-the-art RL algorithms.
Seungyul Han, Youngchul Sung
NeurIPS2
2020 Population-Guided Parallel Policy Search for Reinforcement Learning
Whiyoung Jung, Giseung Park, Youngchul Sung
ICLR3
2020 Fast Beam Search and Refinement for Millimeter-Wave Massive MIMO Based on Two-Level Phased Arrays
abstract
In this paper, a new method of fast beam search and refinement is proposed for millimeter-wave hybrid beamforming systems. The proposed method is based on a two-level phased array approach in which the first-level phased array is formed by actual analog-domain subarrays, and the second-level phased array is virtually formed by the outputs of the first-level subarray phased combiners. Exploiting the fact that the overall beam pattern of the two-level array is the product of the beam patterns of the two levels, the proposed method searches angle-of-arrivals by sweeping coarse-resolution analog training beams at the first-level phased array and matching the first-level subarray outputs with fine-resolution digital training beams by parallel fast Fourier transform (FFT) filtering. In this way, the overall beam search resolution of the proposed method is given by the fine resolution of the second-level virtual array only with the overhead of coarse beam sweeping at the first-level subarrays. Hence, the proposed method provides a very efficient way of beam training and refinement. The performance of the proposed method including the directional ambiguity, beam search latency, and computational complexity is analyzed. Numerical results show the effectiveness of the proposed method.
Song Noh, Jiho Song, Youngchul Sung
IEEE Trans. Wirel. Commun.3
2019 Message-Dropout: An Efficient Training Method for Multi-Agent Deep Reinforcement Learning
abstract
In this paper, we propose a new learning technique named message-dropout to improve the performance for multi-agent deep reinforcement learning under two application scenarios: 1) classical multi-agent reinforcement learning with direct message communication among agents and 2) centralized training with decentralized execution. In the first application scenario of multi-agent systems in which direct message communication among agents is allowed, the messagedropout technique drops out the received messages from other agents in a block-wise manner with a certain probability in the training phase and compensates for this effect by multiplying the weights of the dropped-out block units with a correction probability. The applied message-dropout technique effectively handles the increased input dimension in multi-agent reinforcement learning with communication and makes learning robust against communication errors in the execution phase. In the second application scenario of centralized training with decentralized execution, we particularly consider the application of the proposed messagedropout to Multi-Agent Deep Deterministic Policy Gradient (MADDPG), which uses a centralized critic to train a decentralized actor for each agent. We evaluate the proposed message-dropout technique for several games, and numerical results show that the proposed message-dropout technique with proper dropout rate improves the reinforcement learning performance significantly in terms of the training speed and the steady-state performance in the execution phase.
Woojun Kim, Myungsik Cho, Youngchul Sung
AAAI3
2019 Dimension-Wise Importance Sampling Weight Clipping for Sample-Efficient Reinforcement Learning
abstract
In importance sampling (IS)-based reinforcement learning algorithms such as Proximal Policy Optimization (PPO), IS weights are typically clipped to avoid large variance in learning. However, policy update from clipped statistics induces large bias in tasks with high action dimensions, and bias from clipping makes it difficult to reuse old samples with large IS weights. In this paper, we consider PPO, a representative on-policy algorithm, and propose its improvement by dimension-wise IS weight clipping which separately clips the IS weight of each action dimension to avoid large bias and adaptively controls the IS weight to bound policy update from the current policy. This new technique enables efficient learning for high action-dimensional tasks and reusing of old samples like in off-policy learning to increase the sample efficiency. Numerical results show that the proposed new algorithm outperforms PPO and other RL algorithms in various Open AI Gym tasks.
Seungyul Han, Youngchul Sung
ICML2
2019 A High-Diversity Transceiver Design for MISO Broadcast Channels
abstract
In this paper, the outage behavior and diversity order of the mixture transceiver architecture for multiple-input single-output broadcast channels are analyzed. The mixture scheme groups users with closely-aligned channels and applies superposition coding and successive interference cancellation decoding to each group composed of users with closely-aligned channels, while applying zero-forcing beamforming across semi-orthogonal user groups. In order to enable such analysis, closed-form lower bounds on the achievable rates of a general multiple-input single-output broadcast channel with superposition coding and successive interference cancellation are newly derived. By employing channel-adaptive user grouping and proper power allocation, which ensures that the channel subspaces of user groups have an angle larger than a certain threshold, it is shown that the mixture transceiver architecture achieves full diversity order in multiple-input single-output broadcast channels and opportunistically increases the multiplexing gain while achieving full diversity order. Furthermore, the achieved full diversity order is the same as that of the single-user maximal ratio transmit beamforming. Hence, the mixture scheme can provide reliable communication under channel fading for ultra-reliable low latency communication. The numerical results validate our analysis and show the outage superiority of the mixture scheme over conventional transceiver designs for multiple-input single-output broadcast channels.
Junyeong Seo, Youngchul Sung, Hamid Jafarkhani
IEEE Trans. Wirel. Commun.2
2018 Common Pilot Signal Design for Non-Orthogonal Multiple Access in 5G
abstract
In this paper, common pilot signal design for channel estimation of users with different channel strengths simultaneously served by non-orthogonal multiple access (NOMA) is considered. The problem is approached based on two design criteria: minimization of weighted sum mean-square error (MSE) and maximization of weighted sum conditional mutual information (CMI) between the channel and the received signal given the pilot signal. It is shown that the two design problems reduce to semi-definite programming (SDP). Furthermore, in the practical case of two-user clustering, closed-form solutions to the two problems are obtained under some additional assumption on the channel's statistical information. Numerical result shows that the pilot signal design exploiting the multi-group structure of NOMA yields better channel estimation performance than conventional design.
Jungho So, Youngchul Sung
TENCON2
2018 A New Approach to User Scheduling in Massive Multi-User MIMO Broadcast Channels
abstract
In this paper, a new two-phase-feedback-based user-scheduling-and-beamforming method is proposed for multi-user multiple-input multiple-output downlink in the context of two-stage beamforming. The key ideas of the proposed method are: 1) to use a set of orthogonal reference beams and construct a cone around each reference beam to select “nearly-optimal” semi-orthogonal users based only on channel quality indicator feedback; and 2) to apply post-user-selection beam design with zero-forcing beamforming (ZFBF) based on channel state information (CSI) feedback only from the selected users. It is proven that the proposed scheduling-and-beamforming method is asymptotically optimal as the number of users increase. Furthermore, the proposed scheduling-and-beamforming method almost achieves the performance of the existing semi-orthogonal user selection with ZFBF that requires full CSI for all users, with a significantly reduced amount of required channel information which is even less than that required by random beamforming.
Gilwon Lee, Youngchul Sung
IEEE Trans. Commun.2
2016 Randomly-Directional Beamforming in Millimeter-Wave Multiuser MISO Downlink
abstract
In this paper, the performance of opportunistic random beamforming (RBF) and the multiuser (MU) gain in millimeter-wave (mm-wave) MU multiple-input single-output (MISO) downlink systems are analyzed based on the uniform random single-path (UR-SP) channel model suitable for highly directional mm-wave radio propagation channels. It is shown that under the UR-SP channel model, RBF achieves linear sum rate scaling with respect to (w.r.t.) the number of transmit antennas and, furthermore, yields optimal sum rate performance when the number of transmit antennas is large, if the number of users increases linearly w.r.t. the number of transmit antennas. Several beam training and user selection methods are investigated to yield insights into the most effective beamforming and scheduling choice for mm-wave MU-MISO in various operating conditions. Simulation results validate our analysis based on asymptotic techniques for finite cases.
Gilwon Lee, Youngchul Sung, Junyeong Seo
IEEE Trans. Wirel. Commun.2
2015 Pilot Signal Design for Massive MIMO Systems: A Received Signal-To-Noise-Ratio-Based Approach
abstract
In this letter, the pilot signal design for massive MIMO systems to maximize the training-based received signal-to-noise ratio (SNR) is considered under two channel models: block Gauss-Markov and block independent and identically distributed (i.i.d.) channel models. First, it is shown that under the block Gauss-Markov channel model, the optimal pilot design problem reduces to a semi-definite programming (SDP) problem, which can be solved numerically by a standard convex optimization tool. Second, under the block i.i.d. channel model, an optimal solution is obtained in closed form. Numerical results show that the proposed method yields noticeably better performance than other existing pilot design methods in terms of received SNR.
Jungho So, Donggun Kim 0001, Yuni Lee, Youngchul Sung
IEEE Signal Process. Lett.4
2015 Two-Stage Beamformer Design for Massive MIMO Downlink By Trace Quotient Formulation
abstract
In this paper, the problem of outer beamformer design based only on channel statistic information is considered for two-stage beamforming for multi-user massive MIMO downlink, and the problem is approached based on signal-to-leakage-plus-noise ratio (SLNR). To eliminate the dependence on the instantaneous channel state information, a lower bound on the average SLNR is derived by assuming zero-forcing (ZF) inner beamforming, and an outer beamformer design method that maximizes the lower bound on the average SLNR is proposed. It is shown that the proposed SLNR-based outer beamformer design problem reduces to a trace quotient problem (TQP), which is often encountered in the field of machine learning. An iterative algorithm is presented to obtain an optimal solution to the proposed TQP. The proposed method has the capability of optimally controlling the weighting factor between the signal power to the desired user and the interference leakage power to undesired users according to different channel statistics. Numerical results show that the proposed outer beamformer design method yields significant performance gain over existing methods.
Donggun Kim 0001, Gilwon Lee, Youngchul Sung
IEEE Trans. Commun.3
2014 Training signal design for channel estimation in massive MIMO systems
abstract
In this paper, the design of training signals for channel estimation in massive multiple-input multiple-output (MIMO) systems is considered. Under a stationary, block Gauss-Markov channel model, a method for optimal pilot beam pattern design for enhanced channel estimation is proposed, exploiting both the properties of Kalman filtering and the spatio-temporal channel correlation. First, pilot beam pattern design is considered under the assumption of orthogonal beam patterns within a block. The orthogonality assumption is subsequently relaxed and the design problem is solved via a greedy approach. Numerical results show the efficacy of the proposed algorithm.
Song Noh, Michael D. Zoltowski, Youngchul Sung, David J. Love
ICASSP3
2014 An extended least difference greedy clique-cover algorithm for index coding
abstract
In this paper, linear binary index coding is considered. It is shown that the minimum clique-cover heuristic algorithm can provide an efficient way to solving linear binary index coding problems. Based on the least difference greedy (LDG) clique-cover algorithm, an existing minimum clique-cover algorithm for index coding, proposed by Birk and Kol [1], [2], we develop an extended LDG algorithm by considering a transpose index coding model and cycle detection in the side information graph. Numerical results show that the proposed algorithm considerably outperforms the conventional LDG algorithm in terms of the number of transmissions.
Sangwoon Kwak, Jungho So, Youngchul Sung
ISIT3
2014 Some new results on index coding when the number of data is less than the number of receivers
abstract
In this paper, index coding problems in which the number (m) of receivers is larger than that (n) of data are considered. Unlike the case that the two numbers are same (n = m), index coding problems with n ≤ m are more general and hard to handle. To circumvent this difficulty, problems with n < m are approached via corresponding problems with n = m. It is shown that in certain cases, the symmetric capacity and code construction for index coding problems with n < m can be obtained from the existing symmetric capacity result and codes for index coding problems with n = m. Such cases include cases with n < m ≤ 5.
Jungho So, Sangwoon Kwak, Youngchul Sung
ISIT3
2014 Filter-And-Forward Relay Design for MIMO-OFDM Systems
abstract
In this paper, the filter-and-forward (FF) relay design for multiple-input-multiple-output (MIMO) orthogonal frequency-division multiplexing (OFDM) systems is investigated. Due to the considered MIMO structure, the FF relay design is investigated in the framework of joint design together with the linear MIMO transceiver. As the design criterion, first, the minimization of weighted sum mean square error (MSE) is considered. The joint design in this case is approached based on alternating optimization that iterates between the optimal design of the FF relay for given MIMO precoding and decoding matrices and the optimal design of MIMO precoding and decoding matrices for a given FF relay filter. Second, a more advanced problem of joint design for rate maximization is considered. The second problem is approached based on the obtained result regarding the first problem of weighted sum MSE minimization and the existing result regarding the relationship between weighted MSE minimization and rate maximization. Numerical results show the effectiveness of the proposed FF relay design method and significant performance improvement by the proposed FF relay over widely considered simple AF relays for MIMO-OFDM systems.
Donggun Kim 0001, Youngchul Sung, Jihoon Chung
IEEE Trans. Commun.2
2014 A New Precoder Design for Blind Channel Estimation in MIMO-OFDM Systems
abstract
A new precoder for precoding-based blind channel estimation for MIMO-OFDM systems is proposed. In the proposed scheme, only a small number of data symbols, commensurate with the channel length, are linearly precoded prior to transmission to induce the signal correlation needed for the scheme to blindly estimate the channels. Similar to how pilot symbols are transmitted, the subcarriers carrying these linearly precoded data symbols are equi-spaced across the frequency band. The other subcarriers carry data symbols in the standard way, enabling MLD per subcarrier and also allowing for MIMO linear precoding across antennas at each subcarrier. This is in contrast to previous precoding-based blind channel estimation schemes which precode all of the data symbols so that every subcarrier carries a linear combination of symbols, aking the resulting joint MLD problem infeasible. In addition, this also makes it infeasible to employ MIMO precoding per subcarrier across the transmit antennas. The proposed precoder is designed via a multi-stage optimization process that seeks to minimize both channel estimation error and symbol estimation error. For channel estimation purposes, the resulting optimal design offers low-cost features such as sign change and FFT while providing reasonable channel estimation performance for low mobility applications.
Song Noh, Youngchul Sung, Michael D. Zoltowski
IEEE Trans. Wirel. Commun.2
2013 An efficient parameterization for Pareto-optimal beamformers for k-user MIMO interference channels
abstract
In this paper, Pareto-optimal beamforming in the K-pair Gaussian multiple-input multiple-output (MIMO) interference channel is considered. Under the assumption of Gaussian signaling at transmitters and single-user decoding at receivers, a necessary condition for any transmit signal covariance matrix to achieve a Pareto boundary point of the achievable rate region is derived. Based on the necessary condition for Pareto-optimality, an efficient parameterization for Pareto-optimal transmit signal covariance matrices is obtained. The obtained parameter space is given by the product manifold of a Stiefel manifold and a subset of a hyperplane, which is a low dimensional embedded submanifold of the original high dimensional beam search space. The new parameterization enables us to devise very efficient beam design algorithms for the K-pair MIMO interference channel.
Juho Park, Youngchul Sung
ICASSP2
2013 Guest Editorial: Theories and Methods for Advanced Wireless Relays - Issue II
abstract
The demand for wireless access continues to increase rapidly in both military and civilian communities. The modern internet and modern personal-area devices have made billions of users around the world accustomed to data-hungry applications such as videos. This has an inevitable effect on the users' desire for the same through wireless media. The articles in this special issue focus on new theories and methods for advanced wireless relay technologies.
Yingbo Hua, Daniel W. Bliss, Saeed Gazor, Yue Rong, Youngchul Sung
IEEE J. Sel. Areas Commun.5
2012 The Bahadur efficiency for energy detection of stationary Gaussian processes
abstract
In this paper, the performance and optimization of energy detection of stationary Gaussian signals are considered. Based on the Bahadur asymptotic relative efficiency, the performance of energy detection relative to optimal detection is compared, and the optimal threshold for energy detection is derived. It is shown that the optimal threshold for optimal detection is not optimal for energy detection, and an integral equation for determining the optimal threshold for energy detection is provided. A numerical example of the detection of equi-correlated signals is provided, and the numerical result validates our asymptotic analysis in the finite sample regime.
Yuni Lee, Youngchul Sung
ICASSP2
2012 Guest Editorial Theories and Methods for Advanced Wireless Relays - Issue I
abstract
The 46 papers focusing on the theme of "Theories and Methods for Advanced Wireless Relays" have been divided into two groups to be published in two separate issues. This first issue includes 23 papers on relay performance bound, MIMO relay beamforming, relay channel estimation, two-way and shared relays, full-duplex relays, and security for relay networks. The second issue includes papers on coding for relay networks, medium access control for relays, implementation ans system performance studies.
Yingbo Hua, Daniel W. Bliss, Saeed Gazor, Yue Rong, Youngchul Sung
IEEE J. Sel. Areas Commun.5
2012 Generalized Chernoff Information for Mismatched Bayesian Detection and Its Application to Energy Detection
abstract
In this letter, the performance of mismatched likelihood ratio detectors for binary Bayesian hypothesis testing problems is considered. Based on large deviation theory, a method for achieving the maximum Bayesian error exponent for a mismatched likelihood ratio detector is presented. It is shown that the maximum Bayesian error exponent is given by generalized Chernoff information, which is an extension of the Chernoff information to the case of two mismatched distributions and has similar properties to those of the original Chernoff information. As an application example, energy detection under the Gauss-Markov signal model, is considered. It is shown that the generalized Chernoff information of energy detection, which is achieved by optimally choosing the detection threshold, is close to the original Chernoff information for the considered signal model, and thus, the performance of suboptimal energy detection can be improved significantly simply by choosing the detection threshold judiciously.
Yuni Lee, Youngchul Sung
IEEE Signal Process. Lett.2
2012 A Nonlinear Transceiver Architecture for Overloaded Multiuser MIMO Interference Channels
abstract
A new transceiver architecture for overloaded MIMO interference channels is proposed to fix the rate saturation problem of purely linear beamforming in this case. The proposed scheme is based on a mixture of linear beamforming and multi-user detection. It is shown that non-trivial degrees of freedom can be achieved by the proposed mixed scheme properly dividing the interference signals for linear processing and multiuser detection, and the achievable degrees of freedom of the proposed scheme are obtained. Numerical results show that the proposed scheme outperforms linear beamforming in overloaded MIMO interference channels.
Youngseok Oh, Heejung Yu, Yong Hoon Lee, Youngchul Sung
IEEE Trans. Commun.4
2012 Outage Probability and Outage-Based Robust Beamforming for MIMO Interference Channels with Imperfect Channel State Information
abstract
In this paper, the outage probability and outage-based beam design for multiple-input multiple-output (MIMO) interference channels are considered. First, closed-form expressions for the outage probability in MIMO interference channels are derived under the assumption of Gaussian-distributed channel state information (CSI) error, and the asymptotic behavior of the outage probability as a function of several system parameters is examined by using the Chernoff bound. It is shown that the outage probability decreases exponentially with respect to the quality of CSI measured by the inverse of the mean square error of CSI. Second, based on the derived outage probability expressions, an iterative beam design algorithm for maximizing the sum outage rate is proposed. Numerical results show that the proposed beam design algorithm yields significantly better sum outage rate performance than conventional algorithms such as interference alignment developed under the assumption of perfect CSI.
Juho Park, Youngchul Sung, Donggun Kim 0001, H. Vincent Poor
IEEE Trans. Wirel. Commun.2
2011 Sum Outage-Rate Maximization for MIMO Interference Channels
abstract
In this paper, the weighted sum outage-rate maximization for time-invariant multiple-input multiple-output (MIMO) interference channels is considered. The cumulative distribution function (CDF) of outage events in MIMO interference channels is characterized under the assumption of Gaussian distribution for channel uncertainty. Based on the derived expression for the outage probability an iterative beam design algorithm for maximizing the weighted sum outage-rate under outage constraints is newly proposed. Numerical results show that the proposed algorithm shows good sum outage-rate performance.
Juho Park, Donggun Kim 0001, Youngchul Sung
GLOBECOM3
2011 Adaptive beam tracking for interference alignment in time-varying MIMO interference channels: Conjugate gradient approach
abstract
Based on a linear formulation to interference alignment, an adaptive algorithm for interference-aligning beam tracking in time-varying MIMO interference channels is proposed. It is shown that obtaining the interference-aligning beam vector is equivalent to minimizing a certain Rayleigh quotient, and the conjugate gradient approach is adopted to construct an adaptive algorithm. The convergence and stability of the proposed algorithm are established in static channel case, and numerical results show that the proposed algorithm performs well compared with other existing methods with much less complexity.
Junse Lee, Heejung Yu, Youngchul Sung, Yong Hoon Lee
ICASSP3
2010 Adaptive beam tracking for interference alignment for multiuser time-varying MIMO interference channels
abstract
The problem of interference alignment in time-varying MIMO interference channels is considered. To reduce complexity, an adaptive algorithm for beam vector design is proposed based on our previous work of least squares approach to beam design for interference alignment and matrix perturbation theory. The proposed algorithm calculates interference-aligning beam vectors by additive update of the previous value and reduces complexity significantly. Numerical results are provided to validate the proposed algorithm. It is shown that the proposed adaptive algorithm yields almost the same performance as a non-adaptive method that calculates interference-aligning beam vectors at every time step.
Heejung Yu, Youngchul Sung, Haksoo Kim, Yong Hoon Lee
ICASSP2
2009 Upper bound for the loss factor of energy detection of random signals in multipath fading cognitive radios
abstract
In this paper, the loss of energy detection compared with optimal sensing caused by neglecting signal correlation due to multipath fading is considered. The loss factor or relative performance of energy detection compared with optimal sensing is analyzed using Pitman's asymptotic relative efficiency (ARE) which is defined as the ratio of the required number of samples of one detector to that of the other to yield the same detection performance in large sample scheme. Under the assumption of L-tap finite impulse response (FIR) channel with zero-mean independent and identically distributed (i.i.d.) tap coefficients, it is shown that the loss factor of the energy detection relative to optimal sensing is no larger than 1/2 in large delay spread case (i.e., strong correlation); under the same signal power condition the required number of samples for energy detection neglecting the signal correlation is no more than twice of that required for optimal sensing exploiting the signal correlation fully.
Yirang Lim, Youngchul Sung
ICASSP2
2009 Multi-band CSMA/CA-based cognitive radio networks
abstract
A new flexible multiple access control (MAC) scheme embedding channelization in multi-band carrier sense multiple access / collision avoidance (CSMA/CA) systems is proposed to provide different priority classes among users, and its performance is investigated. In the proposed scheme, two priority classes of users, primary and secondary users, are considered; the classical CSMA/CA protocol is modified to multiband operation for secondary users, whereas channelization is provided for primary users to ensure the quality of service (QoS). The performance of the proposed MAC scheme is analyzed using a new multi-band CSMA/CA model based on a Markov chain capturing the primary user channel activity and the number of bands. The throughput of secondary users is obtained as a function of the primary user activity as well as other CSMA/CA parameters. It is shown that this mixed MAC scheme can guarantee the throughput of primary users certainly with minimal performance degradation to secondary users, especially when the number of available bands is large or the maximum contention window (CW) size for the CSMA/CA secondary users is small compared with the number of secondary users.
Jo Woon Chong, Youngchul Sung, Dan Keun Sung
IWCMC2
2009 Amplify-Forward Relays with Superimposed Pilot Signals for Frequency-Selective Fading Channels
abstract
In this paper, a new amplify-and-forward (AF) relay systems using superimposed pilot signals is proposed for frequency selective fading channel environments. In the proposed scheme, the relay superimposes an additional pilot sequence on the received signal from the source, and forwards the combined signal to the destination. The superimposed pilot signal in this way enables the destination node to estimate the channel between the relay and destination (R-D), and this channel information is used to process the received signal at the destination optimally; the noise correlation due to the frequency selective fading in the R-D channel is whitened. The optimal power ratio of the superimposed pilot to the incoming signal at the relay is obtained via simulations, and it is shown under frequency selective channel environments that the proposed scheme yields better bit-error- rate (BER) performance than the conventional method where all relay power is allocated only to the incoming signal.
Haksoo Kim, Sungho Choi, Heejung Yu, Youngchul Sung, Yong Hoon Lee
VTC Spring4
2009 Upper Bound for the Loss of Energy Detection of Signals in Multipath Fading Channels
abstract
The performance of energy detection under multipath fading is analyzed and compared with locally optimal detection using Pitman's asymptotic relative efficiency. Under the L-tap finite impulse response channel model with zero-mean independent and identically distributed tap coefficients, it is shown that the average performance loss of energy detection is no greater than 50% in sample size for the same performance compared with locally optimal detection exploiting signal correlation. Also, an algorithm exploiting signal correlation and improving the detection performance is proposed based on the estimation of signal correlation. Numerical results show that the proposed algorithm almost achieves the performance of locally optimal detection.
Yirang Lim, Juho Park, Youngchul Sung
IEEE Signal Process. Lett.3
2009 How much information can one get from a wireless ad hoc sensor network over a correlated random field?
abstract
New large-deviations results that characterize the asymptotic information rates for general d-dimensional (d -D) stationary Gaussian fields are obtained. By applying the general results to sensor nodes on a two-dimensional (2-D) lattice, the asymptotic behavior of ad hoc sensor networks deployed over correlated random fields for statistical inference is investigated. Under a 2-D hidden Gauss-Markov random field model with symmetric first-order conditional autoregression and the assumption of no in-network data fusion, the behavior of the total obtainable information [nats] and energy efficiency [nats/J] defined as the ratio of total gathered information to the required energy is obtained as the coverage area, node density, and energy vary. When the sensor node density is fixed, the energy efficiency decreases to zero with rate Theta(area-1/2) and the per-node information under fixed per-node energy also diminishes to zero with rate O(Nt-1/3) as the number Ntof network nodes increases by increasing the coverage area. As the sensor spacing dnincreases, the per-node information converges to its limit D with rate D-radic(dn)e-alphadnfor a given diffusion rate alpha. When the coverage area is fixed and the node density increases, the per-node information is inversely proportional to the node density. As the total energy Et consumed in the network increases, the total information obtainable from the network is given by O(logEt) for the fixed node density and fixed coverage case and by Theta(Et2/3) for the fixed per-node sensing energy and fixed density and increasing coverage case.
Youngchul Sung, H. Vincent Poor, Heejung Yu
IEEE Trans. Inf. Theory1
2008 Analysis of CSMA/CA Systems under Carrier Sensing Error: Throughput, Delay and Sensitivity
abstract
In this paper, the performance of the carrier sense multiple access/ collision avoidance (CSMA/CA) protocol under the presence of carrier sensing error is analyzed. Based on our previous work [1], we extend the results to n-user case (n>2), and analyze the sensitivity of the throughput with respect to the key physical-layer parameter, the sensing threshold, via a cross-layer approach. The result provide guidelines about how to operate the CSMA/CA considering imperfect sensing at physical layer. It is shown that the throughput sensitivity highly depends on the ratio of the contention window size W to the frame length L, and the throughput is sensitive to the design of the sensing threshold when W/L is either small or large.
Jo Woon Chong, Youngchul Sung, Dan Keun Sung
GLOBECOM2
2008 Large deviations analysis for the detection of 2D hidden Gauss-Markov random fields using sensor networks
abstract
The detection of hidden two-dimensional Gauss-Markov random fields using sensor networks is considered. Under a conditional autoregressive model, the error exponent for the Neyman-Pearson detector satisfying a fixed level constraint is obtained using the large deviations principle. For a symmetric first order autoregressive model, the error exponent is given explicitly in terms of the SNR and an edge dependence factor (field correlation). The behavior of the error exponent as a function of correlation strength is seen to divide into two regions depending on the value of the SNR. At high SNR, uncorrected observations maximize the error exponent for a given SNR, whereas there is non-zero optimal correlation at low SNR. Based on the error exponent, the energy efficiency (defined as the ratio of the total information gathered to the total energy required) of ad hoc sensor network for detection is examined for two sensor deployment models: an infinite area model and and infinite density model. For a fixed sensor density, the energy efficiency diminishes to zero at rate 0(area -1/2) as the area is increased. On the other hand, non-zero efficiency is possible for increasing density depending on the behavior of the physical correlation as a function of the link length.
Youngchul Sung, H. Vincent Poor, Heejung Yu
ICASSP1
2008 On optimal operating characteristics of sensing and training for cognitive radios
abstract
The problem of optimal sensing and training in a cognitive radio system is considered when the training signal of the primary transmitter is used for both channel estimation at the primary receiver and sensing for the secondary transmitter. First, the optimal operating characteristics of sensing that maximizes the overall system rate for given training is investigated. It is shown that the optimal false alarm probability at the secondary sensor is monotone increasing as the activity of the primary user increases if the sensing ROC curve is concave. When the primary activity factor is unknown, the max-min criterion is applied to optimal sensing strategy and the resulting max-min optimal solution is given by an equalizer rule for any type of sensing ROC curve. The joint optimization of sensing and training has a unique solution and it can be easily found numerically using a gradient ascent algorithm. By optimal design of sensing and training in such a way, the overall system rate can be improved.
Heejung Yu, Youngchul Sung, Yong Hoon Lee
ICASSP2
2008 Information, energy and density for Ad Hoc sensor networks over correlated random fields: Large deviations analysis
abstract
Using large deviations results that characterize the amount of information per node on a two-dimensional (2-D) lattice, asymptotic behavior of a sensor network deployed over a correlated random field for statistical inference is investigated. Under a 2-D hidden Gauss-Markov random field model with symmetric first order conditional autoregression, the behavior of the total information [nats] and energy efficiency [nats/J] defined as the ratio of total gathered information to the required energy is obtained as the coverage area, node density and energy vary.
Youngchul Sung, Heejung Yu, H. Vincent Poor
ISIT1
2007 Cooperative routing for distributed detection in large sensor networks
abstract
In this paper, the detection of a correlated Gaussian field using a large multi-hop sensor network is investigated. A cooperative routing strategy is proposed by introducing a new link metric that characterizes the detection error exponent. Derived from the Chernoff information and Schweppe's likelihood recursion, this link metric captures the contribution of a given link to the decay rate of error probability and has the form of the capacity of a Gaussian channel with the sender transmitting the innovation of its measurement. For one-dimensional Gauss-Markov fields, the link metric can be represented explicitly as a function of the link length. Cooperative routing is achieved using the Kalman data aggregation and shortest path routing. Numerical simulations show that cooperative routing can be significantly more energy efficient than noncooperative routing for the same detection performance
Youngchul Sung, Saswat Misra, Lang Tong 0001, Anthony Ephremides
IEEE J. Sel. Areas Commun.1
2006 Neyman-pearson detection of gauss-Markov signals in noise: closed-form error exponentand properties
abstract
The performance of Neyman-Pearson detection of correlated random signals using noisy observations is considered. Using the large deviations principle, the performance is analyzed via the error exponent for the miss probability with a fixed false-alarm probability. Using the state-space structure of the signal and observation model, a closed-form expression for the error exponent is derived using the innovations approach, and the connection between the asymptotic behavior of the optimal detector and that of the Kalman filter is established. The properties of the error exponent are investigated for the scalar case. It is shown that the error exponent has distinct characteristics with respect to correlation strength: for signal-to-noise ratio (SNR) /spl ges/1, the error exponent is monotonically decreasing as the correlation becomes strong whereas for SNR<1 there is an optimal correlation that maximizes the error exponent for a given SNR.
Youngchul Sung, Lang Tong 0001, H. Vincent Poor
IEEE Trans. Inf. Theory1
2005 A large deviations approach to sensor scheduling for detection of correlated random fields
abstract
The problem of scheduling sensor transmissions for the detection of correlated random fields using spatially deployed sensors is considered. Using the large deviations principle, a closed-form expression for the error exponent of the miss probability is given as a function of the sensor spacing and signal-to-noise ratio (SNR). It is shown that the error exponent has a distinct characteristic: at high SNR, the error exponent monotonically increases with respect to sensor spacing, while at low SNR, there is an optimal spacing for scheduled sensors.
Youngchul Sung, Lang Tong 0001, H. Vincent Poor
ICASSP (3)1
2005 Sensor configuration and activation for field detection in large sensor arrays
abstract
The problems of sensor configuration and activation for the detection of correlated random fields using large sensor arrays are considered. Using results that characterize the large-array performance of sensor networks in this application, the detection capabilities of different sensor configurations are analyzed and compared. The dependence of the optimal choice of configuration on parameters such as sensor signal-to-noise ratio (SNR), field correlation, etc., is examined, yielding insights into the most effective choices for sensor selection and activation in various operating regimes.
Youngchul Sung, Lang Tong 0001, H. Vincent Poor
IPSN1
2005 Neyman-Pearson detection of Gauss-Markov signals in noise: closed-form error exponent and properties
abstract
The performance of Neyman-Pearson detection of correlated stochastic signals using noisy observations is investigated via the error exponent for the miss probability with a fixed level. Using the state-space structure of the signal and observation model, a closed-form expression for the error exponent is derived, and the connection between the asymptotic behavior of the optimal detector and that of the Kalman filter is established. The properties of the error exponent are investigated for the scalar case. It is shown that the error exponent has distinct characteristics with respect to correlation strength: for signal-to-noise ratio (SNR) > 1 the error exponent decreases monotonically as the correlation becomes stronger, whereas for SNR < 1 there is an optimal correlation that maximizes the error exponent for a given SNR
Youngchul Sung, Lang Tong 0001, H. Vincent Poor
ISIT1
2004 Asymptotic locally optimal detector for large-scale sensor networks under the Poisson regime
abstract
We consider distributed detection with a large number of identical sensors deployed over a region where the phenomenon of interest (POI) has unknown spatially varying strength. Each sensor makes a decision based on its own measurement of the signal at its location and the local decision of each sensor is sent to a fusion center through a multiple access channel. The fusion center decides whether the POI has occurred in the region, under a global size constraint in the Neyman-Pearson formulation. Assuming that the initial distribution of sensors is a homogeneous spatial Poisson process, we show that the Poisson process of 'alarmed' sensors satisfies the locally asymptotic normality (LAN) condition as the number of sensor goes to infinity. We derive a new asymptotically locally most powerful (ALMP) detector jointly over the fusion scheme and the sensor threshold. We also derive the conditions on the spatial signal shape to guarantee the existence of the ALMP detector. We show that the optimal test statistic is a weighted sum of local decisions, the optimal weight function being the shape of the spatial signal, but the exact value of the spatial signal is not required. The optimal threshold for a single sensor is also derived. For the case of independent, identically-distributed (i.i.d.) sensor observation, we show that the counting-based detector is also asymptotic locally optimal.
Youngchul Sung, Lang Tong 0001, Ananthram Swami
ICASSP (2)1
2004 Asymptotic locally optimal detector for large-scale sensor networks under Poisson regime
abstract
We consider the distributed detection problem with a large number if identical sensors deployed over a region where the phenomenon of interest (POI) has different signal strength depending on the location. Each sensor makes a decision based on its own measurement of the spatially varying signal and the local decision of each sensor is sent to a fusion center through a multiple access channel. The fusion center decides whether the POI has occurred in the region, under a global size constraint in the Neyman-Pearson formulation. Assuming that the initial distribution of sensors is a homogeneous spatial Poisson process, we show that the Poisson process of 'alarmed' sensors satisfies the locally asymptotic normality (LAN) condition as the number of sensor goes to infinity and derive a new asymptotically locally most powerful detector for the spatially varying signal. We show that (1) an optimal test statistic is a weighted sum of local decisions, (2) the optimal weight function is the shape of the spatial signal, and (3) the exact value of the spatial signal is not required. For the case of independent, identical distributed (i.i.d.) sensor observation, we show that the counting-based detector is also asymptotic locally optimal.
Youngchul Sung, Lang Tong 0001, Ananthram Swami
ICC1
2003 Semiblind channel estimation for space-time coded WCDMA
abstract
A new semiblind channel estimation technique is proposed for space-time coded wideband CDMA systems using aperiodic and possibly multirate spreading codes. Using a decorrelating matched filter, the received signal is projected onto subspaces from which channel parameters and data symbols can be estimated jointly. Exploiting the subspace structure of the WCDMA signaling and the orthogonality of the space-time code, the proposed algorithm provides the least squares channel estimate in closed form. A new identifiability condition is established. The mean square error of the estimated channel is compared with the Cramer-Rao bound, and a bit error rate (BER) expression for the proposed algorithm is compared with differential schemes.
Youngchul Sung, Lang Tong 0001, Ananthram Swami
ICC1
2002 A projection-based semi-blind channel estimation for long-code WCDMA
abstract
The third generation (3G) wireless systems such as Wideband CDMA(WCDMA) employ coherent detection where pilot symbols for channel estimation are transmitted simultaneously with data using orthogonal channelization codes. The orthogonality of codes, unfortunately, is lost easily due to multipath delays. Furthermore, the use of long scrambling code introduces time variation in channel model. In this paper, based on the projection of time-varying subspaces, we present a semi-blind channel estimation technique for long code CDMA systems that use code division multiplexed pilot symbols.
Youngchul Sung, Lang Tong 0001
ICASSP1