Wonjong Rhee

dblp:37/711 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 since 2021Computer networks · 6 · 3 first-authorTheory of computation · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
abstract
Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation.
Dongnam Byun, Jungwon Park, Jungmin Ko, Changin Choi, Wonjong Rhee
AAAI5
2026 Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
abstract
Knowledge-intensive visual question answering (VQA) requires external knowledge beyond image content, demanding precise visual grounding and coherent integration of visual and textual information.Although multimodal retrieval-augmented generation has achieved notable advances by incorporating external knowledge bases, existing approaches largely adopt single-pass frameworks that often fail to acquire sufficient knowledge and lack mechanisms to revise misdirected reasoning.We propose PMSR (Progressive Multimodal Search and Reasoning), a framework that progressively constructs a structured reasoning trajectory to enhance both knowledge acquisition and synthesis.PMSR uses dual-scope queries conditioned on both the latest record and the trajectory to retrieve diverse knowledge from heterogeneous knowledge bases.The retrieved evidence is then synthesized into compact records via compositional reasoning.This design facilitates controlled iterative refinement, which supports more stable reasoning trajectories with reduced error propagation.Extensive experiments across six diverse benchmarks (Encyclopedic-VQA, InfoSeek, MMSearch, LiveVQA, FVQA, and OK-VQA) demonstrate that PMSR consistently improves both retrieval recall and end-to-end answer accuracy.
Changin Choi, Wonseok Lee 0002, Jungmin Ko, Wonjong Rhee
ACL (1)4
2025 Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CDFSL) methods typically parameterize models with task-agnostic and task-specific parameters. To adapt task-specific parameters, recent approaches have utilized fixed optimization strategies, despite their potential sub-optimality across varying domains or target tasks. To address this issue, we propose a novel adaptation mechanism called Task-Specific Preconditioned gradient descent (TSP). Our method first meta-learns Domain-Specific Preconditioners (DSPs) that capture the characteristics of each meta-training domain, which are then linearly combined using task-coefficients to form the Task-Specific Preconditioner. The preconditioner is applied to gradient descent, making the optimization adaptive to the target task. We constrain our preconditioners to be positive definite, guiding the preconditioned gradient toward the direction of steepest descent. Empirical evaluations on the Meta-Dataset show that TSP achieves state-of-the-art performance across diverse experimental scenarios.
Suhyun Kang, Jungwon Park, Wonseok Lee 0002, Wonjong Rhee
AAAI4
2025 Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
abstract
In a surge of text-to-image (T2I) models and their customization methods that generate new images of a user-provided subject, current works focus on alleviating the costs incurred by a lengthy per-subject optimization. These zero-shot customization methods encode the image of a specified subject into a visual embedding which is then utilized alongside the textual embedding for diffusion guidance. The visual embedding incorporates intrinsic information about the subject, while the textual embedding provides a new context. However, the existing methods often 1) generate images with the same pose as an input image, and 2) exhibit deterioration in the subject's identity when facing a pose variation prompt. We first pin down the problem and show that redundant pose information in the visual embedding interferes with the pose indication in the textual embedding. Conversely, the textual embedding also harms the subject's identity which is tightly entangled with the pose in the visual embedding. As a remedy, we propose text-orthogonal visual embedding which effectively harmonizes with the given textual embedding. We also adopt the visual-only embedding and inject the subject's clear features utilizing a self-attention swap. Our method is both effective and robust, offering highly flexible zero-shot generation while effectively maintaining the subject's identity.
Yeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin, Wonjong Rhee, Nojun Kwak
AAAI5
2025 ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
abstract
Rectified Flow text-to-image models surpass diffusion models in image quality and text alignment, but adapting ReFlow for real-image editing remains challenging. We propose a new real-image editing method for ReFlow by analyzing the intermediate representations of multimodal transformer blocks and identifying three key features. To extract these features from real images with sufficient structural preservation, we leverage mid-step latent, which is inverted only up to the mid-step. We then adapt attention during injection to improve editability and enhance alignment to the target text. Our method is training-free, requires no user-provided mask, and can be applied even without a source prompt. Extensive experiments on two benchmarks with nine baselines demonstrate its superior performance over prior methods, further validated by human evaluations confirming a strong user preference for our approach.
Jimyeong Kim, Jungwon Park, Yeji Song, Nojun Kwak, Wonjong Rhee
ICCV5
2025 Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models
abstract
Recent text-to-image diffusion models leverage cross-attention layers, which have been effectively utilized to enhance a range of visual generative tasks. However, our understanding of cross-attention layers remains somewhat limited. In this study, we introduce a mechanistic interpretability approach for diffusion models by constructing Head Relevance Vectors (HRVs) that align with human-specified visual concepts. An HRV for a given visual concept has a length equal to the total number of cross-attention heads, with each element representing the importance of the corresponding head for the given visual concept. To validate HRVs as interpretable features, we develop an ordered weakening analysis that demonstrates their effectiveness. Furthermore, we propose concept strengthening and concept adjusting methods and apply them to enhance three visual generative tasks. Our results show that HRVs can reduce misinterpretations of polysemous words in image generation, successfully modify five challenging attributes in image editing, and mitigate catastrophic neglect in multi-concept generation. Overall, our work provides an advancement in understanding cross-attention layers and introduces new approaches for fine-controlling these layers at the head level.
Jungwon Park, Jungmin Ko, Dongnam Byun, Jangwon Suh, Wonjong Rhee
ICLR5
2025 Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning
Changin Choi, Wonjong Rhee
INTERSPEECH3
2025 Towards a better evaluation of out-of-domain generalization
Duhun Hwang, Suhyun Kang, Moonjung Eo, Jimyeong Kim, Wonjong Rhee
Neural Networks5
2025 Improving forward compatibility in class incremental learning by increasing representation rank and feature richness
Jaeill Kim, Wonseok Lee 0002, Moonjung Eo, Wonjong Rhee
Neural Networks4
2024 Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
abstract
In text-to-image personalization, a timely and crucial challenge is the tendency of generated images overfitting to the biases present in the reference images. We initiate our study with a comprehensive categorization of the biases into background, nearby-object, tied-object, substance (in style re-contextualization), and pose biases. These biases manifest in the generated images due to their entanglement into the subject embedding. This undesired embedding entanglement not only results in the reflection of biases from the reference images into the generated images but also notably diminishes the alignment of the generated images with the given generation prompt. To address this challenge, we propose SID (Selectively Informative Description), a text description strategy that deviates from the prevalent approach of only characterizing the subject's class identification. SID is generated utilizing multimodal GPT-4 and can be seamlessly integrated into optimization-based models. We present comprehensive experimental results along with analyses of cross-attention maps, subject-alignment, non-subject-disentanglement, and text-alignment.
Jimyeong Kim, Jungwon Park, Wonjong Rhee
CVPR3
2024 A Benchmark Suite for Evaluating Neural Mutual Information Estimators on Unstructured Datasets
abstract
Mutual Information (MI) is a fundamental metric for quantifying dependency between two random variables. When we can access only the samples, but not the underlying distribution functions, we can evaluate MI using sample-based estimators. Assessment of such MI estimators, however, has almost always relied on analytical datasets including Gaussian multivariates. Such datasets allow analytical calculations of the true MI values, but they are limited in that they do not reflect the complexities of real-world datasets. This study introduces a comprehensive benchmark suite for evaluating neural MI estimators on unstructured datasets, specifically focusing on images and texts. By leveraging same-class sampling for positive pairing and introducing a binary symmetric channel trick, we show that we can accurately manipulate true MI values of real-world datasets. Using the benchmark suite, we investigate seven challenging scenarios, shedding light on the reliability of neural MI estimators for unstructured datasets.
Kyungeun Lee, Wonjong Rhee
NeurIPS2
2024 Towards a rigorous analysis of mutual information in contrastive learning
Kyungeun Lee, Jaeill Kim, Suhyun Kang, Wonjong Rhee
Neural Networks4
2023 Diversified and Realistic 3D Augmentation via Iterative Construction, Random Placement, and HPR Occlusion
abstract
In autonomous driving, data augmentation is commonly used for improving 3D object detection. The most basic methods include insertion of copied objects and rotation and scaling of the entire training frame. Numerous variants have been developed as well. The existing methods, however, are considerably limited when compared to the variety of the real world possibilities. In this work, we develop a diversified and realistic augmentation method that can flexibly construct a whole-body object, freely locate and rotate the object, and apply self-occlusion and external-occlusion accordingly. To improve the diversity of the whole-body object construction, we develop an iterative method that stochastically combines multiple objects observed from the real world into a single object. Unlike the existing augmentation methods, the constructed objects can be randomly located and rotated in the training frame because proper occlusions can be reflected to the whole-body objects in the final step. Finally, proper self-occlusion at each local object level and external-occlusion at the global frame level are applied using the Hidden Point Removal (HPR) algorithm that is computationally efficient. HPR is also used for adaptively controlling the point density of each object according to the object's distance from the LiDAR. Experiment results show that the proposed DR.CPO algorithm is data-efficient and model-agnostic without incurring any computational overhead. Also, DR.CPO can improve mAP performance by 2.08% when compared to the best 3D detection result known for KITTI dataset.
Jungwook Shin, Jaeill Kim, Kyungeun Lee, Hyunghun Cho, Wonjong Rhee
AAAI5
2023 Meta-Learning with a Geometry-Adaptive Preconditioner
abstract
Model-agnostic meta-learning (MAML) is one of the most successful meta-learning algorithms. It has a bi-level optimization structure where the outer-loop process learns a shared initialization and the inner-loop process optimizes task-specific weights. Although MAML relies on the standard gradient descent in the inner-loop, recent studies have shown that controlling the inner-loop's gradient descent with a meta-learned preconditioner can be beneficial. Existing preconditioners, however, cannot simultaneously adapt in a task-specific and path-dependent way. Additionally, they do not satisfy the Riemannian metric condition, which can enable the steepest descent learning with preconditioned gradient. In this study, we propose Geometry-Adaptive Preconditioned gradient descent (GAP) that can overcome the limitations in MAML; GAP can efficiently meta-learn a preconditioner that is dependent on task-specific parameters, and its preconditioner can be shown to be a Riemannian metric. Thanks to the two properties, the geometry-adaptive preconditioner is effective for improving the inner-loop optimization. Experiment results show that GAP outperforms the state-of-the-art MAML family and preconditioned gradient descent-MAML (PGD-MAML) family in a variety of few-shot learning tasks. Code is available at: https://github.com/Suhyun777/CVPR23-GAP.
Suhyun Kang, Duhun Hwang, Moonjung Eo, Taesup Kim, Wonjong Rhee
CVPR5
2023 VNE: An Effective Method for Improving Deep Representation by Manipulating Eigenvalue Distribution
abstract
Since the introduction of deep learning, a wide scope of representation properties, such as decorrelation, whitening, disentanglement, rank, isotropy, and mutual information, have been studied to improve the quality of representation. However, manipulating such properties can be challenging in terms of implementational effectiveness and general applicability. To address these limitations, we propose to regularize von Neumann entropy (VNE) of representation. First, we demonstrate that the mathematical formulation of VNE is superior in effectively manipulating the eigenvalues of the representation autocorrelation matrix. Then, we demonstrate that it is widely applicable in improving state-of-the-art algorithms or popular benchmark algorithms by investigating domain-generalization, meta-learning, self-supervised learning, and generative models. In addition, we formally establish theoretical connections with rank, disentanglement, and isotropy of representation. Finally, we provide discussions on the dimension control of VNE and the relationship with Shannon entropy. Code is available at: https://github.com/jaeill/CVPR23-VNE.
Jaeill Kim, Suhyun Kang, Duhun Hwang, Jungwook Shin, Wonjong Rhee
CVPR5
2023 Isotropic Representation Can Improve Dense Retrieval
Euna Jung, Jungwon Park, Jaekeol Choi, Sungyoon Kim, Wonjong Rhee
PAKDD (3)5
2023 AID-purifier: A light auxiliary network for boosting adversarial defense
Duhun Hwang, Wonjong Rhee
Neurocomputing3
2023 An effective low-rank compression with a joint rank selection followed by a compression-friendly training
Moonjung Eo, Suhyun Kang, Wonjong Rhee
Neural Networks3
2022 AID-Purifier: A Light Auxiliary Network for Boosting Adversarial Defense
abstract
In this study, we propose AID-Purifier that can boost the robustness of adversarially-trained networks by purifying their inputs. AID-Purifier is an auxiliary network that works as an add-on to an already trained main classifier. To keep it computationally light, it is trained as a discriminator with a binary cross-entropy loss. To obtain additionally useful information from the adversarial examples, the architecture design is closely related to the information maximization principle where two layers of the main classification network are piped into the auxiliary network. To assist the iterative optimization procedure of purification, the auxiliary network is trained with AVmixup. AID-Purifier can be also used together with other purifiers such as PixelDefend for an extra enhancement. Because input purification has been studied relative less when compared to adversarial training or gradient masking, we conduct extensive attack experiments to validate AID-Purifier’s robustness. The overall results indicate that the best performing adversarially-trained networks can be enhanced further with AID-Purifier. The code is available in https://github.com/yelobean/AIDPurifier.
Duhun Hwang, Wonjong Rhee
ICPR3
2022 Semi-Siamese Bi-encoder Neural Ranking Model Using Lightweight Fine-Tuning
abstract
A BERT-based Neural Ranking Model (NRM) can be either a cross-encoder or a bi-encoder. Between the two, bi-encoder is highly efficient because all the documents can be pre-processed before the actual query time. In this work, we show two approaches for improving the performance of BERT-based bi-encoders. The first approach is to replace the full fine-tuning step with a lightweight fine-tuning. We examine lightweight fine-tuning methods that are adapter-based, prompt-based, and hybrid of the two. The second approach is to develop semi-Siamese models where queries and documents are handled with a limited amount of difference. The limited difference is realized by learning two lightweight fine-tuning modules, where the main language model of BERT is kept common for both query and document. We provide extensive experiment results for monoBERT, TwinBERT, and ColBERT where three performance metrics are evaluated over Robust04, ClueWeb09b, and MS-MARCO datasets. The results confirm that both lightweight fine-tuning and semi-Siamese are considerably helpful for improving BERT-based bi-encoders. In fact, lightweight fine-tuning is helpful for cross-encoder, too.1
Euna Jung, Jaekeol Choi, Wonjong Rhee
WWW3
2021 "I wrote as if I were telling a story to someone I knew.": Designing Chatbot Interactions for Expressive Writing in Mental Health
abstract
Writing about experiences of trauma and other challenges in life is known to provide measurable health benefits. Though writing for an audience may ensure better benefits, confiding one's most troubled memories in others risks a social stigma. Conversational agents can provide a virtual audience that ensures privacy and allows social disclosure. To understand the writing experience with an agent, we created Diarybot, a chatbot assistant for expressive writing. We designed two versions, Basic and Responsive, to explore the writing experience with and without bot follow-up interactions compared to a Google doc baseline. Findings from a 4-day user study with 30 participants reveal that social disclosure with Diarybot can encourage narrative writing, with relative ease and emotional expression in Basic chat. Responsive chat can mediate social acceptance of the bot and provide guidance for self-reflection in the process. We discuss design reflections on social disclosure with agents in pursuit of wellbeing.
SoHyun Park, Anja Thieme, Jeongyun Han, Sungwoo Lee, Wonjong Rhee, Bongwon Suh
Conference on Designing Interactive Systems5
2021 Statistical Characteristics of Deep Representations: An Empirical Investigation
Daeyoung Choi, Kyungeun Lee, Duhun Hwang, Wonjong Rhee
ICANN (5)4
2021 Improving Bi-encoder Document Ranking Models with Two Rankers and Multi-teacher Distillation
abstract
BERT-based Neural Ranking Models (NRMs) can be classified according to how the query and document are encoded through BERT's self-attention layers - bi-encoder versus cross-encoder. Bi-encoder models are highly efficient because all the documents can be pre-processed before the query time, but their performance is inferior compared to cross-encoder models. Both models utilize a ranker that receives BERT representations as the input and generates a relevance score as the output. In this work, we propose a method where multi-teacher distillation is applied to a cross-encoder NRM and a bi-encoder NRM to produce a bi-encoder NRM with two rankers. The resulting student bi-encoder achieves an improved performance by simultaneously learning from a cross-encoder teacher and a bi-encoder teacher and also by combining relevance scores from the two rankers. We call this method TRMD (Two Rankers and Multi-teacher Distillation). In the experiments, TwinBERT and ColBERT are considered as baseline bi-encoders. When monoBERT is used as the cross-encoder teacher, together with either TwinBERT or ColBERT as the bi-encoder teacher, TRMD produces a student bi-encoder that performs better than the corresponding baseline bi-encoder. For [email protected], the maximum improvement was 11.4%, and the average improvement was 6.8%. As an additional experiment, we considered producing cross-encoder students with TRMD, and found that it could also improve the cross-encoders.
Jaekeol Choi, Euna Jung, Jangwon Suh, Wonjong Rhee
SIGIR4
2019 Utilizing Class Information for Deep Network Representation Shaping
abstract
Statistical characteristics of deep network representations, such as sparsity and correlation, are known to be relevant to the performance and interpretability of deep learning. When a statistical characteristic is desired, often an adequate regularizer can be designed and applied during the training phase. Typically, such a regularizer aims to manipulate a statistical characteristic over all classes together. For classification tasks, however, it might be advantageous to enforce the desired characteristic per class such that different classes can be better distinguished. Motivated by the idea, we design two class-wise regularizers that explicitly utilize class information: class-wise Covariance Regularizer (cw-CR) and classwise Variance Regularizer (cw-VR). cw-CR targets to reduce the covariance of representations calculated from the same class samples for encouraging feature independence. cw-VR is similar, but variance instead of covariance is targeted to improve feature compactness. For the sake of completeness, their counterparts without using class information, Covariance Regularizer (CR) and Variance Regularizer (VR), are considered together. The four regularizers are conceptually simple and computationally very efficient, and the visualization shows that the regularizers indeed perform distinct representation shaping. In terms of classification performance, significant improvements over the baseline and L1/L2 weight regularization methods were found for 21 out of 22 tasks over popular benchmark datasets. In particular, cw-VR achieved the best performance for 13 tasks including ResNet-32/110.
Daeyoung Choi, Wonjong Rhee
AAAI2
2019 Subtask Gated Networks for Non-Intrusive Load Monitoring
abstract
Non-intrusive load monitoring (NILM), also known as energy disaggregation, is a blind source separation problem where a household’s aggregate electricity consumption is broken down into electricity usages of individual appliances. In this way, the cost and trouble of installing many measurement devices over numerous household appliances can be avoided, and only one device needs to be installed. The problem has been well-known since Hart’s seminal paper in 1992, and recently significant performance improvements have been achieved by adopting deep networks. In this work, we focus on the idea that appliances have on/off states, and develop a deep network for further performance improvements. Specifically, we propose a subtask gated network that combines the main regression network with an on/off classification subtask network. Unlike typical multitask learning algorithms where multiple tasks simply share the network parameters to take advantage of the relevance among tasks, the subtask gated network multiply the main network’s regression output with the subtask’s classification probability. When standby-power is additionally learned, the proposed solution surpasses the state-of-the-art performance for most of the benchmark cases. The subtask gated network can be very effective for any problem that inherently has on/off states.
Changho Shin, Sunghwan Joo, Jaeryun Yim, Hyoseop Lee, Taesup Moon, Wonjong Rhee
AAAI6
2018 On the Difficulty of DNN Hyperparameter Optimization Using Learning Curve Prediction
abstract
With the recent success of deep learning on a variety of applications, efficiently tuning hyperparameters of Deep Neural Networks (DNNs) with less effort has become a timely and practical topic. As an algorithmic solution, automatic hyperparameter optimization methods like Bayesian optimization have gained popularity for achieving human-comparable or even human-surpassing performance. To further speed up hyperparameter optimization, learning curves of DNNs can be predicted and used to early terminate the training phase of the chosen hyperparameter setting when the expected training performance is not satisfactory. While the previous studies show promising results, it is still unclear if an effective general rule can be derived for a broad spectrum of DNN hyperparameter optimization problems. In this work, we consider hyperparameter optimization of MNIST and CIFAR-10, and for each task, we analyze the characteristics of the 20,000 learning curves that correspond to the 20,000 different hyperparameter configurations. By investigating a large number of learning curves for a given task, we find that the characteristics of learning curve shapes can drastically change depending on the choice and range of hyperparameters. Therefore, utilizing learning curves for speed improvement is not a simple task and can be dependent on many factors. Based on the observations and analyses on the 20,000 learning curves, we design two early termination rules, ETR-1 and ETR-2, and show that the rules can be beneficial in the best case but can be harmful as well. Our observations and experimental results highlight that hyperparameter optimization of DNNs using learning curve prediction is challenging. In particular, the results of recent studies that are based on at most thousands of learning curves of a limited number of tasks should be carefully interpreted depending on the task, DNN model, hyperparameter choice, and hyperparameter range.
Daeyoung Choi, Hyunghun Cho, Wonjong Rhee
TENCON3
2015 N-Polar Visualization: Visual Analytics for Exploring Data Objects with Multiple Interactive Anchors
Taeil Jeon, Wonjong Rhee, Bongwon Suh
IV3
2012 Distributed Crosstalk Management for Upstream VDSL Using Dynamic Power Control
abstract
The severe interference from neighbor copper lines, commonly known as crosstalk, is a well-known limitation that can reduce the upstream rate of a victim user by 50% or more in dense VDSL (Very high bit-rate Digital Subscriber Lines) deployment. This problem has been partially addressed by the use of UPBO (Upstream Power BackOff). UPBO blindly controls transmit PSD (Power Spectral Densities) based on the pre-defined parameters and channel measurements without considering each user's target data rate. Although UPBO can provide fair protection against crosstalk, it still leaves room for practical improvement. This paper proposes a distributed algorithm that dynamically controls transmit power according to each line's need and capacity. The algorithm effectively improves data rates and line reaches in real VDSL deployment.
Wonjong Rhee, Mehdi Mohseni, John M. Cioffi
IEEE Trans. Commun.2
2011 Practical Crosstalk Management for Upstream VDSL Using Dynamic Power Control
abstract
The severe interference from neighbor copper lines, commonly known as crosstalk, is a well-known limitation that can cause 50% or larger reduction in the upstream rates of dense VDSL (Very high bit-rate Digital Subscriber Lines) deployments. To handle this problem, UPBO (Upstream Power BackOff) has been studied in literature, and subsequently standardized. UPBO blindly selects transmit PSD (Power Spectral Densities) based on the pre-defined parameters and channel measurement without considering each line's target data rate. Although UPBO can provide significant protection against crosstalk, it still leaves room for practical improvement. This paper proposes a distributed algorithm that dynamically controls transmit power based on each line's need and capability. The algorithm can achieve significant rate and reach gains in real VDSL deployments.
Wonjong Rhee, Mehdi Mohseni, John M. Cioffi
GLOBECOM2
2006 On the Ergodic Capacity-Achieving Covariance Matrix of Certain Classes of MIMO Channels
abstract
Multiple-input multiple-output (MIMO) flat fading channels with channel state information available at the receiver and channel state distribution available at the transmitter are considered. The ergodic capacity-achieving covariance matrix (ECACM) under an average input power constraint is shown to be not unique in general, and three sufficient conditions for the uniqueness are derived. Then, the ECACM structure is characterized for certain classes of MIMO channels including the following classes of MIMO channels: i) right-permutationally-invariant channels including i.i.d. channels; ii) right-symmetrically-distributed channels; and iii) right-permutationally-invariant and right-symmetrically-distributed channels
Wonjong Rhee, Giorgio Taricco
ISIT1
2006 On the Ergodic Capacity-Achieving Covariance Matrix of Certain Classes of MIMO Channels
abstract
Multiple-input multiple-output (MIMO) flat-fading channels with channel state information available at the receiver and channel state distribution available at the transmitter are considered. The ergodic capacity-achieving covariance matrix (ECACM) under an average input power constraint is shown to be not unique in general, and three sufficient conditions for the uniqueness are derived. Then, the structure of ECACM is characterized for certain classes of MIMO channels including the following classes of MIMO channels: 1) right-permutationally invariant channels including independent and identically distributed (i.i.d.) channels; 2)right-symmetrically distributed channels; and 3) right-permutationally invariant and right-symmetrically distributed channels. Finally, part of the single-user results are extended to the multiple-access channel.
Wonjong Rhee, Giorgio Taricco
IEEE Trans. Inf. Theory1
2005 Sum power iterative water-filling for multi-antenna Gaussian broadcast channels
abstract
In this correspondence, we consider the problem of maximizing sum rate of a multiple-antenna Gaussian broadcast channel (BC). It was recently found that dirty-paper coding is capacity achieving for this channel. In order to achieve capacity, the optimal transmission policy (i.e., the optimal transmit covariance structure) given the channel conditions and power constraint must be found. However, obtaining the optimal transmission policy when employing dirty-paper coding is a computationally complex nonconvex problem. We use duality to transform this problem into a well-structured convex multiple-access channel (MAC) problem. We exploit the structure of this problem and derive simple and fast iterative algorithms that provide the optimum transmission policies for the MAC, which can easily be mapped to the optimal BC policies.
Nihar Jindal, Wonjong Rhee, Sriram Vishwanath, Syed Ali Jafar, Andrea J. Goldsmith
IEEE Trans. Inf. Theory2
2004 Iterative water-filling for Gaussian vector multiple-access channels
abstract
This paper proposes an efficient numerical algorithm to compute the optimal input distribution that maximizes the sum capacity of a Gaussian multiple-access channel with vector inputs and a vector output. The numerical algorithm has an iterative water-filling interpretation. The algorithm converges from any starting point, and it reaches within 1/2 nats per user per output dimension from the sum capacity after just one iteration. The characterization of sum capacity also allows an upper bound and a lower bound for the entire capacity region to be derived.
Wei Yu 0001, Wonjong Rhee, Stephen P. Boyd, John M. Cioffi
IEEE Trans. Inf. Theory2
2004 The optimality of beamforming in uplink multiuser wireless systems
abstract
This paper considers the optimal uplink transmission strategy that achieves the sum-capacity in a multiuser multi-antenna wireless system. Assuming an independent identically distributed block-fading model with transmitter channel side information, beamforming for each remote user is shown to be necessary for achieving sum-capacity when there is a large number of users in the system. This result stands even in the case where each user is equipped with a large number of transmit antennas, and it can be readily extended to channels with intersymbol interference if an orthogonal frequency division multiplexing modulation is assumed. This result is obtained by deriving a rank bound on the transmit covariance matrices, and it suggests that all users should cooperate by each user using only a small portion of available dimensions. Based on the result, a suboptimal transmit scheme is proposed for the situation where only partial channel side information is available at each transmitter. Simulations show that the suboptimal scheme is not only able to achieve a sum rate very close to the capacity, but also insensitive to channel estimation error.
Wonjong Rhee, Wei Yu 0001, John M. Cioffi
IEEE Trans. Wirel. Commun.1
2003 On the capacity of multiuser wireless channels with multiple antennas
abstract
The advantages of multiuser communication, where many users are allowed to simultaneously transmit or receive in a common bandwidth, are considered for multiple-antenna systems in a high signal-to-noise ratio (SNR) regime. Assuming channel state information at receiver (CSIR) to be available, the ergodic capacity is characterized for both unbiased and biased channels, and the quantitative capacity gain of a multiple-antenna multiuser system is analyzed for multiple-access channels. For highly biased (correlated) channels, a multiuser system is shown to be inherently superior to a single-user system (a time- or frequency-division multiple-access (TDMA or FDMA) based system) due to the underlying multiuser diversity, and the sum capacity is shown to scale linearly with the number of antennas. For unbiased channels, the characteristics of ergodic capacity are shown to transfer to outage capacity when a large degree of space diversity exists, and to deterministic capacity when the number of receive antennas is large. Also, a brief discussion on the multiuser multiple-antenna communication in broadcast channel is provided.
Wonjong Rhee, John M. Cioffi
IEEE Trans. Inf. Theory1
2001 On the asymptotic optimality of beam-forming in multi-antenna Gaussian multiple access channels
abstract
In this paper, transmit schemes for multi-antenna Gaussian multiple access channels are considered. For the case of a large number of users, a short term power constraint, and slowly fading channels, asymptotic optimality of beamforming is shown under the sum rate maximization criterion. This is an interesting result because the optimal transmit scheme for each user is to beamform to one direction even if each user is equipped with a large number of transmit antennas. Also, this result extends to ISI channels assuming OFDM modulation. A suboptimal transmit scheme based on this result follows for a system with partial channel side information.
Wonjong Rhee, John M. Cioffi
GLOBECOM1
2001 Optimal power control in multiple access fading channels with multiple antennas
abstract
This paper characterizes the optimal power control method for maximum sum capacity in a multiple access fading channel with multiple transmitter and receiver antennas when perfect channel side information is available at both the transmitters and the receiver. The profound benefit of multi-antenna diversity is demonstrated by a dimension counting argument. The optimal power allocation strategy in a system with n transmit antennas for each user and m receive antennas is a combination of successive cancellation and a TDMA-like scheme where in each time slot the rank of the transmit signals r/sub k/ for all users must satisfy /spl Sigma//sub k/r/sub k/(r/sub k/+1)/spl les/m(m+1). Thus, the total number of users that are allowed to transmit simultaneously is constrained by the number of receiver antennas. Receiver diversity increases the total number of dimensions thus allowing more users to transmit at the same time. By contrast, transmitter diversity allows a single user to occupy multiple dimensions as to benefit its own transmission, thus having the effect of precluding simultaneous transmission by other users.
Wei Yu 0001, Wonjong Rhee, John M. Cioffi
ICC2
2000 Utilizing multiuser diversity for multiple antenna systems
abstract
Previous research has shown that the capacity of a multiple antenna system grows linearly with increasing number of antennas for rich-scattering environments. However, this is not true for wireless channels with a small number of independent paths. To overcome this problem, this paper investigates the possibility of exploiting the multiuser dimension with and without channel side information at the transmitter. First, the single user capacity per antenna is shown to converge to zero with increasing number of antennas for channels with a finite number of independent paths. Then multiuser capacity per antenna at the limit is shown to be positive. Simulation results are presented for a single user system and a multiuser uplink system.
Wonjong Rhee, Wei Yu 0001, John M. Cioffi
WCNC1