Chien-Yi Wang

dblp:12/6741 · DBLP profile ↗
← Back
39ranked-venue papers
14as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-authorTheory of computation · 8 · 5 first-authorComputer networks · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 MCPNet++: Interpretable Classification Models via Multi-Level Concept Prototypes
abstract
Post-hoc and inherently interpretable methods have shown great success in uncovering the inner workings of black-box models, whether by examining them after training or by explicitly designing for interpretability. While these approaches effectively narrow the semantic gap between a model's latent space and human understanding, they typically extract only high-level semantics from the model's final feature map. As a result, they provide a limited perspective on the decision-making process. We argue that explanations lacking insight into both lower- and mid-level semantics cannot be considered fully faithful or genuinely useful. To address this issue, we introduce the Multi-Level Concept Prototypes Classifier (MCPNet), which offers a more holistic interpretation by drawing on information from multiple levels within the model. Rather than relying on predefined concept labels, MCPNet autonomously discovers meaningful concepts from feature maps. To increase versatility, we further propose MCPNet++, which can be seamlessly applied to both CNN and transformer backbones, allowing it to learn meaningful concepts from their respective features. Building on these learned concepts, we also introduce a large language model (LLM)-based method to bridge the gap between these concepts and human perception. Experimental results show that MCPNet++ provides more comprehensive explanations without sacrificing model performance, with the discovered concepts aligning closely with human understanding.
Bor-Shiun Wang, Chien-Yi Wang, Walon Wei-Chen Chiu
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP
abstract
Large-scale vision-language models, such as CLIP, are known to contain societal bias regarding protected attributes (e.g., gender, age). This paper aims to address the problems of societal bias in CLIP. Although previous studies have proposed to debias societal bias through adversarial learning or test-time projecting, our comprehensive study of these works identifies two critical limitations: 1) loss of attribute information when it is explicitly disclosed in the input and 2) use of the attribute annotations during debiasing process. To mitigate societal bias in CLIP and overcome these limitations simultaneously, we introduce a simple-yet-effective debiasing method called SANER (societal attribute neutralizer) that eliminates attribute information from CLIP text features only of attribute-neutral descriptions. Experimental results show that SANER, which does not require attribute annotations and preserves original information for attribute-specific descriptions, demonstrates superior debiasing ability than the existing methods.
Yusuke Hirota, Min-Hung Chen, Chien-Yi Wang, Yuta Nakashima, Yu-Chiang Frank Wang, Ryo Hachiuma
ICLR3
2025 BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL
abstract
Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in the context of single-objective BO, learning-based AFs witnessed promising empirical results given its favorable non-myopic nature. Despite this, the direct extension of these approaches to multi-objective Bayesian optimization (MOBO) suffer from the hypervolume identifiability issue, which results from the non-Markovian nature of MOBO problems. To tackle this, inspired by the non-Markovian RL literature and the success of Transformers in language modeling, we present a generalized deep Q-learning framework and propose BOFormer, which substantiates this framework for MOBO via sequence modeling. Through extensive evaluation, we demonstrate that BOFormer constantly achieves better performance than the benchmark rule-based and learning-based algorithms in various synthetic MOBO and real-world multi-objective hyperparameter optimization problems.
Yu-Heng Hung, Kai-Jie Lin, Yu-Heng Lin, Chien-Yi Wang, Ping-Chun Hsieh
ICLR4
2025 Semantic Prompt Learning for Weakly-Supervised Semantic Segmentation
abstract
Weakly-Supervised Semantic Segmentation (WSSS) aims to train segmentation models using image data with only image-level supervision. Since precise pixel-level annotations are not accessible, existing methods typically focus on producing pseudo masks for training segmentation models by refining CAM-like heatmaps. However, the produced heatmaps may capture only the discriminative image regions of object categories or the associated co-occurring backgrounds. To address the issues, we propose a Semantic Prompt Learning for WSSS (SemPLeS) framework, which learns to effectively prompt the CLIP latent space to enhance the semantic alignment between the segmented regions and the target object categories. More specifically, we propose Contrastive Prompt Learning and Prompt-guided Semantic Refinement to learn the prompts that adequately describe and suppress the co-occurring backgrounds associated with each object category. In this way, SemPLeS can perform better semantic alignment between object regions and class labels, resulting in desired pseudo masks for training segmentation models. The proposed SemPLeS framework achieves competitive performance on standard WSSS benchmarks, PASCAL VOC 2012 and MS COCO 2014, and shows compatibility with other WSSS methods. Project page: https://projectdisr.github.io/semples/
Ci-Siang Lin, Chien-Yi Wang, Yu-Chiang Frank Wang, Min-Hung Chen
WACV2
2024 MCPNet: An Interpretable Classifier via Multi-Level Concept Prototypes
abstract
Recent advancements in post-hoc and inherently inter-pretable methods have markedly enhanced the explanations of black box classifier models. These methods operate either through post-analysis or by integrating concept learning during model training. Although being effective in bridging the semantic gap between a model's latent space and human interpretation, these explanation methods only partially reveal the model's decision-making process. The outcome is typically limited to high-level semantics derived from the last feature map. We argue that the expla-nations lacking insights into the decision processes at low and mid-level features are neither fully faithful nor useful. Addressing this gap, we introduce the Multi-Level Concept Prototypes Classifier (MCPNet), an inherently interpretable model. MCPNet autonomously learns meaningful concept prototypes across multiple feature map levels using Cen-tered Kernel Alignment (CKA) loss and an energy-based weighted PCA mechanism, and it does so without reliance on predefined concept labels. Further, we propose a novel classifier paradigm that learns and aligns multilevel concept prototype distributions for classification purposes via Class-aware Concept Distribution (CCD) loss. Our experiments reveal that our proposed MCPNet while being adapt-able to various model architectures, offers comprehensive multilevel explanations while maintaining classification accuracy. Additionally, its concept distribution-based classification approach shows improved generalization capabilities in few-shot classification scenarios. Project page is available11https://eddie221.github.io/MCPNet/.
Bor-Shiun Wang, Chien-Yi Wang, Walon Wei-Chen Chiu
CVPR2
2024 RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question Answering
abstract
Natural Language Explanation (NLE) in vision and language tasks aims to provide human-understandable explanations for the associated decision-making process. In practice, one might encounter explanations which lack informativeness or contradict visual-grounded facts, known as implausibility and hallucination problems, respectively. To tackle these challenging issues, we consider the task of visual question answering (VQA) and introduce Rapper, a two-stage Reinforced Rationale-Prompted Paradigm. By knowledge distillation, the former stage of Rapper infuses rationale-prompting via large language models (LLMs), encouraging the rationales supported by language-based facts. As for the latter stage, a unique Reinforcement Learning from NLE Feedback (RLNF) is introduced for injecting visual facts into NLE generation. Finally, quantitative and qualitative experiments on two VL-NLE benchmarks show that Rapper surpasses state-of-the-art VQA-NLE methods while providing plausible and faithful NLE.
Kai-Po Chang, Chi-Pin Huang, Wei-Yuan Cheng, Fu-En Yang, Chien-Yi Wang, Yung-Hsuan Lai, Yu-Chiang Frank Wang
ICLR5
2024 DoRA: Weight-Decomposed Low-Rank Adaptation
abstract
Among the widely used parameter-efficient fine-tuning (PEFT) methods, LoRA and its variants have gained considerable popularity because of avoiding additional inference costs. However, there still often exists an accuracy gap between these methods and full fine-tuning (FT). In this work, we first introduce a novel weight decomposition analysis to investigate the inherent differences between FT and LoRA. Aiming to resemble the learning capacity of FT from the findings, we propose Weight-Decomposed Low-Rank Adaptation (DoRA). DoRA decomposes the pre-trained weight into two components, magnitude and direction, for fine-tuning, specifically employing LoRA for directional updates to efficiently minimize the number of trainable parameters. By employing DoRA, we enhance both the learning capacity and training stability of LoRA while avoiding any additional inference overhead. DoRA consistently outperforms LoRA on fine-tuning LLaMA, LLaVA, and VL-BART on various downstream tasks, such as commonsense reasoning, visual instruction tuning, and image/video-text understanding. The code is available at https://github.com/NVlabs/DoRA.
Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov 0001, Yu-Chiang Frank Wang, Kwang-Ting Cheng, Min-Hung Chen
ICML2
2024 Probabilistic 3D Multi-Object Cooperative Tracking for Autonomous Driving via Differentiable Multi-Sensor Kalman Filter
abstract
Current state-of-the-art autonomous driving vehicles mainly rely on each individual sensor system to perform perception tasks. Such a framework’s reliability could be limited by occlusion or sensor failure. To address this issue, more recent research proposes using vehicle-to-vehicle (V2V) communication to share perception information with others. However, most relevant works focus only on cooperative detection and leave cooperative tracking an underexplored research field. A few recent datasets, such as V2V4Real, provide 3D multi-object cooperative tracking benchmarks. However, their proposed methods mainly use cooperative detection results as input to a standard single-sensor Kalman Filter-based tracking algorithm. In their approach, the measurement uncertainty of different sensors from different connected autonomous vehicles (CAVs) may not be properly estimated to utilize the theoretical optimality property of Kalman Filter-based tracking algorithms. In this paper, we propose a novel 3D multi-object cooperative tracking algorithm for autonomous driving via a differentiable multi-sensor Kalman Filter. Our algorithm learns to estimate measurement uncertainty for each detection that can better utilize the theoretical property of Kalman Filter-based tracking methods. The experiment results show that our algorithm improves the tracking accuracy by 17% with only 0.037x communication costs compared with the state-of-the-art method in V2V4Real. Our code and videos are available at the URL and the URL.
Hsu-Kuang Chiu, Chien-Yi Wang, Min-Hung Chen, Stephen F. Smith
ICRA2
2023 MixFairFace: Towards Ultimate Fairness via MixFair Adapter in Face Recognition
abstract
Although significant progress has been made in face recognition, demographic bias still exists in face recognition systems. For instance, it usually happens that the face recognition performance for a certain demographic group is lower than the others. In this paper, we propose MixFairFace framework to improve the fairness in face recognition models. First of all, we argue that the commonly used attribute-based fairness metric is not appropriate for face recognition. A face recognition system can only be considered fair while every person has a close performance. Hence, we propose a new evaluation protocol to fairly evaluate the fairness performance of different approaches. Different from previous approaches that require sensitive attribute labels such as race and gender for reducing the demographic bias, we aim at addressing the identity bias in face representation, i.e., the performance inconsistency between different identities, without the need for sensitive attribute labels. To this end, we propose MixFair Adapter to determine and reduce the identity bias of training samples. Our extensive experiments demonstrate that our MixFairFace approach achieves state-of-the-art fairness performance on all benchmark datasets.
Fu-En Wang, Chien-Yi Wang, Min Sun 0001, Shang-Hong Lai
AAAI2
2023 Generalized Face Anti-Spoofing via Multi-Task Learning and One-Side Meta Triplet Loss
abstract
With the increasing variations of face presentation attacks, model generalization becomes an essential challenge for a practical face anti-spoofing system. This paper presents a generalized face anti-spoofing framework that consists of three tasks: depth estimation, face parsing, and live/spoof classification. With the pixel-wise supervision from the face parsing and depth estimation tasks, the regularized features can better distinguish spoof faces. While simulating domain shift with meta-learning techniques, the proposed one-side triplet loss can further improve the generalization capability by a large margin. Extensive experiments on four public datasets demonstrate that the proposed framework and training strategies are more effective than previous works for model generalization to unseen domains.
Chu-Chun Chuang, Chien-Yi Wang, Shang-Hong Lai
FG2
2023 Efficient Model Personalization in Federated Learning via Client-Specific Prompt Generation
abstract
Federated learning (FL) emerges as a decentralized learning framework which trains models from multiple distributed clients without sharing their data to preserve privacy. Recently, large-scale pre-trained models (e.g., Vision Transformer) have shown a strong capability of deriving robust representations. However, the data heterogeneity among clients, the limited computation resources, and the communication bandwidth restrict the deployment of large-scale models in FL frameworks. To leverage robust representations from large-scale models while enabling efficient model personalization for heterogeneous clients, we propose a novel personalized FL framework of client-specific Prompt Generation (pFedPG), which learns to deploy a personalized prompt generator at the server for producing client-specific visual prompts that efficiently adapts frozen backbones to local data distributions. Our proposed framework jointly optimizes the stages of personalized prompt adaptation locally and personalized prompt generation globally. The former aims to train visual prompts that adapt foundation models to each client, while the latter observes local optimization directions to generate personalized prompts for all clients. Through extensive experiments on benchmark datasets, we show that our pFedPG is favorable against state-of-the-art personalized FL methods under various types of data heterogeneity, allowing computation and communication efficient model personalization.
Fu-En Yang, Chien-Yi Wang, Yu-Chiang Frank Wang
ICCV2
2022 FedFR: Joint Optimization Federated Framework for Generic and Personalized Face Recognition
abstract
Current state-of-the-art deep learning based face recognition (FR) models require a large number of face identities for central training. However, due to the growing privacy awareness, it is prohibited to access the face images on user devices to continually improve face recognition models. Federated Learning (FL) is a technique to address the privacy issue, which can collaboratively optimize the model without sharing the data between clients. In this work, we propose a FL based framework called FedFR to improve the generic face representation in a privacy-aware manner. Besides, the framework jointly optimizes personalized models for the corresponding clients via the proposed Decoupled Feature Customization module. The client-specific personalized model can serve the need of optimized face recognition experience for registered identities at the local device. To the best of our knowledge, we are the first to explore the personalized face recognition in FL setup. The proposed framework is validated to be superior to previous approaches on several generic and personalized face recognition benchmarks with diverse FL scenarios. The source codes and our proposed personalized FR benchmark under FL setup are available at https://github.com/jackie840129/FedFR.
Chih-Ting Liu, Chien-Yi Wang, Shao-Yi Chien, Shang-Hong Lai
AAAI2
2022 PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition
abstract
Face anti-spoofing (FAS) plays a critical role in securing face recognition systems from different presentation attacks. Previous works leverage auxiliary pixel-level supervision and domain generalization approaches to address unseen spoof types. However, the local characteristics of image captures, i.e., capturing devices and presenting materials, are ignored in existing works and we argue that such information is required for networks to discriminate between live and spoof images. In this work, we propose PatchNet which reformulates face anti-spoofing as a fine-grained patch-type recognition problem. To be specific, our framework recognizes the combination of capturing devices and presenting materials based on the patches cropped from non-distorted face images. This reformulation can largely improve the data variation and enforce the network to learn discriminative feature from local capture patterns. In addition, to further improve the generalization ability of the spoof feature, we propose the novel Asymmetric Margin-based Classification Loss and Self-supervised Similarity Loss to regularize the patch embedding space. Our experimental results verify our assumption and show that the model is capable of recognizing unseen spoof types robustly by only looking at local regions. Moreover, the fine-grained and patch-level reformulation of FAS outperforms the existing approaches on intra-dataset, cross-dataset, and domain generalization benchmarks. Furthermore, our PatchNet framework can enable practical applications like FewShot Reference-based FAS and facilitate future exploration of spoof-related intrinsic cues.
Chien-Yi Wang, Yu-Ding Lu, Shang-Ta Yang, Shang-Hong Lai
CVPR1
2022 Local-Adaptive Face Recognition via Graph-based Meta-Clustering and Regularized Adaptation
abstract
Due to the rising concern of data privacy, it's reasonable to assume the local client data can't be transferred to a centralized server, nor their associated identity label is provided. To support continuous learning and fill the last-mile quality gap, we introduce a new problem setup called “local-adaptive face recognition (LaFR)”. Leveraging the environment-specific local data after the deployment of the initial global model, LaFR aims at getting optimal performance by training local-adapted models automatically and un-supervisely, as opposed to fixing their initial global model. We achieve this by a newly proposed embedding cluster model based on Graph Convolution Network (GCN), which is trained via meta-optimization procedure. Compared with previous works, our meta-clustering model can generalize well in unseen local environments. With the pseudo identity labels from the clustering results, we further introduce novel regularization techniques to improve the model adaptation performance. Extensive experiments on racial and internal sensor adaptation demonstrate that our proposed solution is more effective for adapting face recognition models in each specific environment. Meanwhile, we show that LaFR can further improve the global model by a simple federated aggregation over the updated local models.
Chien-Yi Wang, Kuan-Lun Tseng, Shang-Hong Lai, Baoyuan Wang
CVPR2
2022 Disentangled Representation with Dual-stage Feature Learning for Face Anti-spoofing
abstract
As face recognition is widely used in diverse security-critical applications, the study of face anti-spoofing (FAS) has attracted more and more attention. Several FAS methods have achieved promising performance if the attack types in the testing data are included in the training data, while the performance significantly degrades for unseen attack types. It is essential to learn more generalized and discriminative features to prevent overfitting to pre-defined spoof attack types. This paper proposes a novel dual-stage disentangled representation learning method that can efficiently untangle spoof-related features from irrelevant ones. Un-like previous FAS disentanglement works with one-stage architecture, we found that the dual-stage training design can improve the training stability and effectively encode the features to detect unseen attack types. Our experiments show that the proposed method provides superior accuracy than the state-of-the-art methods on several cross-type FAS benchmarks.
Yu-Chun Wang, Chien-Yi Wang, Shang-Hong Lai
WACV2
2021 High-Accuracy RGB-D Face Recognition via Segmentation-Aware Face Depth Estimation and Mask-Guided Attention Network
abstract
Deep learning approaches have achieved highly accurate face recognition by training the models with very large face image datasets. Unlike the availability of large 2D face image datasets, there is a lack of large 3D face datasets available to the public. Existing public 3D face datasets were usually collected with few subjects, leading to the over-fitting problem. This paper proposes two CNN models to improve the RGB-D face recognition task. The first is a segmentation-aware depth estimation network, called DepthNet, which estimates depth maps from RGB face images by including semantic segmentation information for more accurate face region localization. The other is a novel mask-guided RGB-D face recognition model that contains an RGB recognition branch, a depth map recognition branch, and an auxiliary segmentation mask branch with a spatial attention module. Our DepthN et is used to augment a large 2D face image dataset to a large RGB-D face dataset, which is used for training an accurate RGB-D face recognition model. Furthermore, the proposed mask-guided RGB-D face recognition model can fully exploit the depth map and segmentation mask information and is more robust against pose variation than previous methods. Our experimental results show that DepthNet can produce more reliable depth maps from face images with the segmentation mask. Our mask-guided face recognition model outperforms state-of-the-art methods on several public 3D face datasets.
Meng-Tzu Chiu, Hsun-Ying Cheng, Chien-Yi Wang, Shang-Hong Lai
FG3
2020 Unified Representation Learning for Cross Model Compatibility
Chien-Yi Wang, Ya-Liang Chang, Shang-Ta Yang, Shang-Hong Lai
BMVC1
2020 Construction of Obstacle-Avoiding Delay-Driven GNR Routing Tree
abstract
It is known that graphene nanoribbon (GNR) based devices and interconnects can be treated to be better alternative in nano-scale designs. In this paper, given a source pin and a set of target pins inside a GNR routing plane with a set of rectangular obstacles, based on the selection of the possible obstacle-avoiding delay-driven routing paths on the target pins, an efficient routing algorithm can be proposed to construct an obstacle-avoiding delay-driven GNR routing tree with minimizing the total wirelength for the target pins. Compared with Das's algorithm with obstacle avoidance on total wirelength and maximum source-to-target delay, the experimental results show that our proposed routing algorithm only uses 7.87% of the extra wirelength to reduce 23.42% of the maximum delay in the construction of an obstacle-avoiding delay-driven GNR routing tree for 6 tested examples on the average.
Jin-Tai Yan, Po-Yuan Huang, Chien-Yi Wang
TENCON3
2018 A 3D Dynamic Scene Analysis Framework for Development of Intelligent Transportation Systems
abstract
Holistic driving scene understanding is a critical step toward intelligent transportation systems. It involves different levels of analysis, interpretation, reasoning and decision making. In this paper, we propose a 3D dynamic scene analysis framework as the first step toward driving scene understanding. Specifically, given a sequence of synchronized 2D and 3D sensory data, the framework systematically integrates different perception modules to obtain 3D position, orientation, velocity and category of traffic participants and the ego car in a reconstructed 3D semantically labeled traffic scene. We implement this framework and demonstrate the effectiveness in challenging urban driving scenarios. The proposed framework builds a foundation for higher level driving scene understanding problems such as intention and motion prediction of surrounding entities, ego motion planning, and decision making.
Chien-Yi Wang, Athma Narayanan, Abhishek Patil, Yi-Ting Chen 0001
Intelligent Vehicles Symposium1
2018 Improved Converses and Gap Results for Coded Caching
abstract
Improved lower bounds are derived on the average and worst case rate-memory tradeoffs of the Maddah-Ali and Niesen-coded caching scenario. For any number of users and files and for arbitrary cache sizes, the multiplicative gap between the exact rate-memory tradeoff and the new lower bound is shown to be less than 2.315 in the worst case scenario and 2.507 in the average-case scenario.
Chien-Yi Wang, Shirin Saeedi Bidokhti, Michèle Wigger
IEEE Trans. Inf. Theory1
2018 On Achievability for Downlink Cloud Radio Access Networks With Base Station Cooperation
abstract
This paper investigates the downlink of a cloud radio access network (C-RAN) in which a central processor communicates with two mobile users through two base stations (BSs). The BSs act as relay nodes and cooperate with each other through error-free rate-limited links. We develop and analyze two coding schemes for this scenario. The first coding scheme modifies the Liu-Kang scheme (to make it amenable to a rigorous analysis) and extends it to introduce common codewords and to apply for downlink C-RAN with BS-to-BS cooperation. This first coding scheme enables arbitrary correlation among the auxiliary codewords that are recovered by the BSs. We show that this scheme improves over previous schemes for various instances of Gaussian C-RAN channels. In particular, in many scenarios, the scheme can better exploit the possibility of BS-to-BS cooperation than other schemes. The second coding scheme extends the distributed decode and forward (DDF) scheme by means of Gray-Wyner compression and by exploiting the cooperation links between BSs. In addition and as a separate extension, we provide an improved capacity approximation for the DDF strategy for the capacity of a general N-BS L-user C-RAN model in the memoryless Gaussian case.
Chien-Yi Wang, Michèle Wigger, Abdellatif Zaidi
IEEE Trans. Inf. Theory1
2017 Improved converses and gap-results for coded caching
abstract
Improved lower bounds on the worst-case and the average-case rate-memory tradeoffs for the Maddah-Ali&Niesen coded-caching scenario are presented. For any number of users and files and for arbitrary cache sizes, the multiplicative gap between the exact rate-memory tradeoff and the new lower bound is less than 2.315 in the worst-case scenario and less than 2.507 in the average-case scenario.
Chien-Yi Wang, Shirin Saeedi Bidokhti, Michèle Wigger
ISIT1
2017 Rate-distortion regions of instances of cascade source coding with side information
abstract
In this work, we study a three-terminal cascade source coding problem with side information Yi known to the source encoder and the first user, and side information Y2 known only to the second user. Each user wants to reconstruct some desired function of the source, lossily, to within some fidelity level. We establish single-letter characterization of the rate-distortion region of this model in some important special cases, including when the reconstruction is lossless at the first user. We then establish a connection among the studied model and the so-called side information-scalable source coding problem (i.e., Heegard-Berger problem with side information and successive refinement) to infer single-letter characterization of the rate-distortion region of some instances of the latter problem. In contrast with most previous related works, the results of this paper hold irrespective of the ordering among the source and side information sequences, which are then arbitrarily correlated.
Chien-Yi Wang, Abdellatif Zaidi
ISIT1
2017 On Achievability for Downlink Cloud Radio Access Networks with Base Station Cooperation
abstract
This work investigates the downlink of cloud radio access networks (C-RANs), assuming digital cooperation links among the base stations (BSs). A generalization of the data-sharing scheme is proposed for the case of two BSs and two mobile users. The generalized data-sharing scheme includes a common part and allows full exploitation of correlation among auxiliary codewords. The cooperation links between the BSs are used to exchange and to redirect indices precomputed at the central processor. On the other hand, by simplifying the achievable rate region of the distributed decode-forward (DDF) scheme, it is shown that the DDF scheme for broadcast achieves the capacity region of a downlink $N$-BS $L$-user C- RAN with BS cooperation under the memoryless Gaussian model to within a gap of $\frac{L}{2}#x002B;\frac{\min\ {N,L\log_2 N\}}{2}$ bits per dimension. Numerical evaluations for the memoryless Gaussian model indicate that the generalized data-sharing scheme 1) outperforms the DDF scheme in the low-power regime and when the channel gain matrix is ill-conditioned and 2) benefits more from BS cooperation.
Chien-Yi Wang, Michèle Wigger, Abdellatif Zaidi
WCNC1
2017 Information-Theoretic Caching: The Multi-User Case
abstract
In this paper, we consider a cache aided network in which each user is assumed to have individual caches, while upon users' requests, an update message is sent through a common link to all users. First, we formulate a general information theoretic setting that represents the database as a discrete memoryless source, and the users' requests as side information that is available everywhere except at the cache encoder. The decoders' objective is to recover a function of the source and the side information. By viewing cache aided networks in terms of a general distributed source coding problem and through information theoretic arguments, we present inner and outer bounds on the fundamental tradeoff of cache memory size and update rate. Then, we specialize our general inner and outer bounds to a specific model of content delivery networks: file selection networks, in which the database is a collection of independent equal-size files and each user requests one of the files independently. For file selection networks, we provide an outer bound and two inner bounds (for centralized and decentralized caching strategies). For the case when the user request information is uniformly distributed, we characterize the rate versus cache size tradeoff to within a multiplicative gap of 4. By further extending our arguments to the framework of Maddah-Ali and Niesen, we also establish a new outer bound and two new inner bounds in which it is shown to recover the centralized and decentralized strategies, previously established by Maddah-Ali and Niesen. Finally, in terms of rate versus cache size tradeoff, we improve the previous multiplicative gap of 72 to 4.7 for the average case with uniform requests.
Sung Hoon Lim, Chien-Yi Wang, Michael Gastpar
IEEE Trans. Inf. Theory2
2016 Information theoretic caching: The multi-user case
abstract
In this paper, we present information theoretic inner and outer bounds on the fundamental tradeoff between cache memory size and update rate in a multi-user cache network. Each user is assumed to have individual caches, while upon users' requests, an update message is sent though a common link to all users. The database is represented as a discrete memoryless source and the user request information is represented as side information that is available at the decoders and the update encoder, but oblivious to the cache encoder. We establish two inner bounds, the first based on a centralized caching strategy and the second based on a decentralized caching strategy. For the case when the user requests are i.i.d. with the uniform distribution, we show that the performance of the decentralized inner bound is within a multiplicative gap of 4 from the optimal cache-rate tradeoff. For general request distributions, we numerically compare the bounds and the baseline uncoded strategy, caching the most popular files.
Sung Hoon Lim, Chien-Yi Wang, Michael Gastpar
ISIT2
2016 Information-Theoretic Caching: Sequential Coding for Computing
abstract
Under the paradigm of caching, partial data are delivered before the actual requests of users are known. In this paper, this problem is modeled as a canonical distributed source coding problem with side information, where the side information represents the users' requests. For the single-user case, a singleletter characterization of the optimal rate region is established, and for several important special cases, closed-form solutions are given, including the scenario of uniformly distributed user requests. In this case, it is shown that the optimal caching strategy is closely related to total correlation and Wyner's common information. Using the insight gained from the single-user case, three two-user scenarios admitting single-letter characterization are considered, which draw connections to existing source coding problems in the literature: the Gray-Wyner system and distributed successive refinement. Finally, the model studied by Maddah-Ali and Niesen is rephrased to make a comparison with the considered information-theoretic model. Although the two caching models have a similar behavior for the single-user case, it is shown through a two-user example that the two caching models behave differently in general.
Chien-Yi Wang, Sung Hoon Lim, Michael Gastpar
IEEE Trans. Inf. Theory1
2015 Robust Image Segmentation Using Contour-Guided Color Palettes
abstract
The contour-guided color palette (CCP) is proposed for robust image segmentation. It efficiently integrates contour and color cues of an image. To find representative colors of an image, color samples along long contours between regions, similar in spirit to machine learning methodology that focus on samples near decision boundaries, are collected followed by the mean-shift (MS) algorithm in the sampled color space to achieve an image-dependent color palette. This color palette provides a preliminary segmentation in the spatial domain, which is further fine-tuned by post-processing techniques such as leakage avoidance, fake boundary removal, and small region mergence. Segmentation performances of CCP and MS are compared and analyzed. While CCP offers an acceptable standalone segmentation result, it can be further integrated into the framework of layered spectral segmentation to produce a more robust segmentation. The superior performance of CCP-based segmentation algorithm is demonstrated by experiments on the Berkeley Segmentation Dataset.
Xiang Fu 0006, Chien-Yi Wang, Chen Chen 0021, Changhu Wang, C.-C. Jay Kuo
ICCV2
2015 Information-theoretic caching
abstract
Motivated by the caching problem introduced by Maddah-Ali and Niesen, a problem of distributed source coding with side information is formulated, which captures a distinct interesting aspect of caching. For the single-user case, a single-letter characterization of the optimal rate region is presented. For the cases where the source is composed of either independent or nested components, the exact optimal rate regions are found and some intuitive caching strategies are confirmed to be optimal. When the components are arbitrarily correlated with uniform requests, the optimal caching strategy is found to be closely related to total correlation and Wyner's common information. For the two-user case, some subproblems are solved which draw connections to the Gray-Wyner system and distributed successive refinement. Finally, inner and outer bounds are given for the case of two private caches with a common update.
Chien-Yi Wang, Sung Hoon Lim, Michael Gastpar
ISIT1
2015 Interactive Computation of Type-Threshold Functions in Collocated Gaussian Networks
abstract
In wireless sensor networks, various applications involve learning one or multiple functions of the measurements observed by sensors, rather than the measurements themselves. This paper focuses on the class of type-threshold functions, e.g., the maximum and the indicator functions. A simple network model capturing both the broadcast and superposition properties of wireless channels is considered: the collocated Gaussian network. A general multiround coding scheme exploiting superposition and interaction (through broadcast) is developed. Through careful scheduling of concurrent transmissions to reduce redundancy, it is shown that given any independent measurement distribution, all type-threshold functions can be computed reliably with a nonvanishing rate in the collocated Gaussian network, even if the number of sensors tends to infinity.
Chien-Yi Wang, Sang-Woon Jeon, Michael Gastpar
IEEE Trans. Inf. Theory1
2014 On distributed successive refinement with lossless recovery
abstract
The problem of successive refinement in distributed source coding and in joint source-channel coding is considered. The emphasis is placed on the case where the sources have to be recovered losslessly in the second stage. In distributed source coding, it is shown that all sources are successively refinable in sum rate, with respect to any (joint) distortion measure in the first stage. In joint source-channel coding, the sources are assumed independent and only a (per letter) function is to be recovered losslessly in the first stage. For a class of multiple access channels, it is shown that all sources are successively refinable with respect to a class of linear functions. Finally, when the sources have equal entropy, a simple sufficient condition of successive refinability is provided for partially invertible functions.
Chien-Yi Wang, Michael Gastpar
ISIT1
2014 Approximate Ergodic Capacity of a Class of Fading Two-User Two-Hop Networks
abstract
The fading AWGN two-user two-hop network is considered where the channel coefficients are independent and identically distributed (i.i.d.) according to a continuous distribution and vary over time. For a broad class of channel distributions, the ergodic sum capacity is characterized to within a constant number of bits/second/hertz, independent of the signal-to-noise ratio. The achievability follows from the analysis of an interference neutralization scheme where the relays are partitioned into M pairs, and interference is neutralized separately by each pair of relays. When M = 1, the proposed ergodic interference neutralization characterizes the ergodic sum capacity to within 4 bits/sec/Hz for i.i.d. uniform phase fading and approximately 4.7 bits/sec/Hz for i.i.d. Rayleigh fading. It is further shown that this gap can be tightened to 4 log π-4 bits/sec/Hz (approximately 2.6) for i.i.d. uniform phase fading and 4-4 log(3π/8) bits/sec/Hz (approximately 3.1) for i.i.d. Rayleigh fading in the limit of large M1.
Sang-Woon Jeon, Chien-Yi Wang, Michael Gastpar
IEEE Trans. Inf. Theory2
2014 Computation Over Gaussian Networks With Orthogonal Components
abstract
Function computation over Gaussian networks with orthogonal components is studied for arbitrarily correlated discrete memoryless sources. Two classes of functions are considered: 1) the arithmetic sum function and 2) the type function. The arithmetic sum function in this paper is defined as a set of multiple weighted arithmetic sums, which includes averaging of the sources and estimating each of the sources as special cases. The type or frequency histogram function counts the number of occurrences of each argument, which yields various fundamental statistics, such as mean, variance, maximum, minimum, median, and so on. The proposed computation coding first abstracts Gaussian networks into the corresponding modulo sum multiple-access channels via nested lattice codes and linear network coding and then computes the desired function using linear Slepian-Wolf source coding. For orthogonal Gaussian networks (with no broadcast and multiple-access components), the computation capacity is characterized for a class of networks. For Gaussian networks with multiple-access components (but no broadcast), an approximate computation capacity is characterized for a class of networks.
Sang-Woon Jeon, Chien-Yi Wang, Michael Gastpar
IEEE Trans. Inf. Theory2
2013 Computation over Gaussian networks with orthogonal components
abstract
Function computation of arbitrarily correlated discrete sources over Gaussian networks with multiple access components but no broadcast is studied. Two classes of functions are considered: the arithmetic sum function and the frequency histogram function. The arithmetic sum function in this paper is defined as a set of multiple weighted arithmetic sums, which includes averaging of sources and estimating each of the sources as special cases. The frequency histogram function counts the number of occurrences of each argument, which yields many important statistics such as mean, variance, maximum, minimum, median, and so on. For a class of networks, an approximate computation capacity is characterized. The proposed approach first abstracts Gaussian networks into the corresponding modulo-sum multiple-access channels via lattice codes and linear network coding and then computes the desired function by using linear Slepian-Wolf source coding.
Sang-Woon Jeon, Chien-Yi Wang, Michael Gastpar
ISIT2
2013 Multi-round computation of type-threshold functions in collocated Gaussian networks
abstract
In wireless sensor networks, various applications involve learning one or multiple functions of the measurements observed by sensors, rather than the measurements themselves. This paper focuses on the computation of type-threshold functions which include the maximum, minimum, and indicator functions as special cases. Previous work studied this problem under the collocated collision network model and showed that under many probabilistic models for the measurements, the achievable computation rates tend to zero as the number of sensors increases. In this paper, wireless sensor networks are modeled as fully connected Gaussian networks with equal channel gains, which are termed collocated Gaussian networks. A general multi-round coding scheme exploiting not only the broadcast property but also the superposition property of Gaussian networks is developed. Through careful scheduling of concurrent transmissions to reduce redundancy, it is shown that given any independent measurement distribution, all type-threshold functions can be computed reliably with a non-vanishing rate even if the number of sensors tends to infinity.
Chien-Yi Wang, Sang-Woon Jeon, Michael Gastpar
ISIT1
2012 Approximate ergodic capacity of a class of fading 2-user 2-hop networks
abstract
We consider a fading AWGN 2-user 2-hop network in which the channel coefficients are independently and identically distributed (i.i.d.) drawn from a continuous distribution and vary over time. For a broad class of channel distributions, we characterize the ergodic sum capacity within a constant number of bits/sec/Hz, independent of signal-to-noise ratio. The achievability follows from the analysis of an interference neutralization scheme where the relays are partitioned into K pairs, and interference is neutralized separately by each pair of relays. For K = 1, we previously proved a gap of 4 bits/sec/Hz for i.i.d. uniform phase fading and approximately 4.7 bits/sec/Hz for i.i.d. Rayleigh fading. In this paper, we give a result for general K. In the limit of large K, we characterize the ergodic sum capacity within 4((log π) - 1) ≃ 2.6 bits/sec/Hz for i.i.d. uniform phase fading and 4(4 - log3π) ≃ 3.1 bits/sec/Hz for i.i.d. Rayleigh fading.
Sang-Woon Jeon, Chien-Yi Wang, Michael Gastpar
ISIT2
2012 Asymptotic Coded BER Analysis for MIMO BICM-ID with Quantized Extrinsic LLR
abstract
In this paper, we derive a closed-form expression for the probability density/mass function (PDF/PMF) and the moment generating function (MGF) of the quantized detector soft output, i.e., extrinsic log-likelihood ratio (LLR), for multiple-input multiple-output (MIMO) bit-interleaved coded modulation with iterative decoding (BICM-ID) systems. The effect of either LLR clipping or LLR clipping with rounding, often applied in practical implementations, are considered. Using the derived expression, we analyze the asymptotic coded bit error rate (BER) for MIMO BICM-ID systems with the two quantization operations under a flat Rayleigh fading channel. The error rate degradation caused by quantizing the extrinsic LLR is then interpreted as an additional signal-to-noise ratio (SNR) loss. Rather than Monte Carlo simulations, this theoretical treatment provides a more convenient alternative to determining the clipping level and optimal word-length (the number of bits needed to represent signals) for LLR in a BICM-ID implementation. Finally, several other applications of this proposed theoretical analysis are also demonstrated.
I-Wei Lai, Chien-Yi Wang, Tzi-Dar Chiueh, Gerd Ascheid, Heinrich Meyr
IEEE Trans. Commun.2
2010 BER analysis for MIMO BICM-ID assuming finite precision of extrinsic LLR
abstract
In this paper, we analyze the error rate degradation caused by the finite-precision extrinsic log-likelihood ratio (LLR), i.e., demapper output, for multiple-input multiple-output (MIMO) bit-interleaved coded modulation with iterative decoding (BICM-ID). A closed-form expression for the probability density function (pdf) of the metric difference with finite precision is derived. This pdf is well-approximated by two approaches, namely the pseudo quantization noise (PQN) model and the pdf tail truncation. Moreover, we interpret such performance degradation as an additional signal-to-noise ratio (SNR) loss. As validated by Monte Carlo simulations, this SNR loss can be directly applied at the first iteration, extending our analysis to non-iterative scenario. This work provides a convenient indicator for deciding the LLR precision in real implementations.
Chien-Yi Wang, I-Wei Lai, Tzi-Dar Chiueh, Gerd Ascheid, Heinrich Meyr
ISITA1
2006 A Token-based Distributed Scheduling for Mesh Networks with Chain Topologies
abstract
The IEEE 802.11 DCF protocol used in multihop networks with chain topologies does not provide fair throughput to each node due to the hidden/exposed terminal problem and edge node effect. Eliminating the contention from hidden nodes is the key to providing high performance in the mesh network. We propose a token-based distributed scheduling scheme that not only guarantees throughput fairness among nodes but also allows allocation of proportional throughput among nodes.
Ting-Chao Hou, Chien-Yi Wang, Ming-Chieh Chan
AINA (1)2