Yizhuo Chen

dblp:315/8760 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
20since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Computer networks · 7 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 DiffPhys: Differential Physics Augmentations for Enhanced Representations
abstract
Foundation Models (FMs) have revolutionized representation learning in IoT sensing applications. However, these models face a critical limitation: while their generalized representations excel at detection despite environmental distortions, they struggle to differentiate between fine-grained variations in these distortions—a capability essential for many IoT tasks like proximity assessment and dynamic tracking. Traditional augmentation approaches exacerbate this problem by focusing on invariance to these distortions, teaching models to ignore rather than distinguish meaningful environmental variations. To address these limitations, we introduce DiffPhys, a novel framework that fundamentally shifts how models learn from augmentations. DiffPhys incorporates two key innovations: (i) progressive physics-guided augmentations modeling environmental effects at varying intensities, and (ii) an ordinal consistency constraint structuring the embedding space to preserve physical relationships. To support this framework, we implement differentiable physics-guided augmentations that model progressive environmental effects like attenuation, Doppler shifts, and scattering. DiffPhys is designed as a pluggable module that seamlessly integrates into existing IoT-driven ML pipelines without requiring architectural changes to the underlying models. Evaluations across vehicle classification, speed estimation, and distance tracking tasks show that DiffPhys can enhance models to achieve up to 9% improvement in environmental differentiation tasks compared to standard versions.
Denizhan Kara, Tomoyoshi Kimura, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Hongjue Zhao, Joydeep Bhattacharyya, Tarek F. Abdelzaher
ICCCN5
2025 PASS: Private Attributes Protection with Stochastic Data Substitution
abstract
The growing Machine Learning (ML) services require extensive collections of user data, which may inadvertently include people’s private information irrelevant to the services. Various studies have been proposed to protect private attributes by removing them from the data while maintaining the utilities of the data for downstream tasks. Nevertheless, as we theoretically and empirically show in the paper, these methods reveal severe vulnerability because of a common weakness rooted in their adversarial training based strategies. To overcome this limitation, we propose a novel approach, PASS, designed to stochastically substitute the original sample with another one according to certain probabilities, which is trained with a novel loss function soundly derived from information-theoretic objective defined for utility-preserving private attributes protection. The comprehensive evaluation of PASS on various datasets of different modalities, including facial images, human activity sensory signals, and voice recording datasets, substantiates PASS’s effectiveness and generalizability.
Yizhuo Chen, Chun-Fu Chen 0001, Hsiang Hsu, Shaohan Hu, Tarek F. Abdelzaher
ICML1
2025 DynaGen: Conditional Diffusion Models for Enhancing Acoustic and Seismic-Based Vehicle Detection
Tianshi Wang 0002, Jinyang Li 0004, Qikai Yang, Ruijie Wang 0004, Yizhuo Chen, Dachun Sun, Yigong Hu, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher
INFOCOM5
2025 On Network-Efficient Multimodal Multi-Vantage Foundation Models for Distributed Sensing
abstract
The rise of multi-modal, multi-node foundation models has revolutionized intelligent IoT sensing systems by enabling general-purpose inference from distributed sensing sources to support diverse downstream applications. However, the high communication cost of transmitting raw sensor data from distributed nodes to a central inference model remains a critical bottleneck, particularly in bandwidth- or energy-constrained environments. While existing compression methods can reduce data volume, they often lack the adaptability needed to handle variations in data relevance and redundancy across sources, modalities, and time. To address this challenge, we introduce ZipFM, a lightweight, plug-and-play middleware that dynamically configures sensor data compression strategies on a per-node, per-modality, and per-time-step basis to minimize network traffic while preventing model degradation, taking model sensitivity to different data sources into account. ZipFM is (i) compatible with different pre-trained foundation models without requiring access to their internal mechanisms or retraining, (ii) agnostic to the underlying tools available for data compression, and (iii) independent of the specific downstream inference tasks performed. At its core, ZipFM uses the compression-induced latent representation shift, produced by the foundation model's backbone, as a proxy for downstream accuracy degradation, and enforces a system-wide optimal representation shift (in the sense of minimizing compression-related degradation) through a lightweight feedback control mechanism. Experiments on three real-world IoT sensing datasets demonstrate that ZipFM significantly reduces communication costs while preserving model performance.
Yizhuo Chen, Hongjue Zhao, You Lyu, Jinyang Li 0004, Tomoyoshi Kimura, Yigong Hu, Denizhan Kara, Maggie B. Wigness, Jeffrey N. Twigg, Tarek F. Abdelzaher
MASS2
2025 AdaTS: Learning Adaptive Time Series Representations via Dynamic Soft Contrasts
abstract
Learning robust representations from unlabeled time series is crucial, and contrastive learning offers a promising avenue. However, existing contrastive learning approaches for time series often struggle with defining meaningful similarities, tending to overlook inherent physical correlations and diverse, sequence-varying non-stationarity. This limits their representational quality and real-world adaptability. To address these limitations, we introduce AdaTS, a novel adaptive soft contrastive learning strategy. AdaTS offers a compute-efficient solution centered on dynamic instance-wise and temporal assignments to enhance time series representations, specifically by: (i) leveraging Time-Frequency Coherence for robust physics-guided similarity measurement; (ii) preserving relative instance similarities through ordinal consistency learning; and (iii) dynamically adapting to sequence-specific non-stationarity with dynamic temporal assignments. AdaTS is designed as a pluggable module to standard contrastive frameworks, achieving up to 13.7% accuracy improvements across diverse time series datasets and three state-of-the-art contrastive frameworks while enhancing robustness against label scarcity. The code will be publicly available upon acceptance.
Denizhan Kara, Tomoyoshi Kimura, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Hongjue Zhao, Shengzhong Liu, Tarek F. Abdelzaher
NeurIPS5
2025 Results of the Big ANN: NeurIPS'23 competition
abstract
The 2023 Big ANN Challenge, held at NeurIPS 2023, focused on advancing the state-of-the-art in indexing data structures and search algorithms for practical variants of Approximate Nearest Neighbor (ANN) search that reflect its the growing complexity and diversity of workloads. Unlike prior challenges that emphasized scaling up classical ANN search (Simhadri et al., NeurIPS 2021), this competition addressed sparse, filtered, out-of-distribution, and streaming variants of ANNS. Participants developed and submitted innovative solutions that were evaluated on new standard datasets with constrained computational resources. The results showcased significant improvements in search accuracy and efficiency, with notable contributions from both academic and industrial teams. This paper summarizes the competition tracks, datasets, evaluation metrics, and the innovative approaches of the top-performing submissions, providing insights into the current advancements and future directions in the field of approximate nearest neighbor search.
Harsha Vardhan Simhadri, Martin Aumüller 0001, Matthijs Douze, Dmitry Baranchuk, Amir Ingber, Edo Liberty, Benjamin Landrum, Magdalen Dobson, Mazin Karjikar, Laxman Dhulipala, Yuzheng Cai, Jiayang Shi, Weiguo Zheng, Yizhuo Chen, Ben Huang
NeurIPS19
2025 SCRAG: Social Computing-Based Retrieval Augmented Generation for Community Response Forecasting in Social Media Environments
abstract
This paper introduces SCRAG, a prediction frame-work inspired by social computing, designed to forecast community responses to real or hypothetical social media posts. SCRAG can be used by public relations specialists (e.g., to craft messaging in ways that avoid unintended misinterpretations) or public figures and influencers (e.g., to anticipate social responses), among other applications related to public sentiment prediction, crisis management, and social what-if analysis. While large language models (LLMs) have achieved remarkable success in generating coherent and contextually rich text, their reliance on static training data and susceptibility to hallucinations limit their effectiveness at response forecasting in dynamic social media environments. SCRAG overcomes these challenges by integrating LLMs with a Retrieval-Augmented Generation (RAG) technique rooted in social computing. Specifically, our framework retrieves (i) historical responses from the target community to capture their ideological, semantic, and emotional makeup, and (ii) external knowledge from sources such as news articles to inject time-sensitive context. This information is then jointly used to forecast the responses of the target community to new posts or narratives. Extensive experiments across six scenarios on the X platform (formerly Twitter), tested with various embedding models and LLMs, demonstrate over 10% improvements on average in key evaluation metrics. A concrete example further shows its effectiveness in capturing diverse ideologies and nuances. Our work provides a social computing tool for applications where accurate and concrete insights into community responses are crucial.
Dachun Sun, You Lyu, Jinning Li 0001, Yizhuo Chen, Tianshi Wang 0002, Tomoyoshi Kimura, Tarek F. Abdelzaher
SMARTCOMP4
2025 InfoMAE: Pair-Efficient Cross-Modal Alignment for Multimodal Time-Series Sensing Signals
abstract
Standard multimodal self-supervised learning (SSL) algorithms regard cross-modal synchronization as implicit supervisory labels during pretraining, thus posing high requirements on the scale and quality of multimodal samples. These constraints significantly limit the performance of sensing intelligence in IoT applications, as the heterogeneity and the non-interpretability of time-series signals result in abundant unimodal data but scarce high-quality multimodal pairs. This paper proposes InfoMAE, a cross-modal alignment framework that tackles the challenge of multimodal pair efficiency under the SSL setting by facilitating efficient cross-modal alignment of pretrained unimodal representations. InfoMAE achieves efficient cross-modal alignment with limited data pairs through a novel information theory-inspired formulation that simultaneously addresses distribution-level and instance-level alignment. Extensive experiments on two real-world IoT applications are performed to evaluate InfoMAE's pairing efficiency to bridge pretrained unimodal models into a cohesive joint multimodal model. InfoMAE enhances downstream multimodal tasks by over 60% with significantly improved multimodal pairing efficiency. It also improves unimodal task accuracy by an average of 22%.
Tomoyoshi Kimura, Osama A. Hanna, Yatong Chen 0001, Yizhuo Chen, Denizhan Kara, Tianshi Wang 0002, Jinyang Li 0004, Xiaomin Ouyang, Shengzhong Liu, Mani Srivastava 0001, Suhas N. Diggavi, Tarek F. Abdelzaher
WWW5
2024 Acies-OS: A Content-Centric Platform for Edge AI Twinning and Orchestration
abstract
This paper describes Acies-OS, a content-centric platform for edge AI twinning and orchestration that allows easy deployment, re-configuration, and control of edge AI services, augmented by a digital twin. The work is motivated by the proliferation of edge AI in a plethora of IoT applications, ranging from home automation to military defense, and the emergence of digital twins that go beyond monitoring and emulation into configuration management and optimization of edge capabilities. While past work focused on either the edge capabilities themselves or the digital twin, this work focuses on their seamless interactions, offering abstractions that enable the digital twin to manage and optimize an increasingly diverse edge AI system. Acies-OS features a structured namespace, a thin client library with flexible pub/sub-based communication, health monitoring support, and a control plane for twin-based value-added analysis and optimization. To illustrate the use of Acies-OS, we implemented a multi-node multi-modality vehicle classification application and used Acies-OS to interface it to a digital twin. We then deployed the system in the field to showcase run-time twin-based optimizations of inference latency, classification accuracy, and robustness to failures in noisy and challenging conditions.
Jinyang Li 0004, Yizhuo Chen, Tomoyoshi Kimura, Tianshi Wang 0002, Ruijie Wang 0004, Denizhan Kara, Yigong Hu, Walid A. Hanafy, Abel Souza, Prashant J. Shenoy, Maggie B. Wigness, Joydeep Bhattacharyya, Jae Kim, Guijun Wang, Greg Kimberly, Josh D. Eckhardt, Denis Osipychev, Tarek F. Abdelzaher
ICCCN2
2024 Data Augmentation for Human Activity Recognition via Condition Space Interpolation within a Generative Model
abstract
This paper presents a generative data augmentation approach for human activity recognition (HAR) to close the distribution gap between laboratory training and real-world deployment. Despite the recent success of deep learning methods in wearable sensor-based HAR tasks, performance degradation occurs during real-world deployment due to training data scarcity and the vast variability in human activities. In light of this, we aim to enhance the diversity of training datasets by generating new data points within the vicinity of existing samples, as informed by domain expertise. Unlike the commonly utilized methods that augment data by interpolating in data space or feature space, we innovate by applying interpolation in the condition space of a conditional generative model to augment HAR datasets. We use domain-specific knowledge to extract statistical metrics from sensor data, which serve as conditions to direct the generation process. We demonstrate how a conditional generative diffusion model, steered by interpolated conditions, can synthesize realistic new data with various high-level features that benefit the robustness of the downstream HAR models. Our methodology advances the use of interpolation in data augmentation by exploring the capability of a state-of-the-art generative model, offering novel perspectives for bolstering the robustness and generalizability of HAR systems. Experimental results demonstrate that condition space interpolation outperforms the conventional interpolation-based and generative model-based augmentation methods across various datasets and downstream classifier combinations.
Tianshi Wang 0002, Yizhuo Chen, Qikai Yang, Dachun Sun, Ruijie Wang 0004, Jinyang Li 0004, Tomoyoshi Kimura, Tarek F. Abdelzaher
ICCCN2
2024 MaSS: Multi-attribute Selective Suppression for Utility-preserving Data Transformation from an Information-theoretic Perspective
abstract
The growing richness of large-scale datasets has been crucial in driving the rapid advancement and wide adoption of machine learning technologies. The massive collection and usage of data, however, pose an increasing risk for people’s private and sensitive information due to either inadvertent mishandling or malicious exploitation. Besides legislative solutions, many technical approaches have been proposed towards data privacy protection. However, they bear various limitations such as leading to degraded data availability and utility, or relying on heuristics and lacking solid theoretical bases. To overcome these limitations, we propose a formal information-theoretic definition for this utility-preserving privacy protection problem, and design a data-driven learnable data transformation framework that is capable of selectively suppressing sensitive attributes from target datasets while preserving the other useful attributes, regardless of whether or not they are known in advance or explicitly annotated for preservation. We provide rigorous theoretical analyses on the operational bounds for our framework, and carry out comprehensive experimental evaluations using datasets of a variety of modalities, including facial images, voice audio clips, and human activity motion sensor signals. Results demonstrate the effectiveness and generalizability of our method under various configurations on a multitude of tasks. Our source code is available at this URL.
Yizhuo Chen, Chun-Fu Chen 0001, Hsiang Hsu, Shaohan Hu, Marco Pistoia, Tarek F. Abdelzaher
ICML1
2024 Fine-grained Control of Generative Data Augmentation in IoT Sensing
abstract
Internet of Things (IoT) sensing models often suffer from overfitting due to data distribution shifts between training dataset and real-world scenarios. To address this, data augmentation techniques have been adopted to enhance model robustness by bolstering the diversity of synthetic samples within a defined vicinity of existing samples. This paper introduces a novel paradigm of data augmentation for IoT sensing signals by adding fine-grained control to generative models. We define a metric space with statistical metrics that capture the essential features of the short-time Fourier transformed (STFT) spectrograms of IoT sensing signals. These metrics serve as strong conditions for a generative model, enabling us to tailor the spectrogram characteristics in the time-frequency domain according to specific application needs. Furthermore, we propose a set of data augmentation techniques within this metric space to create new data samples. Our method is evaluated across various generative models, datasets, and downstream IoT sensing models. The results demonstrate that our approach surpasses the conventional transformation-based data augmentation techniques and prior generative data augmentation models.
Tianshi Wang 0002, Qikai Yang, Ruijie Wang 0004, Dachun Sun, Jinyang Li 0004, Yizhuo Chen, Yigong Hu, Chaoqi Yang, Tomoyoshi Kimura, Denizhan Kara, Tarek F. Abdelzaher
NeurIPS6
2024 PhyMask: An Adaptive Masking Paradigm for Efficient Self-Supervised Learning in IoT
abstract
This paper introduces PhyMask, an adaptive masking paradigm designed to enhance the efficiency and interpretability of Masked Autoencoders (MAEs) in analyzing IoT sensing signals. Different from all mainstream MAEs, which rely on random masking techniques, PhyMask employs an adaptive masking strategy that aligns with critical signal information. Its main contributions are threefold. First, PhyMask leverages the energy significance of frequency components to prioritize information-rich time-frequency regions, improving the reconstruction of original signals. Second, it includes a coherence-based masking component to identify and preserve essential temporal dynamics within the data. Finally, PhyMask integrates these components into an adaptive masking paradigm tailored to optimize the sensing context awareness within the masking configuration, focusing on the most informative parts of the data. This allows PhyMask to mask up to 96% of the input, reducing memory requirements by 14% and accelerating pre-training. Evaluations across two sensing applications, four datasets, and two real-world deployments demonstrate PhyMask's superior performance. PhyMask improves MAE accuracy by 7%, reduces pre-training data requirements by up to 75%, and enhances robustness to domain shifts and signal quality variations, making it of great value to robust and efficient intelligent IoT deployments.
Denizhan Kara, Tomoyoshi Kimura, Yatong Chen 0001, Jinyang Li 0004, Ruijie Wang 0004, Yizhuo Chen, Tianshi Wang 0002, Shengzhong Liu, Tarek F. Abdelzaher
SenSys6
2024 MetaHKG: Meta Hyperbolic Learning for Few-shot Temporal Reasoning
abstract
This paper investigates the few-shot temporal reasoning capability within the hyperbolic space. The goal is to forecast future events for newly emerging entities within temporal knowledge graphs (TKGs), leveraging only a limited set of initial observations. Hyperbolic space is advantageous for modeling emerging graph entities for two reasons: First, its geometric property of exponential expansion aligns with the rapid growth of new entities in real-world graphs; Second, it excels in capturing power-law patterns and hierarchical structures, well-suitable for new entities distributed at the peripheries of graph hierarchies and loosely connected with others through few links. We therefore propose a meta-learning framework, MetaHKG, to enable few-shot temporal reasoning within a hyperbolic space. Unlike prior hyperbolic learning works, MetaHKG addresses the challenges of effectively representing new entities in TKGs and adapting model parameters by incorporating novel hyperbolic time encodings and temporal attention networks that achieve translational invariance. We also introduce a meta hyperbolic optimization algorithm to enhance model adaptation by learning both global and entity-specific parameters through bi-level optimization. Comprehensive experiments conducted on three real-world temporal knowledge graphs demonstrate the superiority of MetaHKG over a diverse range of baselines, which achieves average 5.2% relative improvements. Compared to its Euclidean counterpart, MetaHKG operates in a lower-dimensional space but yields a more stable and efficient adaptability towards new entities.
Ruijie Wang 0004, Yutong Zhang 0011, Jinyang Li 0004, Shengzhong Liu, Dachun Sun, Tianshi Wang 0002, Yizhuo Chen, Denizhan Kara, Tarek F. Abdelzaher
SIGIR8
2024 FreqMAE: Frequency-Aware Masked Autoencoder for Multi-Modal IoT Sensing
abstract
This paper presents FreqMAE, a novel self-supervised learning framework that synergizes masked autoencoding (MAE) with physics-informed insights to capture feature patterns in multi-modal IoT sensor data. FreqMAE enhances latent space representation of sensor data, reducing reliance on data labeling and improving accuracy for AI tasks. Differing from data augmentation-based methods like contrastive learning, FreqMAE's approach eliminates the need for handcrafted transformations. Adapting MAE for IoT sensing signals, we present three contributions from frequency domain insights: First, a Temporal-Shifting Transformer (TS-T) encoder that enables temporal interactions while distinguishing different frequency bands; Second, a factorized multi-modal fusion mechanism for leveraging cross-modal correlations and preserving unique modality features; Third, a hierarchically weighted loss function that emphasizes important frequency components and high Signal-to-Noise Ratio (SNR) samples. Comprehensive evaluations on two sensing applications validate FreqMAE's proficiency in reducing labeling needs and enhancing resilience against domain shifts.
Denizhan Kara, Tomoyoshi Kimura, Shengzhong Liu, Jinyang Li 0004, Dongxin Liu, Tianshi Wang 0002, Ruijie Wang 0004, Yizhuo Chen, Yigong Hu, Tarek F. Abdelzaher
WWW8
2024 Navigating Labels and Vectors: A Unified Approach to Filtered Approximate Nearest Neighbor Search
abstract
Given a query vector, approximate nearest neighbor search (ANNS) aims to retrieve similar vectors from a set of high-dimensional base vectors. However, many real-world applications jointly query both vector data and structured data, imposing label constraints such as attributes and keywords on the search, known as filtered ANNS. Effectively incorporating filtering conditions with vector similarity presents significant challenges, including index for dynamically filtered search space, agnostic query labels, computational overhead for label-irrelevant vectors, and potential inadequacy in returning results. To tackle these challenges, we introduce a novel approach called the Label Navigating Graph, which encodes the containment relationships of label sets for all vectors. Built upon graph-based ANNS methods, we develop a general framework termed Unified Navigating Graph (UNG) to bridge the gap between label set containment and vector proximity relations. UNG offers several advantages, including versatility in supporting any query label size and specificity, fidelity in exclusively searching filtered vectors, completeness in providing sufficient answers, and adaptability in integration with most graph-based ANNS algorithms. Extensive experiments on real datasets demonstrate that the proposed framework outperforms all baselines, achieving 10x speedups at the same accuracy.
Yuzheng Cai, Jiayang Shi, Yizhuo Chen, Weiguo Zheng
Proc. ACM Manag. Data3
2023 A Unified Knowledge Distillation Framework for Deep Directed Graphical Models
abstract
Knowledge distillation (KD) is a technique that transfers the knowledge from a large teacher network to a small student network. It has been widely applied to many different tasks, such as model compression and federated learning. However, existing KD methods fail to generalize to general deep directed graphical models (DGMs) with arbitrary layers of random variables. We refer by deep DGMs to DGMs whose conditional distributions are parameterized by deep neural networks. In this work, we propose a novel unified knowledge distillation framework for deep DGMs on various applications. Specifically, we leverage the reparameterization trick to hide the intermediate latent variables, resulting in a compact DGM. Then we develop a surrogate distillation loss to reduce error accumulation through multiple layers of random variables. Moreover, we present the connections between our method and some existing knowledge distillation approaches. The proposed framework is evaluated on four applications: data-free hierarchical variational autoencoder (VAE) compression, data-free variational recurrent neural networks (VRNN) compression, data-free Helmholtz Machine (HM) compression, and VAE continual learning. The results show that our distillation method out-performs the baselines in data-free model compression tasks. We further demonstrate that our method significantly improves the performance of KD-based continual learning for data generation. Our source code is available at https://github.com/YizhuoChen99/KD4DGM-CVPR.
Yizhuo Chen, Kaizhao Liang, Zhe Zeng 0001, Shuochao Yao, Huajie Shao
CVPR1
2023 TwinSync: A Digital Twin Synchronization Protocol for Bandwidth-Limited IoT Applications
abstract
Digital Twins are evolving as a key component in modern systems with diverse applications like remote prognostics, optimizing run-time operation, anomaly detection, and more. The essential elements of a digital twin are a virtual representation, a physical asset, and the transfer of data/information between the two. IoT deployments are generally characterized by resource constraints, making synchronization of digital twins with IoT devices more challenging. There is a pressing need to optimize the bandwidth of the data transferred between the system and the twin, while ensuring that the twin is able to capture selected key aspects of the current operational state accurately. In this paper, we present TwinSync, a framework that can be utilized to construct flexible real-time representations of deployed IoT systems and efficiently synchronize relevant system states with the twin, over a communication bottleneck, within a configurable application-specific notion of error (henceforth referred to as approximate synchronization). Our approach is optimized to achieve data transfers utilizing less bandwidth without compromising the ability of the twin to replicate real-time system states within the specified approximate synchronization semantics. We evaluate the efficacy of TwinSync's synchronization by conducting both a synthetic analysis and a case study based on a real-life application prototype. Our evaluation indicates that using TwinSync can provide the same or greater accuracy (in many cases) while sending significantly fewer bytes than a bandwidth-insensitive synchronization approach. The result is attributed to a more judicial selection of data to transmit over bottlenecks, compared to bandwidth-insensitive approaches.
Deepti Kalasapura, Jinyang Li 0004, Shengzhong Liu, Yizhuo Chen, Ruijie Wang 0004, Tarek F. Abdelzaher, Matthew Caesar 0001, Joydeep Bhattacharyya, Jae Kim, Guijun Wang, Greg Kimberly, Josh D. Eckhardt, Denis Osipychev
ICCCN4
2022 Rethinking Controllable Variational Autoencoders
abstract
The Controllable Variational Autoencoder (ControlVAE) combines automatic control theory with the basic VAE model to manipulate the KL-divergence for overcoming posterior collapse and learning disentangled representations. It has shown success in a variety of applications, such as image generation, disentangled representation learning, and language modeling. However, when it comes to disentangled representation learning, ControlVAE does not delve into the rationale behind it. The goal of this paper is to develop a deeper understanding of ControlVAE in learning disentangled representations, including the choice of a desired KL-divergence (i.e, set point), and its stability during training. We first fundamentally explain its ability to disentangle latent variables from an information bottleneck perspective. We show that KL-divergence is an upper bound of the variational information bottleneck. By controlling the KL-divergence gradually from a small value to a target value, ControlVAE can disentangle the latent factors one by one. Based on this finding, we propose a new DynamicVAE that leverages a modified incremental PI (proportionalintegral) controller, a variant of the proportional-integralderivative (PID) algorithm, and employs a moving average as well as a hybrid annealing method to evolve the value of KL-divergence smoothly in a tightly controlled fashion. In addition, we analytically derive a lower bound of the set point for disentangling. We then theoretically prove the stability of the proposed approach. Evaluation results on multiple benchmark datasets demonstrate that DynamicVAE achieves a good trade-off between the disentanglement and reconstruction quality. We also discover that it can separate disentangled representation learning and re-construction via manipulating the desired KL-divergence.
Huajie Shao, Haohong Lin, Longzhong Lin, Yizhuo Chen, Qinmin Yang, Han Zhao 0002
CVPR5
2022 An Entropy-Weighted Network for Polar Sea Ice Open Lead Detection From Sentinel-1 SAR Images
abstract
Sea ice leads in the Arctic Ocean and Antarctic Ocean are of great significance to polar ecology, climate change, and ship navigation. While rule-based remote-sensing classification methods, such as the threshold method, have been widely used for surveying and mapping polar sea ice leads, they have difficulty in overcoming the problems of noise and poor generalization. In this article, we presented an automated, deep learning sea ice open lead detection network in low wind speed conditions, entropy-weighted network (EW-Net). In EW-Net, a dense block was introduced into a U-Net baseline network to strengthen feature propagation, and an entropy sampling structure was designed to improve the pertinence of the network to leads and alleviate the influence of noise. Furthermore, an entropy-weighted feature fusion block was designed to better restore features. EW-Net was trained on the Sentinel-1 synthetic aperture radar (SAR) dataset in order to obtain lead maps of different polar regions from Sentinel-1 images. The results showed that the proposed EW-Net model was more effective and more generalizable for polar sea ice lead detection than the threshold method and the deep learning methods U-Net and DeepLab V3 Plus, with an overall accuracy (OA) exceeding 0.97. In addition, EW-Net showed advantages in terms of outstanding generalization from cross-sensor image monitoring of leads. The method proposed in this article could be beneficial for providing accurate maps of polar sea ice leads in open water and contribute to further research on polar science.
Xiaoping Pang, Xi Zhao 0003, Guoyuan Li, Yizhuo Chen
IEEE Trans. Geosci. Remote. Sens.6