VLDB 2026 Research / reviewers in the wild / expert
Shao-Lun Huang
dblp:64/2243
· DBLP profile ↗
112ranked-venue papers
18as first author
67since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 8 first-author · 11 since 2021Computer networks · 19 · 1 first-author · 7 since 2021Theory of computation · 16 · 6 first-author · 9 since 2021Systems, architecture and hardware · 8 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive DiffusionabstractCurrent video generation models perform well at single-shot synthesis but struggle with multi-shot videos, facing critical challenges in maintaining character and background consistency across shots and flexibly generating videos of arbitrary length and shot count. To address these limitations, we introduce \textbf{FilmWeaver}, a novel framework designed to generate consistent, multi-shot videos of arbitrary length. First, it employs an autoregressive diffusion paradigm to achieve arbitrary-length video generation. To address the challenge of consistency, our key insight is to decouple the problem into inter-shot consistency and intra-shot coherence. We achieve this through a dual-level cache mechanism: a shot memory caches keyframes from preceding shots to maintain character and scene identity, while a temporal memory retains a history of frames from the current shot to ensure smooth, continuous motion. The proposed framework allows for flexible, multi-round user interaction to create multi-shot videos. Furthermore, due to this decoupled design, our method demonstrates high versatility by supporting downstream tasks such as multi-concept injection and video extension. To facilitate the training of our consistency-aware method, we also developed a comprehensive pipeline to construct a high-quality multi-shot video dataset. Extensive experimental results demonstrate that our method surpasses existing approaches on metrics for both consistency and aesthetic quality, opening up new possibilities for creating more consistent, controllable, and narrative-driven video content. Xiaokun Liu, Wenyu Qin, Meng Wang 0001, Pengfei Wan 0001, Di Zhang 0026, Kun Gai, Shao-Lun Huang |
AAAI | 10 |
| 2026 | OSDTW: Optimal Shared Depth and Task Weighting for Long-Tailed Recognition
Chang Chu, Shao-Lun Huang, Junxiong Zheng |
ICIC (5) | 3 |
| 2026 | PSformer: Parameter-Efficient Transformer with Segment Shared Attention for Time Series Forecasting
Hongkang Zhang, Shao-Lun Huang, Danny Dongning Sun, Xiao-Ping Zhang 0002 |
ICPR (1) | 5 |
| 2026 | On the Equivalence Relationships among Fisher Information, Shannon Measures and Variance
Yuan Xinjie, Tianren Peng, Zhenyu Liu 0003, Shao-Lun Huang |
ISIT | 4 |
| 2025 | Efficient Rectification Signal Validation for Optimal Functional ECO Patch GenerationabstractSynthesis-based functional Engineering Change Order (ECO) algorithms, as classified in [1], are particularly effective for addressing functional bugs. These algorithms typically involve two primary steps: (1) identifying rectification signals to address functional mismatches, and (2) generating patch circuits based on these signals. While much of the existing research focuses on enhancing step (2), step (1) often remains somewhat ad hoc and inefficient.In this paper, we propose a novel approach for systematically collecting and validating all possible sets of rectification signals from a given set of candidates. Leveraging a heuristic for grouping and ranking rectification signals, our ECO flow efficiently identifies minimal patches while achieving a highly competitive runtime. Our contributions include three key innovations: a heuristic for identifying high-quality rectification candidates, an efficient algorithm for validating all feasible sets of rectification signals, and a signal grouping and ranking technique that ensures minimal patch size. When integrated with an open-source patch generation tool, our method demonstrates an average reduction of 44% in patch sizes compared to a leading commercial ECO tool on benchmark circuits. Tzu-Yu Tung, Yu-Ling Hsu, Shao-Lun Huang, Chung-Yang Huang |
DAC | 3 |
| 2025 | Efficient Global Attention and Correlation-Aware Fusion for Hyperspectral Image ClassificationabstractHyperspectral imaging offers extensive spectral and spatial information. However, effectively utilizing this data for accurate classification remains a challenge. This study introduced the CASSX-Net, a novel framework designed to capture both short- and long-range dependencies in HSI data for land cover classification. The network combined a dual spectral-spatial feature extraction mechanism with a multi-head cross-attention module to leverage local and global feature interactions. By combining convolutional layers for short-range feature extraction with cross-attention mechanisms for long-range dependencies, the CASSX-Net addressed the intricate spectral-spatial correlations often missed by traditional CNNs. In addition, the maximal correlation fusion strategy optimally integrated the features from various pathways, improving the ability of the model to distinguish between classes with similar spectral signatures. The rigorous evaluation of four benchmark HSI datasets, including Pavia University, Pavia Centre, Salinas, and Houston 2018, demonstrated that the proposed framework consistently achieved the state-of-the-art performance, surpassing the existing methods in terms of classification accuracy and advancing the HSI land cover classification. Hongkang Zhang, Shao-Lun Huang, Ercan E. Kuruoglu |
ICASSP | 2 |
| 2025 | Canonswap: High-Fidelity and Consistent Video Face Swapping Via Canonical Space Modulation
Ye Zhu 0003, Yunfei Liu 0001, Lijian Lin, Cong Wan, Zijian Cai, Yu Li 0003, Shao-Lun Huang |
ICCV | 8 |
| 2025 | Content-Style Disentangled Audio Style Transfer via Diffusion ModelabstractDeep generative models have advanced the synthesis of high-quality audio signals, shifting the focus from audio fidelity to user-specific customization. Despite significant progress, current models struggle to generate style-consistent audio. Audio style transfer offers a more intuitive approach for capturing user intent but faces challenges in the disentanglement and interpretation of content and style. This paper introduces a novel framework for content-style disentangled audio style transfer. We introduce an interpretable, formula-based style distance that effectively disentangles content and style within the language-audio feature space. The proposed QwenAudio-Contrastive Language Audio Pretraining (Qwen-CLAP) content extraction module and the CLAP-based style disentanglement loss coordinated with the style reconstruction loss, enable interpretable disentanglement and stylization. Comprehensive experiments on our new dataset, BBCreatures, demonstrate superior stylization quality, preserving fine style details and original content. Jiasheng Lu, Yingshan Liang, Zhicheng Du, Qingyang Shi, Shao-Lun Huang |
ICME | 8 |
| 2025 | Quantum Entropy ProverabstractInformation inequalities govern the ultimate limitations in information theory and as such play a pivotal role in characterizing what values the entropy of multipartite states can take. Proving an information inequality, however, quickly becomes arduous when the number of involved parties increases. For classical systems, [Yeung, IEEE Trans. Inf. Theory (1997)] proposed a framework to prove Shannon-type inequalities via linear programming. Here, we derive an analogous framework for quantum systems, based on the strong sub-additivity and weak monotonicity inequalities for the von-Neumann entropy. Importantly, this also allows us to handle constrained inequalities, which - in the classical case - served as a crucial tool in proving the existence of non-standard, so-called non-Shannon-type inequalities [Zhang & Yeung, IEEE Trans. Inf. Theory (1998)]. Our main contribution is the Python package qITIP, for which we present the theory and demonstrate its capabilities with several illustrative examples. Shao-Lun Huang, Tobias Rippchen, Mario Berta |
ISIT | 1 |
| 2025 | On the Optimal Second-Order Convergence Rate of Minimax Estimation Under Weighted Mse
Tianren Peng, Shao-Lun Huang |
ISIT | 3 |
| 2025 | Information-Geometric Analysis of the Optimal Error Exponent in Fixed-Length Hypothesis TestingabstractHypothesis testing has emerged as a significant research area due to its applications in various domains. In many real-world scenarios, we can only obtain training samples of both hypotheses instead of the underlying distributions, and the decision rule shall be restricted to the form of testing sample statistics due to the computational requirement of highdimensional data. In this paper, we study the fixed-length hypothesis testing problem under such constraints. By applying the information-geometric method, we provide the corresponding asymptotic optimal error exponent with respect to the sequence lengths, and propose a valid decision rule. Moreover, we present a geometric interpretation of the trade-off between the sampling processes of training and testing. Qingyue Zhang 0003, Xinyi Tong 0002, Tianren Peng, Shao-Lun Huang |
ISIT | 4 |
| 2025 | A High-Dimensional Statistical Method for Optimizing Transfer Quantities in Multi-Source Transfer LearningabstractMulti-source transfer learning provides an effective solution to data scarcity in real-world supervised learning scenarios by leveraging multiple source tasks. In this field, existing works typically use all available samples from sources in training, which constrains their training efficiency and may lead to suboptimal results. To address this, we propose a theoretical framework that answers the question: what is the optimal quantity of source samples needed from each source task to jointly train the target model? Specifically, we introduce a generalization error measure based on K-L divergence, and minimize it based on high-dimensional statistical analysis to determine the optimal transfer quantity for each source task. Additionally, we develop an architecture-agnostic and data-efficient algorithm OTQMS to implement our theoretical results for target model training in multi-source transfer learning. Experimental studies on diverse architectures and two real-world benchmark datasets show that our proposed algorithm significantly outperforms state-of-the-art approaches in both accuracy and data efficiency. The code is available at https://github.com/zqy0126/OTQMS. Qingyue Zhang 0003, Haohao Fu, Guanbo Huang, Yaoyuan Liang, Chang Chu, Tianren Peng, Yanru Wu, Qi Li 0002, Yang Li 0104, Shao-Lun Huang |
NeurIPS | 10 |
| 2025 | Multi-Kernel Correlation-Attention Vision Transformer for Enhanced Contextual Understanding and Multi-Scale IntegrationabstractSignificant progress has been achieved using Vision Transformers (ViTs) in computer vision. However, challenges persist in modeling multi-scale spatial relationships, hindering effective integration of fine-grained local details and long-range global dependencies. To address this limitation, a Multi-Kernel Correlation-Attention Vision Transformer (MK-CAViT) grounded in the Hirschfeld-Gebelein-Rényi (HGR) theory was proposed, introducing three key innovations. A parallel multi-kernel architecture was utilized to extract multi-scale features through small, medium, and large kernels, overcoming the single-scale constraints of conventional ViTs. The cross-scale interactions were enhanced through the Fast-HGR attention mechanism, which models nonlinear dependencies and applies adaptive scaling to weigh connections and refine contextual reasoning. Additionally, a stable multi-scale fusion strategy was adopted, integrating dynamic normalization and staged learning to mitigate gradient variance, progressively fusing local and global contexts, and improving training stability. The experimental results on ImageNet, COCO, and ADE20K validated the superiority of MK-CAViT in classification, detection, and segmentation, surpassing state-of-the-art baselines in capturing complex spatial relationships while maintaining efficiency. These contributions can establish a theoretically grounded framework for visual representation learning and address the longstanding limitations of ViTs. Hongkang Zhang, Shao-Lun Huang, Ercan E. Kuruoglu |
NeurIPS | 2 |
| 2025 | Stabilizing and improving federated learning with highly non-iid data and client dropout
Jian Xu 0016, Meilin Yang, Wenbo Ding 0001, Shao-Lun Huang |
Appl. Intell. | 4 |
| 2025 | Transferability-Guided Cross-Domain Cross-Task Transfer LearningabstractWe propose two novel transferability metrics fast optimal transport-based conditional entropy (F-OTCE) and joint correspondence OTCE (JC-OTCE) to evaluate how much the source model (task) can benefit the learning of the target task and to learn more generalizable representations for cross-domain cross-task transfer learning. Unlike the original OTCE metric that requires evaluating the empirical transferability on auxiliary tasks, our metrics are auxiliary-free such that they can be computed much more efficiently. Specifically, F-OTCE estimates transferability by first solving an optimal transport (OT) problem between source and target distributions and then uses the optimal coupling to compute the negative conditional entropy (NCE) between the source and target labels. It can also serve as an objective function to enhance downstream transfer learning tasks including model finetuning and domain generalization (DG). Meanwhile, JC-OTCE improves the transferability accuracy of F-OTCE by including label distances in the OT problem, though it incurs additional computation costs. Extensive experiments demonstrate that F-OTCE and JC-OTCE outperform state-of-the-art auxiliary-free metrics by 21.1% and 25.8%, respectively, in correlation coefficient with the ground-truth transfer accuracy. By eliminating the training cost of auxiliary tasks, the two metrics reduce the total computation time of the previous method from 43 min to 9.32 and 10.78 s, respectively, for a pair of tasks. When applied in the model finetuning and DG tasks, F-OTCE shows significant improvements in the transfer accuracy in few-shot classification experiments, with up to 4.41% and 2.34% accuracy gains, respectively. Yang Tan 0004, Enming Zhang, Yang Li 0104, Shao-Lun Huang, Xiao-Ping Zhang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video GroundingabstractThis paper studies the spatio-temporal video grounding task, which aims to localize a spatio-temporal tube in an untrimmed video based on the given text description of an event. Existing one-stage approaches suffer from insufficient space-time interaction in two aspects: i) less precise prediction of event temporal boundaries, and ii) inconsistency in object prediction for the same event across adjacent frames. To address these issues, we propose a framework of Comprehensive Space-Time entAnglement (CoSTA) to densely entangle space-time multi-modal features for spatio-temporal localization. Specifically, we propose a space-time collaborative encoder to extract comprehensive video features and leverage Transformer to perform spatio-temporal multi-modal understanding. Our entangled decoder couples temporal boundary prediction and spatial localization via an entangled query, boasting an enhanced ability to capture object-event relationships. We conduct extensive experiments on the challenging benchmarks of HC-STVG and VidSTG, where CoSTA outperforms existing state-of-the-art methods, demonstrating its effectiveness for this task. Yaoyuan Liang, Yansong Tang, Zhao Yang 0002, Ziran Li, Jingang Wang, Wenbo Ding 0001, Shao-Lun Huang |
AAAI | 8 |
| 2024 | Task Oriented In-Domain Data AugmentationabstractLarge Language Models (LLMs) have shown superior performance in various applications and fields.To achieve better performance on specialized domains such as law and advertisement, LLMs are often continue pre-trained on in-domain data.However, existing approaches suffer from two major issues.First, in-domain data are scarce compared with general domainagnostic data.Second, data used for continual pre-training are not task-aware, such that they may not be helpful to downstream applications.We propose TRAIT, a task-oriented in-domain data augmentation framework.Our framework is divided into two parts: in-domain data selection and task-oriented synthetic passage generation.The data selection strategy identifies and selects a large amount of in-domain data from general corpora, and thus significantly enriches domain knowledge in the continual pre-training data.The synthetic passages contain guidance on how to use domain knowledge to answer questions about downstream tasks.By training on such passages, the model aligns with the need of downstream applications.We adapt LLMs to two domains: advertisement and math.On average, TRAIT improves LLM performance by 8% in the advertisement domain and 7.5% in the math domain. Simiao Zuo, Yeyun Gong, Qiang Lou, Yi Liu 0071, Shao-Lun Huang, Jian Jiao 0007 |
EMNLP | 7 |
| 2024 | NAC: Mitigating Noisy Correspondence in Cross-Modal Matching Via Neighbor Auxiliary CorrectorabstractThe presence of noisy correspondence within cross-modal matching has significantly undermined the performance of existing matching methods. In this paper, we introduce a robust framework named Neighbor Auxiliary Corrector (NAC) for alleviating noise by utilizing the neighbors, which are indicative of similar textual targets. NAC is inspired by an observation that similar texts tend to correspond to similar images. Leveraging the zero-shot capabilities of Pre-trained Language Models (PLMs), we identify the top-k nearest neighbors for each positive image-text pair. Subsequently, the side information provided by these neighbors is harnessed for both sample verification and sample rectification. Extensive experiments on benchmark datasets demonstrate that our framework can significantly boost the performance and is more robust to various levels of noisy correspondence. Haoming Huang, Jian Xu 0016, Shao-Lun Huang |
ICASSP | 4 |
| 2024 | On the Second Order Asymptotics of Covert Communications over AWGN ChannelsabstractThis work tackles the asymptotics of the maximal throughput of covert communications over AWGN channels when the covert metric is Kullback-Leibler divergence (KL divergence). It is shown that the first and second order asymptotics of the maximal throughput are$\sqrt{n\delta\log e}$and (2)${ }^{\frac{1}{2}}(n \delta)^{\frac{1}{4}}(\log e)^{\frac{3}{4}} \cdot Q^{-1}(\epsilon)$, respectively by$n$channel uses, where$\delta$and$\epsilon$are constraints imposed on covertness and channel decoding error probabilities, respectively. The technique we use in the achievability is quasi-$\varepsilon$-neighborhood notion from information geometry. For finite blocklength$n$, the generating distributions are chosen to be a family of truncated Gaussian distributions with decreasing variances. The law of decreasing is carefully designed so that it maximizes the throughput at the main channel in the asymptotic sense under the condition that the output distributions satisfy the covert constraint. For the converse, the optimality of Gaussian distribution for minimizing KL divergence under second order moment constraint is extended from dimension 1 to dimension$n$, which further leads to the direct converse bound in terms of covert metric. Xinchun Yu, Shuangqing Wei, Shao-Lun Huang, Xiao-Ping Zhang 0003 |
ICC | 3 |
| 2024 | AdaForensics: Learning A Characteristic-aware Adaptive Deepfake DetectorabstractIn this paper, we propose a characteristic-aware adaptive network named AdaForensics for deepfake detection. Most existing methods learn a fixed network to detect deepfakes based on carefully-designed network architectures. However, these methods employ the same deepfake detector for all the images despite of various facial characteristic, which fail to provide customized forgery detection for different individuals. To address this, our AdaForensics simultaneously learns characteristic-agnostic and characteristic-specific embeddings, where the detector dynamically adapts to varying faces with our designed hypernetwork on the fly. More specifically, our AdaForensics not only explores the shareable abstractions from various deepfake images, but also adapts the detector to the given characteristic at test time. To achieve this, we propose a two-branch HyperNetwork to learn an adaptive deepfake detector, which automatically adjusts the parameters based on characteristic of the input. Extensive experiments on widely-used datasets including FaceForensics, Celeb-DF and DFDC demonstrate our AdaForensics outperforms the state-of-the-art works. Xiaoke Yang, Haixu Song, Shao-Lun Huang, Yueqi Duan |
ICME | 4 |
| 2024 | Second-Order Characterization of Minimax Parameter Estimation in Restricted Parameter SpaceabstractEstimating unknown parameters in restricted parameter space is an important problem with applications in communication, statistics, and machine learning. In this paper, we adopt the conventional minimax formulation to investigate such problems. In particular, we focus on studying the second-order characterizations of the minimax risk in the asymptotic regime. We first show that the second-order convergence rate of the minimax risk depends on the local flatness of the Fisher information around its global optimum. Then, we demonstrate that the second-order terms can be computed by solving certain ordinary differential equations, where the coefficients of the second-order terms can be explicitly expressed in some cases. Finally, the estimators achieving the minimax risk are also given, which provides potential guidance for machine learning designs. Tianren Peng, Xinyi Tong 0002, Shao-Lun Huang |
ISIT | 3 |
| 2024 | On the Asymptotic HGR Maximal Correlation of Gaussian Markov ChainabstractThe Hirschfeld-Gebelein-Renyi (HGR) maximal correlation shows widespread applications in statistics and machine learning fields. This paper explores the HGR maximal correlation among two discrete time random processes that form Markov chains with infinite chain lengths. Under the specific form of Gaussian random variables, the optimal correlation functions are linear to the data. Therefore, this problem can be reduced to solving the largest singular value of a particular matrix. Then, we present the analytical expression of the asymptotic HGR maximal correlation, where a geometric interpretation is also provided. This study offers insights into the effective design of feature extraction in machine learning tasks. Tianren Peng, Xinyi Tong 0002, Shao-Lun Huang |
ITW | 3 |
| 2024 | The Second-Order Perspectives of Minimax Parameter Estimation in Restricted Space with Weighted Squared Error LossabstractIn this paper, we investigate the parameter estimation problems, where the unknown parameter is assumed to be within certain restricted parameter space. To this end, we adopt the conventional minimax formulation and apply the weighted mean-squared error as the loss function. In particular, we focus on analyzing the minimax risk in the asymptotic regime, where the second-order convergence rate, and the ordinary differential equation for computing the second-order terms are presented. Moreover, our results are applied to some widely considered probability distribution models, where the analytical expressions of the second-order terms and the corresponding estimators are explicitly provided. Finally, some numerical simulations are also presented when the second-order term cannot be explicitly expressed, which further supports our theoretical results. Tianren Peng, Xinyi Tong 0002, Shao-Lun Huang |
ITW | 3 |
| 2024 | A Non-asymptotic Framework for Characterizing Dependency Structures in Multimodal LearningabstractDependency structures between modalities have been utilized explicitly and implicitly in multimodal learning to enhance classification performance, particularly when the training samples are insufficient. Recent efforts have con-centrated on developing mathematical frameworks utilizing conditional dependency structures, but the non-asymptotic relations between the training sample size and various structures are not sufficiently addressed. To address this issue, we propose a mathematical framework that can be utilized to characterize conditional dependency structures in analytic ways. It provides an explicit description of the sample size in learning various structures in a non-asymptotic regime. Additionally, it demonstrates how task complexity and a fitness evaluation of conditional dependence structures affect the results. Furthermore, we develop an autonomously updated coefficient algorithm auto-CODES based on the theoretical framework and conduct experiments on multimodal emotion recognition tasks using the MELD dataset. The experimental results validate our theory and show the effectiveness of the proposed algorithm. Weida Wang, Yaoyuan Liang, Xinyi Tong 0002, Shao-Lun Huang |
ITW | 5 |
| 2024 | Unleashing Region Understanding in Intermediate Layers for MLLM-based Referring Expression GenerationabstractThe Multi-modal Large Language Model (MLLM) based Referring Expression Generation (REG) task has gained increasing popularity, which aims to generate an unambiguous text description that applies to exactly one object or region in the image by leveraging foundation models. We empirically found that there exists a potential trade-off between the detailedness and the correctness of the descriptions for the referring objects. On the one hand, generating sentences with more details is usually required in order to provide more precise object descriptions. On the other hand, complicated sentences could easily increase the probability of hallucinations. To address this issue, we propose a training-free framework, named ``unleash-then-eliminate'', which first elicits the latent information in the intermediate layers, and then adopts a cycle-consistency-based decoding method to alleviate the production of hallucinations. Furthermore, to reduce the computational load of cycle-consistency-based decoding, we devise a Probing-based Importance Estimation method to statistically estimate the importance weights of intermediate layers within a subset. These importance weights are then incorporated into the decoding process over the entire dataset, intervening in the next token prediction from intermediate layers.
Extensive experiments conducted on the RefCOCOg and PHD benchmarks show that our proposed framework could outperform existing methods on both semantic and hallucination-related metrics. Code will be made available in https://github.com/Glupayy/unleash-eliminate. Yaoyuan Liang, Zhuojun Cai, Guanbo Huang, Ziran Li, Jingang Wang, Shao-Lun Huang |
NeurIPS | 10 |
| 2023 | MultiEMO: An Attention-Based Correlation-Aware Multimodal Fusion Framework for Emotion Recognition in ConversationsabstractEmotion Recognition in Conversations (ERC) is an increasingly popular task in the Natural Language Processing community, which seeks to achieve accurate emotion classifications of utterances expressed by speakers during a conversation.Most existing approaches focus on modeling speaker and contextual information based on the textual modality, while the complementarity of multimodal information has not been well leveraged, few current methods have sufficiently captured the complex correlations and mapping relationships across different modalities.Furthermore, existing state-ofthe-art ERC models have difficulty classifying minority and semantically similar emotion categories.To address these challenges, we propose a novel attention-based correlation-aware multimodal fusion framework named MultiEMO, which effectively integrates multimodal cues by capturing cross-modal mapping relationships across textual, audio and visual modalities based on bidirectional multi-head crossattention layers.The difficulty of recognizing minority and semantically hard-to-distinguish emotion classes is alleviated by our proposed Sample-Weighted Focal Contrastive (SWFC) loss.Extensive experiments on two benchmark ERC datasets demonstrate that our MultiEMO framework consistently outperforms existing state-of-the-art approaches in all emotion categories on both datasets, the improvements in minority and semantically similar emotions are especially significant. Shao-Lun Huang |
ACL (1) | 2 |
| 2023 | A Joint Training-Calibration Framework for Test-Time Personalization with Label Shift in Federated LearningabstractThe data heterogeneity has been a challenging issue in federated learning in both training and inference stages, which motivates a variety of approaches to learn either personalized models for participating clients or test-time adaptations for unseen clients. One such approach is employing a shared feature representation and a customized classifier head for each client. However, previous works either neglect the global head with rich knowledge or assume the new clients have enough labeled data, which significantly limit their broader practicality. In this work, we propose a lightweight framework to tackle with the label shift issue during the model deployment by test priors estimation and model prediction calibration. We also demonstrate the importance of training a balanced global model in FL so as to guarantee the general effectiveness of prior estimation approaches. Evaluation results on benchmark datasets demonstrate the superiority of our framework for model adaptation in unseen clients with unknown label shifts. Jian Xu 0016, Shao-Lun Huang |
CIKM | 2 |
| 2023 | Robust Deep Joint Source Channel Coding with Time-Varying NoiseabstractDeep Joint Source-Channel Coding (JSCC) has gained increased attention, asserting its significance in the communication field. However, existing Deep JSCC techniques struggle to mitigate time-varying noise due to the deep neural networks being trained beforehand and fixed. To address this issue, we propose a robust deep JSCC scheme. Firstly, a multi-network parallel structure, as well as error-correcting codes, is introduced to effectively exploit label information. Secondly, a closed-form linear encoder and decoder pair is employed at the input and output ends of the channel to deal with the varying noise, which releases the neural network from dealing with a large range of varying noise levels. Thirdly, a transfer learning algorithm is utilized for estimating real-time noise statistics, which outperforms conventional estimation methods when noise statistics are time-dependent. These three components are effectively integrated as a comprehensive transmission system. Experimental results demonstrate that our optimized scheme outperforms existing approaches in the literature. Weida Wang, Xinchun Yu, Xinyi Tong 0002, Xiao-Ping Zhang 0002, Shao-Lun Huang |
GLOBECOM | 6 |
| 2023 | RCA-NOC: Relative Contrastive Alignment for Novel Object CaptioningabstractIn this paper, we introduce a novel approach to novel object captioning which employs relative contrastive learning to learn visual and semantic alignment. Our approach maximizes compatibility between regions and object tags in a contrastive manner. To set up a proper contrastive learning objective, for each image, we augment tags by leveraging the relative nature of positive and negative pairs obtained from foundation models such as CLIP. We then use the rank of each augmented tag in a list as a relative relevance label to contrast each top-ranked tag with a set of lower-ranked tags. This learning objective encourages the top-ranked tags to be more compatible with their image and text context than lower-ranked tags, thus improving the discriminative ability of the learned multi-modality representation. We evaluate our approach on two datasets and show that our proposed RCA-NOC approach outperforms state-of-the-art methods by a large margin, demonstrating its effectiveness in improving vision-language representation for novel object captioning. Jiashuo Fan, Yaoyuan Liang, Leyao Liu, Shao-Lun Huang |
ICCV | 4 |
| 2023 | Personalized Federated Learning with Feature Alignment and Classifier Collaboration
Jian Xu 0016, Xinyi Tong 0002, Shao-Lun Huang |
ICLR | 3 |
| 2023 | Mitigating Model Poisoning Attacks on Distributed Learning with Heterogeneous DataabstractGradient-based distributed learning techniques have been essential for machine model training on distributed samples without collecting raw data. However, such learning systems are vulnerable to both internal failures and external attacks. In this work, we study the Byzantine robustness of distributed learning over heterogeneous data for classification tasks, where each worker only contains data samples belonging to some specific classes (aka., label skew) and a fraction of workers are corrupted by a Byzantine adversary to conduct attacks by sending malicious gradients. Existing defenses usually fail under such heterogeneous cases. To remedy this, we propose a gradient decomposition scheme called DeSGD to achieve more robust distributed model training. The key idea for mitigating the impact of data heterogeneity on the Byzantine robustness is to divide the full global gradient into individual gradients of each data class and conduct resilient aggregation in a class-wise manner. The proposed framework can easily integrate existing advanced defense methods and local momentum mechanism. Evaluation results on the Fashion-MNIST dataset with various strong attacks demonstrate the improved robustness of learning over distributed data in the presence of both label skew and attacks. Jian Xu 0016, Guangfeng Yan, Ziyan Zheng, Shao-Lun Huang |
ICMLA | 4 |
| 2023 | An Information Theoretic Approach for Collaborative Distributed Parameter EstimationabstractIn many federated learning scenarios, the distributed nodes represent and exchange information in the form of functions or statistics of data, and the computation and communication are often restricted by the dimensionality of the functions. In this paper, we explore the collaborative distributed parameter estimation under such constraints. Specifically, we assume that each node can observe a sequence of i.i.d. sampled data and communicate some statistics of the observed data with dimensionality constraints. We characterize the Cramer-Rao lower bound (CRLB) and construct the asymptotic efficient estimator that achieves CRLB. In addition, we provide the information geometric interpretation of the CRLB as projecting the score function onto the functional subspaces spanned by the distributed nodes. Finally, we present the neural estimator to compute the optimal statistics that the nodes shall transmit to each other for continuous variables. Xinyi Tong 0002, Tianren Peng, Shao-Lun Huang |
ISIT | 3 |
| 2023 | LUNA: Language as Continuing Anchors for Referring Expression ComprehensionabstractReferring expression comprehension aims to localize a natural language description in an image. Using location priors to help reduce inaccuracies in cross-modal alignments is the state of the art for CNN-based methods tackling this problem. Recent Transformer-based models cast aside this idea, making the case for steering away from hand-designed components. In this work, we propose LUNA, which uses language as continuing anchors to guide box prediction in a Transformer decoder, and thus show that language-guided location priors can be effectively exploited in a Transformer-based architecture. Our method first initializes an anchor box from the input expression via a small "proto-decoder,'' and then uses this anchor and its refined successors as location guidance in a modified Transformer decoder. At each decoder layer, the anchor box is first used as a query for gathering multi-modal context, and then updated based on the gathered context (producing the next, refined anchor). In the end, a lightweight assessment pathway evaluates the quality of all produced anchors, yielding the final prediction in a dynamic way. This approach allows box decoding to be conditioned on learned anchors, which facilitates accurate grounding, as we shown in the experiments. Our method outperforms existing state-of-the-art methods on the datasets of ReferIt Game, RefCOCO/+/g, and Flickr30K Entities. Yaoyuan Liang, Zhao Yang 0002, Yansong Tang, Jiashuo Fan, Ziran Li, Jingang Wang, Philip Torr 0001, Shao-Lun Huang |
ACM Multimedia | 8 |
| 2023 | Predicting Events in MOBA Games: Prediction, Attribution, and EvaluationabstractThe multiplayer online battle arena (MOBA) games have become increasingly popular in recent years. Consequently, many efforts have been devoted to providing pregame or in-game predictions for them. These predictions can be used in many MOBA esports-related applications, such as artificial intelligence commentator systems, in-game data analysis, and game-assistant bots. However, these works are limited in the following two aspects: the lack of sufficient in-game features and the absence of interpretability in the prediction results. These two limitations greatly restrict the practical performance and industrial application of the current works. In this work, we collect a large-scale dataset containing rich in-game features for the popular MOBA gameHonor of Kings. We then propose to predict four types of prediction tasks in an interpretable way by attributing the predictions to the input features using two gradient-based attribution methods:Integrated GradientsandSmoothGrad. To evaluate the explanatory power of different models and attribution methods, a fidelity-based evaluation metric is further proposed. Finally, we evaluate the accuracy and fidelity of several competitive methods to assess how well machines predict events in MOBA games. Zelong Yang 0002, Yan Wang 0060, Piji Li, Shaobin Lin, Shuming Shi 0001, Shao-Lun Huang, Wei Bi |
IEEE Trans. Games | 6 |
| 2022 | Regularization Penalty Optimization for Addressing Data Quality Variance in OoD AlgorithmsabstractDue to the poor generalization performance of traditional empirical risk minimization (ERM) in the case of distributional shift, Out-of-Distribution (OoD) generalization algorithms receive increasing attention. However, OoD generalization algorithms overlook the great variance in the quality of training data, which significantly compromises the accuracy of these methods. In this paper, we theoretically reveal the relationship between training data quality and algorithm performance, and analyze the optimal regularization scheme for Lipschitz regularized invariant risk minimization. A novel algorithm is proposed based on the theoretical results to alleviate the influence of low quality data at both the sample level and the domain level. The experiments on both the regression and classification benchmarks validate the effectiveness of our method with statistical significance. Runpeng Yu 0001, Hong Zhu 0003, Kaican Li, Lanqing Hong, Rui Zhang 0003, Nanyang Ye 0001, Shao-Lun Huang, Xiuqiang He 0001 |
AAAI | 7 |
| 2022 | Finding Influential Instances for Distantly Supervised Relation ExtractionabstractDistant supervision (DS) is a strong way to expand the datasets for enhancing relation extraction (RE) models but often suffers from high label noise. Current works based on attention, reinforcement learning, or GAN are black-box models so they neither provide meaningful interpretation of sample selection in DS nor stability on different domains. On the contrary, this work proposes a novel model-agnostic instance sampling method for DS by influence function (IF), namely REIF. Our method identifies favorable/unfavorable instances in the bag based on IF, then does dynamic instance sampling. We design a fast influence sampling algorithm that reduces the computational complexity from \mathcal{O}(mn) to \mathcal{O}(1), with analyzing its robustness on the selected sampling function. Experiments show that by simply sampling the favorable instances during training, REIF is able to win over a series of baselines which have complicated architectures. We also demonstrate that REIF can support interpretable instance selection. Zifeng Wang 0008, Rui Wen 0001, Xi Chen 0003, Shao-Lun Huang, Ningyu Zhang 0001, Yefeng Zheng 0001 |
COLING | 4 |
| 2022 | A General Framework For Incomplete Cross-Modal Retrieval With Missing Labels And Missing ModalitiesabstractAmong various cross-modal retrieval methods, the supervised methods achieve the best performance by exploiting the semantic labels. However, in realistic applications, the data are not always complete with labels and full multi-modal data, which makes these methods hard to be used. In this paper, we propose a general framework for handling cross-modal retrieval tasks with both missing labels and missing modalities. To be more specific, in our framework we embed the data in each modality and labels all into a common feature space and maximize their correlation altogether. When labels or data in some modalities are missing, we can still maximize the correlation between the remaining data or labels. Combined with the label prediction and data reconstruction modules, our model can effectively extract useful information from the incomplete data for cross-modal retrieval tasks. In the extensive experiments, our model outperforms many other methods on different datasets, which proves the effectiveness and flexibility for handling incomplete data of our model. Shao-Lun Huang, Lin Zhang 0001 |
ICASSP | 2 |
| 2022 | Byzantine-Resilient Decentralized Collaborative LearningabstractDecentralized learning techniques have been increasingly popular for model training on distributed worker nodes. Unfortunately, such learning systems are vulnerable to failures and attacks. In this paper, we consider the decentralized learning problem over communication networks, in which worker nodes collaboratively train a machine learning model by exchanging model parameters with neighbors, but a fraction of nodes are corrupted by a Byzantine attacker and could conduct malicious attacks. Our key idea for mitigating Byzantine attacks is to check the direction and magnitude of the cross-update vectors (the difference between each received model and local model from the previous round) at each consensus round. We propose a similarity-based reweighting scheme to obtain a robust local model update for each worker. Our proposed method does not need to know the exact number of Byzantine nodes and can be employed in both static and time-varying networks. We evaluate our method on the Fashion-MNIST dataset with different Byzantine attacks and system sizes. Numerical results demonstrate the robustness of our proposed method against Byzantine attacks and superior performance than existing methods. Jian Xu 0016, Shao-Lun Huang |
ICASSP | 2 |
| 2022 | Communication-Efficient and Byzantine-Robust Distributed Stochastic Learning with Arbitrary Number of Corrupted WorkersabstractDistributed implementations of gradient-based algorithms have been essential for training large machine learning models on massive datasets. However, distributed learning algorithms are confronted with several challenges, including communication costs, straggler issues, and attacks from Byzantine adversaries. Existing works on attack-resilient distributed learning, e.g., the coordinate-wise median of gradients, usually neglect communication and/or straggler issues, and fail to defend against well-crafted attacks. Moreover, those methods are ineffective when more than half of workers are corrupted by a Byzantine adversary. To tackle those challenges simultaneously, we develop a robust gradient aggregation framework that is compatible with gradient compression and straggler mitigation techniques. Our proposed framework requires the parameter server to maintain an honest gradient as a reference at each iteration, thus can compute trust-score and similarity for each received gradient and tolerate arbitrary number of corrupted workers. We also provide convergence analysis of our method for non-convex optimization problems. Finally, experiments of image classification task on Fashion-MNIST dataset are conducted under various Byzantine attacks and gradient sparsification operations, and the numerical results demonstrate the effectiveness of our proposed strategy. Jian Xu 0016, Xinyi Tong 0002, Shao-Lun Huang |
ICC | 3 |
| 2022 | Multi-source Transfer Learning for Signal Detection over a Fading Channel with Co-channel InterferenceabstractFor signal detection tasks in wireless communications, most of the existing algorithms either ignore the co-channel interference or treat it as Gaussian noise, which may result in unsatisfactory accuracy when the interference is non-negligible and complex-distributed. When neural networks are motivated in the design of the detectors, difficulty arises in the training due to the fact that there are few accessible pilots in each packet. In this paper, we consider a data-driven detector based on multi-source transfer learning (MSTL) for signal detection in a fading channel with interference. The MSTL detector transfers channel knowledge of previous packets into the latest detection. In particular, we consider a linear combination of the pilots and historical symbols in the distribution space, and design the optimal combination coefficients based on the number of those symbols as well as distributions similarity. Numerical simulations on a Gauss-Markov flat Rayleigh fading channel with co-channel interference validate the advantages of our algorithms, compared with several existing training schemes including directly applying the fully connected deep neural network (FCDNN) and conventional linear minimum mean square error (LMMSE) detector. Ziyan Zheng, Xinyi Tong 0002, Xinchun Yu, Xiangxiang Xu 0001, Shao-Lun Huang |
ICC | 5 |
| 2022 | Byzantine-robust Federated Learning through Collaborative Malicious Gradient FilteringabstractGradient-based training in federated learning is known to be vulnerable to faulty/malicious clients, which are often modeled as Byzantine clients. To this end, previous work either makes use of auxiliary data at parameter server to verify the received gradients (e.g., by computing validation error rate) or leverages statistic-based methods (e.g. median and Krum) to identify and remove malicious gradients from Byzantine clients. In this paper, we remark that auxiliary data may not always be available in practice and focus on the statistic-based approach. However, recent work on model poisoning attacks has shown that well-crafted attacks can circumvent most of median- and distance-based statistical defense methods, making malicious gradients indistinguishable from honest ones. To tackle this challenge, we show that the element-wise sign of gradient vector can provide valuable insight in detecting model poisoning attacks. Based on our theoretical analysis of the Little is Enough attack, we propose a novel approach called SignGuard to enable Byzantine-robust federated learning through collaborative malicious gradient filtering. More precisely, the received gradients are first processed to generate relevant magnitude, sign, and similarity statistics, which are then collaboratively utilized by multiple filters to eliminate malicious gradients before final aggregation. Finally, extensive experiments of image and text classification tasks are conducted under recently proposed attacks and defense strategies. The numerical results demonstrate the effectiveness and superiority of our proposed approach. Jian Xu 0016, Shao-Lun Huang, Linqi Song, Tian Lan 0001 |
ICDCS | 2 |
| 2022 | PAC-Bayes Information Bottleneck
Zifeng Wang 0008, Shao-Lun Huang, Ercan E. Kuruoglu, Jimeng Sun 0001, Xi Chen 0003, Yefeng Zheng 0001 |
ICLR | 2 |
| 2022 | FedPer++: Toward Improved Personalized Federated Learning on Heterogeneous and Imbalanced DataabstractFederated learning is an emerging technique to collaboratively train machine learning models over multiple clients without exposing private data but suffers from heterogeneous data distributions across clients, which results in slow convergence and degraded model performance. This challenge has motivated several approaches to learn a personalized model that has better model accuracy than the global model for each participating client. On the other hand, the success of multi-task learning suggests that data in similar tasks often share a common feature representation, while the output layer (classifier) for classification is more task-correlated. Existing methods usually neglect the collaboration between locally trained classifiers, which makes each client failed to fully benefit from data in other clients. Therefore, we perform a linear combination based classifier collaboration method for achieving better personalized model performance. Moreover, training a local feature extractor based on local data is prone to over-fitting when local data is insufficient and cannot be fully corrected by global aggregation. To tackle this issue, we propose a feature-regularized training strategy to mitigate the local over-fitting risk and reduce the parameter divergence across clients, facilitating the global feature extractor aggregation. With extensive experiments performed on Fashion-MNIST and CIFAR-10 datasets, we demonstrate the performance improvement and robustness of our method over existing methods. Jian Xu 0016, Shao-Lun Huang |
IJCNN | 3 |
| 2022 | A Mathematical Framework to Characterize the Dependency Structures in Multimodal Learning with Minimax PrincipleabstractMultimodal learning is an increasingly important research topic. Exploiting conditional dependency across multiple modalities has been shown useful for estimating of the multimodal joint distribution, especially when the number of training samples is insufficient. However, it is difficult to theoretically characterize such conditional dependency structure. To address this issue, we establish a mathematical framework and formulate the estimation problem based on the minimax principle. Then, we propose an estimator which is close to the analytical solution of the problem under a mild assumption on the sample size. Moreover, the proposed estimator is a linear combination of the learning results from the true dependency structure and the conditional one. The combining coefficient is related to three aspects: the number of training samples, the fitness of the conditional dependency structure, and the cardinality of each modality. Finally, numerical simulations are provided to verify our theoretical results that the proposed estimator is close to the optimal solution of the formulated problem. Tianren Peng, Weida Wang, Shao-Lun Huang |
ISIT | 3 |
| 2022 | Communicating Type Classes Through Channels: An Information Geometric ViewabstractIn this paper, we study a binary hypothesis testing problem where in each hypothesis, a length-n sequence is uniformly sampled from a particular type class, and the output of transmitting this sequence through a discrete memoryless channel is observed. The goal is to characterize the error exponent of detecting the true hypothesis from this observation. In order to obtain theoretical insights, we focus on the regime that the type classes are similar. In such regime, we show that the problem can be analytically solved by a local information geometric approach, which provides fundamental geometric insights for the optimal decision rule. Moreover, we characterize the error exponents for mismatched detectors, and also provide a universal lower bound when the sequences are non-uniformly sampled from the type classes. Finally, we present a potential application of this problem by analyzing a coding scheme of the unequal error protection (UEP), which encodes the special bit by different input distributions. Our result provides an achievability lower bound for the UEP. Shao-Lun Huang |
ITW | 1 |
| 2022 | An Information-theoretic Method for Collaborative Distributed Learning with Limited CommunicationabstractIn this paper, we study the information transmission problem under the distributed learning framework, where each worker node is merely permitted to transmit a m-dimensional statistic to improve learning results of the target node. Specifically, we evaluate the corresponding expected population risk (EPR) under the regime of large sample sizes. We prove that the performance can be enhanced since the transmitted statistics contribute to estimating the underlying distribution under the mean square error measured by the EPR norm matrix. Accordingly, the transmitted statistics correspond to the eigenvectors of this matrix, and the desired transmission allocates these eigenvectors among the statistics such that the EPR is minimal. Moreover, we provide the analytical solution of the desired statistics for single-node and two-node transmission, where a geometrical interpretation is given to explain the eigenvector selection. For the general case, an efficient algorithm that can output the allocation solution is developed based on the node partitions. Xinyi Tong 0002, Jian Xu 0016, Shao-Lun Huang |
ITW | 3 |
| 2022 | Revisiting Sparse Convolutional Model for Visual RecognitionabstractDespite strong empirical performance for image classification, deep neural networks are often regarded as ``black boxes'' and they are difficult to interpret. On the other hand, sparse convolutional models, which assume that a signal can be expressed by a linear combination of a few elements from a convolutional dictionary, are powerful tools for analyzing natural images with good theoretical interpretability and biological plausibility. However, such principled models have not demonstrated competitive performance when compared with empirically designed deep networks. This paper revisits the sparse convolutional modeling for image classification and bridges the gap between good empirical performance (of deep learning) and good interpretability (of sparse convolutional models). Our method uses differentiable optimization layers that are defined from convolutional sparse coding as drop-in replacements of standard convolutional layers in conventional deep neural networks. We show that such models have equally strong empirical performance on CIFAR-10, CIFAR-100 and ImageNet datasets when compared to conventional neural networks. By leveraging stable recovery property of sparse modeling, we further show that such models can be much more robust to input corruptions as well as adversarial perturbations in testing through a simple proper trade-off between sparse regularization and data reconstruction terms. Xili Dai, Pengyuan Zhai, Shengbang Tong, Xingjian Gao, Shao-Lun Huang, Zhihui Zhu, Chong You, Yi Ma 0001 |
NeurIPS | 6 |
| 2022 | AC-SGD: Adaptively Compressed SGD for Communication-Efficient Distributed LearningabstractGradient compression (e.g., gradient quantization and gradient sparsification) is a core technique in reducing communication costs in distributed learning systems. The recent trend of gradient compression is to use a varying number of bits across iterations, however, relying on empirical observations or engineering heuristics without a systematic treatment and analysis. To the best of our knowledge, a general dynamic gradient compression that leverages both quantization and sparsification techniques is still far from understanding. This paper proposes a novel Adaptively-Compressed Stochastic Gradient Descent (AC-SGD) strategy to adjust the number of quantization bits and the sparsification size with respect to the norm of gradients, the communication budget, and the remaining number of iterations. In particular, we derive an upper bound, tight in some cases, of the convergence error for arbitrary dynamic compression strategy. Then we consider communication budget constraints and propose an optimization formulation - denoted as theAdaptive Compression Problem (ACP)- for minimizing the deep model’s convergence error under such constraints. By solving the ACP, we obtain an enhanced compression algorithm that significantly improves model accuracy under given communication budget constraints. Finally, through extensive experiments on computer vision and natural language processing tasks on MNIST, CIFAR-10, CIFAR-100 and AG-News datasets, respectively, we demonstrate that our compression scheme significantly outperforms the state-of-the-art gradient compression methods in terms of mitigating communication costs. Guangfeng Yan, Tan Li 0002, Shao-Lun Huang, Tian Lan 0001, Linqi Song |
IEEE J. Sel. Areas Commun. | 3 |
| 2022 | Interpretable Real-Time Win Prediction for Honor of Kings - A Popular Mobile MOBA EsportabstractWith the rapid prevalence and explosive development of Multiplayer Online Battle Arena electronic sports (MOBA esports), much research effort has been devoted to automatically predicting game results (win predictions). While this task has great potential in various applications, such as esports live streaming and game commentator artificial intelligence systems, previous studies fail to investigate the methods tointerpretthese win predictions. To mitigate this issue, we collected a large-scale dataset that contains real-time game records with rich input features of the popular MOBA gameHonor of Kings. For interpretable predictions, we proposed a two-stage spatial–temporal network (TSSTN) that can not only provide accurate real-time win predictions but also attribute the ultimate prediction results to the contributions of different features for interpretability. Experiment results and applications in real-world live streaming scenarios showed that the proposed TSSTN model is effective in both prediction accuracy and interpretability. Zelong Yang 0002, Zhufeng Pan, Yan Wang 0060, Deng Cai 0002, Shuming Shi 0001, Shao-Lun Huang, Wei Bi, Xiaojiang Liu |
IEEE Trans. Games | 6 |
| 2021 | OTCMR: Bridging Heterogeneity Gap with Optimal Transport for Cross-modal RetrievalabstractCross-modal retrieval is a classic task in the multimedia community, which aims to search for semantically similar results from different modalities. The core of cross-modal retrieval is to learn the most correlated features in a common feature space for the multi-modal data so that the similarity can be directly measured. In this paper, we propose a novel model using optimal transport for bridging the heterogeneity gap in cross-modal retrieval tasks. Specifically, we calculate the optimal transport plans between feature distributions of different modalities and then minimize the transport cost by optimizing the feature embedding functions. In this way, the feature distributions of multi-modal data can be well aligned in the common feature space. In addition, our model combines the complementary losses in different levels: 1) semantic level, 2) distributional level, and 3) pairwise level for improving cross-modal retrieval performance. In extensive experiments, our method outperforms many other cross-modal retrieval methods, which proves the efficacy of using optimal transport in cross-modal retrieval tasks. Shao-Lun Huang, Lin Zhang 0001 |
CIKM | 2 |
| 2021 | OTCE: A Transferability Metric for Cross-Domain Cross-Task RepresentationsabstractTransfer learning across heterogeneous data distributions (a.k.a. domains) and distinct tasks is a more general and challenging problem than conventional transfer learning, where either domains or tasks are assumed to be the same. While neural network based feature transfer is widely used in transfer learning applications, finding the optimal transfer strategy still requires time-consuming experiments and domain knowledge. We propose a transferability metric called Optimal Transport based Conditional Entropy (OTCE), to analytically predict the transfer performance for supervised classification tasks in such cross-domain and cross-task feature transfer settings. Our OTCE score characterizes transferability as a combination of domain difference and task difference, and explicitly evaluates them from data in a unified framework. Specifically, we use optimal transport to estimate domain difference and the optimal coupling between source and target distributions, which is then used to derive the conditional entropy of the target task (task difference). Experiments on the largest cross-domain dataset DomainNet and Office31 demonstrate that OTCE shows an average of 21% gain in the correlation with the ground truth transfer accuracy compared to state-of-the-art methods. We also investigate two applications of the OTCE score including source model selection and multi-source feature fusion. Yang Tan 0004, Yang Li 0104, Shao-Lun Huang |
CVPR | 3 |
| 2021 | Semi-Supervised Multimodal Image Translation for Missing Modality ImputationabstractMissing data is a common problem in multimodal and multi-view learning. It raises a critical challenge for most multimodal algorithms, which are unable to deal with incomplete datasets. Rather than discarding entries with missing modalities, this paper aims to reconstruct the complete image-based multimodal data by imputing missing modalities. We solve the imputation problem as an image translation task, which transforms images in one domain to other domains. Existing image translation techniques either can not fully utilize the information contained in partially complete entries or are limited to the bimodal situation. We propose a semi-supervised algorithm for multimodal learning with missing data, namely Cyclic Autoencoder (CycAE). Specifically, a novel cyclical structure, as well as the correlation among modalities, is integrated to leverage infoπnation from complete entries to incomplete ones. Experiments on two multimodal datasets show that our model outperforms state-of-the-art models. Downstream tasks can also benefit from the completed datasets. Wangbin Sun, Fei Ma 0006, Yang Li 0104, Shao-Lun Huang, Shiguang Ni, Lin Zhang 0001 |
ICASSP | 4 |
| 2021 | 2021 CAD Contest Problem A: Functional ECO with Behavioral Change Guidance Invited PaperabstractFunctional ECO is an essential solution in the VLSI design flow. The technique is to realize the functional changes with a minimal patch netlist in the gate-level netlist. As the increasing of the design complexity, functional ECO becomes more and more difficult to generate a minimal patch. ICCAD 2021 CAD contest calls for a feasible and efficient ECO algorithm with behavioral change guidance. More than ordinary functional ECO problems, the RTL designs are provided. Contestants can utilize the behavioral change in RTL designs to minimize the patch for G1. Yen-Chun Fang, Shao-Lun Huang, Chi-An Wu, Chung-Han Chou, Chih-Jen Hsu, WoeiTzy Jong, Kei-Yong Khoo |
ICCAD | 2 |
| 2021 | An Efficient Approach for Audio-Visual Emotion Recognition With Missing Labels And Missing ModalitiesabstractAudio-visual emotion recognition is important for human-machine interaction systems by combining the information of audio and visual modalities. Although great progress has been made by previous works using multimodal learning compared with unimodal learning, they still cannot effectively deal with two key challenges. Firstly, it is difficult or expensive to acquire labeled emotional data, which results in a large amount of data with missing labels. Secondly, emotional data often has missing modalities. To address these problems, we propose a unified deep learning framework to efficiently handle missing labels and missing modalities for audio-visual emotion recognition through correlation analysis. Specifically, we consider four types of emotional data during the training stage: complete, label missing, visual missing, and audio missing. We propose a correlation loss based on Hirschfeld-Gebelein-Ŕenyi (HGR) maximal correlation to effectively capture the common information in different types of training data for emotion prediction. Experiments on the eNTERFACE’05 and RAVDESS datasets show that our deep learning approach has high effectiveness for audio-visual emotion recognition. Fei Ma 0006, Shao-Lun Huang, Lin Zhang 0001 |
ICME | 2 |
| 2021 | Dual Feature Distributional Regularization for Defending Against Adversarial Attacks
Xiangxiang Xu 0001, Shao-Lun Huang, Lin Zhang 0001 |
ICONIP (6) | 3 |
| 2021 | Enhancing Neural Network Based Hybrid Learning with Empirical Wavelet Transform for Time Series ForecastingabstractOver the past decades, the hybrid models which integrate parametric and non-parametric learning models have proven to be a viable method in time series forecasting. Several structures were proposed as a combination of parametric models such as autoregressive moving average (ARIMA) and non-parametric models such as artificial neural network (ANN). Although these models show superior performance than a single model, there is the scope of further improvement if the underlying feature laid in the original data is augmented before applying models. In this work, empirical wavelet transform (EWT) is implemented to decompose the original series into several sub-series which contain unique features at different frequency horizon. The sub-series are then used along with moving average filter (MA), ARIMA and ANN to perform forecasting tasks. The experiments were performed on four public real-world data sets and compared to the other six benchmark models. The results showed that the proposed model achieved considerably higher forecast accuracy. Bunchalit Eua-Arporn, Shao-Lun Huang, Ercan E. Kuruoglu |
ICTAI | 2 |
| 2021 | Live Gradient Compensation for Evading Stragglers in Distributed LearningabstractThe training efficiency of distributed learning systems is vulnerable to stragglers, namely, those slow worker nodes. A naive strategy is performing the distributed learning by incor-porating the fastest K workers and ignoring these stragglers, which may induce high deviation for non-IID data. To tackle this, we develop a Live Gradient Compensation (LGC) strategy to incorporate the one-step delayed gradients from stragglers, aiming to accelerate learning process and utilize the stragglers simultaneously. In LGC framework, mini-batch data are divided into smaller blocks and processed separately, which makes the gradient computed based on partial work accessible. In addition, we provide theoretical convergence analysis of our algorithm for non-convex optimization problem under non-IID training data to show that LGC-SGD has almost the same convergence error as full synchronous SGD. The theoretical results also allow us to quantify a novel tradeoff in minimizing training time and error by selecting the optimal straggler threshold. Finally, extensive simulation experiments of image classification on CIFAR-10 dataset are conducted, and the numerical results demonstrate the effectiveness of our proposed strategy. Jian Xu 0016, Shao-Lun Huang, Linqi Song, Tian Lan 0001 |
INFOCOM | 2 |
| 2021 | An Information Theoretic Framework for Distributed Learning AlgorithmsabstractDistributed learning is recently an important research topic, while the information theoretic optimality of the distributed learning algorithms is often not sufficiently addressed. This paper studies the distributed learning problems such that each node observes i.i.d. samples and sends a feature function of observed samples to the central machine for decision making. Both the binary hypothesis testing in information theory and the classification problems in machine learning are considered, and the optimal error exponent and the set of optimal features are characterized. By exploiting an information theoretic framework, we show that these two problems share the same set of optimal features, from which the information theoretic optimality of some machine learning algorithms can be established. Finally, we generalize our analyses to$M$-ary distributed hypothesis testing and classification problems. A full version of this paper is accessible at: https://xiangxiangxu.com/media/documents/isit2021.pdf Xiangxiang Xu 0001, Shao-Lun Huang |
ISIT | 2 |
| 2021 | Exact Recovery in the Balanced Stochastic Block Model with Side InformationabstractThe role that side information plays in improving the exact recovery threshold in the stochastic block model (SBM) has been studied in many aspects. This paper studies exact recovery in n node balanced binary symmetric SBM with side information, given in the form of $O(\log n)$ i.i.d. samples at each node. A sharp exact recovery threshold is obtained and turns out to coincide with an existing threshold result, where no balanced constraint is imposed. Our main contribution is an efficient semi-definite programming (SDP) algorithm that achieves the optimal exact recovery threshold. Compared to the existing works on SDP algorithm for SBM with constant number of samples as side information, the challenge in this paper is to deal with the number of samples increasing in n. Jin Sima, Shao-Lun Huang |
ITW | 3 |
| 2021 | On Distributed Hypothesis Testing with Constant-Bit Communication ConstraintsabstractIn this paper, we consider the distributed hypothesis testing (DHT) problem where two nodes are constrained to transmit constant bits to a central decoder. In such cases, we show that in order to achieve the optimal error exponents, it suffices to consider the empirical distributions of observed data sequences and encode them to the transmission bits. With such a coding strategy, we develop a geometric approach in the distribution spaces to show the optimal achievable error exponents and coding scheme for the following cases: (i) both nodes can transmit $\log_{2}3$ bits; (ii) one of the nodes can transmit 1 bit, and the other node is not constrained; (iii) the joint distribution of the nodes are conditionally independent given one hypothesis. Our approach essentially reveals new potentials for characterizing the precise error exponents for DHT with general communication constraints. Xiangxiang Xu 0001, Shao-Lun Huang |
ITW | 2 |
| 2021 | On the Optimal Error Rate of Stochastic Block Model with Symmetric Side InformationabstractSide information improves the accuracy in community detection problems. While experimental results demonstrate the superior performance of many detection methods based on both the node attributes and graph structure, the question of the fundamental limit of the error rate for exact recovery remains open. In this paper, we obtain the asymptotic optimal error rate in the sense of exact recovery for a special two-community symmetric stochastic block model (SSBM) with side information consisting of multiple features. Our result provides insight on the number of features and nodes in the graph needed for community detection. Jin Sima, Shao-Lun Huang |
ITW | 3 |
| 2021 | DQ-SGD: Dynamic Quantization in SGD for Communication-Efficient Distributed LearningabstractGradient quantization is an emerging technique in reducing communication costs in distributed learning. Existing gradient quantization algorithms often rely on engineering heuristics or empirical observations, lacking a systematic approach to dynamically quantize gradients. This paper addresses this issue by proposing a novel dynamically quantized SGD (DQ-SGD) framework, enabling us to dynamically adjust the quantization scheme for each gradient descent step by exploring the trade-off between communication cost and convergence error. We derive an upper bound, tight in some cases, of the convergence error for a restricted family of quantization schemes and loss functions. We design our DQSGD algorithm via minimizing the communication cost under the convergence error constraints. Finally, through extensive experiments on large-scale natural language processing and computer vision tasks on AG-News, CFAR-10, and CIFAR-100 datasets, we demonstrate that our quantization scheme achieves better tradeoffs between the communication cost and learning performance than other state-of-the-art gradient quantization methods. Guangfeng Yan, Shao-Lun Huang, Tian Lan 0001, Linqi Song |
MASS | 2 |
| 2021 | A Mathematical Framework for Quantifying Transferability in Multi-source Transfer LearningabstractCurrent transfer learning algorithm designs mainly focus on the similarities between source and target tasks, while the impacts of the sample sizes of these tasks are often not sufficiently addressed. This paper proposes a mathematical framework for quantifying the transferability in multi-source transfer learning problems, with both the task similarities and the sample complexity of learning models taken into account. In particular, we consider the setup where the models learned from different tasks are linearly combined for learning the target task, and use the optimal combining coefficients to measure the transferability. Then, we demonstrate the analytical expression of this transferability measure, characterized by the sample sizes, model complexity, and the similarities between source and target tasks, which provides fundamental insights of the knowledge transferring mechanism and the guidance for algorithm designs. Furthermore, we apply our analyses for practical learning tasks, and establish a quantifiable transferability measure by exploiting a parameterized model. In addition, we develop an alternating iterative algorithm to implement our theoretical results for training deep neural networks in multi-source transfer learning tasks. Finally, experiments on image classification tasks show that our approach outperforms existing transfer learning algorithms in multi-source and few-shot scenarios. Xinyi Tong 0002, Xiangxiang Xu 0001, Shao-Lun Huang, Lizhong Zheng |
NeurIPS | 3 |
| 2021 | Lifelong Learning Based Disease Diagnosis on Clinical Notes
Zifeng Wang 0008, Yifan Yang 0006, Rui Wen 0001, Xi Chen 0003, Shao-Lun Huang, Yefeng Zheng 0001 |
PAKDD (1) | 5 |
| 2021 | A Semi-supervised Learning Approach for Visual Question Answering based on Maximal CorrelationabstractIn this paper, we propose a semi-supervised learning approach for the Visual Question Answering (VQA) task based on maximal correlation. Instead of training the VQA model with just classification loss like cross-entropy, we propose a semi-supervised loss function to incorporate Soft-HGR, a training approach based on Hirschfeld-Gebelein-Rényi (HGR) maximal correlation, to realize semi-supervised model training. With Soft-HGR, the high-order correlation from cross-modal common information of VQA image-question pairs is utilized to improve VQA model performance even without discriminative supervision from answer labels. We conduct experiments on the VQA v2 dataset by training the VQA model with different percentages of unlabeled samples. Experimental results show that our approach is efficient and model-agnostic for this semi-supervised learning task. Sikai Yin, Fei Ma 0006, Shao-Lun Huang |
SMC | 3 |
| 2021 | Online Disease Diagnosis with Inductive Heterogeneous Graph Convolutional NetworksabstractWe propose a Healthcare Graph Convolutional Network (HealGCN) to offer disease self-diagnosis service for online users based on Electronic Healthcare Records (EHRs). Two main challenges are focused in this paper for online disease diagnosis: (1) serving cold-start users via graph convolutional networks and (2) handling scarce clinical description via a symptom retrieval system. To this end, we first organize the EHR data into a heterogeneous graph that is capable of modeling complex interactions among users, symptoms and diseases, and tailor the graph representation learning towards disease diagnosis with an inductive learning paradigm. Then, we build a disease self-diagnosis system with a corresponding EHR Graph-based Symptom Retrieval System (GraphRet) that can search and provide a list of relevant alternative symptoms by tracing the predefined meta-paths. GraphRet helps enrich the seed symptom set through the EHR graph when confronting users with scarce descriptions, hence yield better diagnosis accuracy. At last, we validate the superiority of our model on a large-scale EHR dataset. Zifeng Wang 0008, Rui Wen 0001, Xi Chen 0003, Shilei Cao 0001, Shao-Lun Huang, Buyue Qian, Yefeng Zheng 0001 |
WWW | 5 |
| 2021 | On the Sample Complexity of HGR Maximal Correlation Functions for Large DatasetsabstractThe Hirschfeld-Gebelein-Rényi (HGR) maximal correlation and the corresponding functions have been shown useful in many machine learning scenarios. In this paper, we study the sample complexity of estimating the HGR maximal correlation functions by the alternating conditional expectation (ACE) algorithm using training samples from large datasets. Specifically, we develop a mathematical framework to characterize the learning errors between the maximal correlation functions computed from the true distribution, and the functions estimated from the ACE algorithm. For both supervised and semi-supervised learning scenarios, we establish the analytical expressions for the error exponents of the learning errors. Furthermore, we demonstrate that for large datasets, the upper bounds for the sample complexity of learning the HGR maximal correlation functions by the ACE algorithm can be expressed using the established error exponents. Moreover, with our theoretical results, we investigate the sampling strategy for different types of samples in semi-supervised learning with a total sampling budget constraint, and an optimal sampling strategy is developed to maximize the error exponent of the learning error. Finally, the numerical simulations are presented to support our theoretical results. Shao-Lun Huang, Xiangxiang Xu 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2020 | Less Is Better: Unweighted Data Subsampling via Influence FunctionabstractIn the time of Big Data, training complex models on large-scale data sets is challenging, making it appealing to reduce data volume for saving computation resources by subsampling. Most previous works in subsampling are weighted methods designed to help the performance of subset-model approach the full-set-model, hence the weighted methods have no chance to acquire a subset-model that is better than the full-set-model. However, we question that how can we achieve better model with less data? In this work, we propose a novel Unweighted Influence Data Subsampling (UIDS) method, and prove that the subset-model acquired through our method can outperform the full-set-model. Besides, we show that overly confident on a given test set for sampling is common in Influence-based subsampling methods, which can eventually cause our subset-model's failure in out-of-sample test. To mitigate it, we develop a probabilistic sampling scheme to control the worst-case risk over all distributions close to the empirical distribution. The experiment results demonstrate our methods superiority over existed subsampling methods in diverse tasks, such as text classification, image classification, click-through prediction, etc. Zifeng Wang 0008, Hong Zhu 0003, Zhenhua Dong, Xiuqiang He 0001, Shao-Lun Huang |
AAAI | 5 |
| 2020 | Semantically Supervised Maximal Correlation For Cross-Modal RetrievalabstractWith the rapid growth of multimedia data, the cross-modal retrieval problem has attracted a lot of interest in both research and industry in recent years. However, the inconsistency of data distribution from different modalities makes such task challenging. In this paper, we propose Semantically Supervised Maximal Correlation (S2MC) method for cross-modal retrieval by incorporating semantic label information into the traditional maximal correlation framework. Combining with maximal correlation based method for extracting unsupervised pairing information, our method effectively exploits supervised semantic information on both common feature space and label space. Extensive experiments show that our method outperforms other current state-of-the-art methods on cross-modal retrieval tasks on three widely used datasets. Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001 |
ICIP | 3 |
| 2020 | Person Recognition with HGR Maximal Correlation on Multimodal DataabstractMultimodal person recognition is a common task in video analysis and public surveillance, where information from multiple modalities, such as images and audio extracted from videos, are used to jointly determine the identity of a person. Previous person recognition techniques either use only uni-modal data or only consider shared representations between different input modalities, while leaving the extraction of their relationship with identity information to downstream tasks. Furthermore, real-world data often contain noise, which makes recognition more challenging practical situations. In our work, we propose a novel correlation-based multimodal person recognition framework that is relatively simple but can efficaciously learn supervised information in multimodal data fusion and resist noise. Specifically, our framework learns a discriminative embeddings of persons by joint learning visual features and audio features while maximizing HGR maximal correlation among multimodal input and persons' identities. Experiments are done on a subset of Voxceleb2. Compared with state-of-the-art methods, the proposed method demonstrates an improvement of accuracy and robustness to noise. Yihua Liang, Fei Ma 0006, Yang Li 0104, Shao-Lun Huang |
ICPR | 4 |
| 2020 | On the Sample Complexity of Estimating Small Singular ModesabstractWhile it is commonly believed that estimating the small singular modes for a nearly low-rank matrix requires more samples, the sample size needed is generally unclear. In this paper, we investigate this sample complexity by considering the difference between the estimation errors of estimating a matrix with or without estimating these small singular modes. Specifically, we develop a mathematical framework based on the matrix perturbation analysis to characterize the noise level of estimating small singular modes by n samples. In particular, we show that under mild assumptions on the sample noise, it requires at least n = O(η-2) samples to well estimate the singular modes with the singular value in the order of some small η. More importantly, our results are applied to the channel state estimation and Hirschfeld-Gebelein-Rényi (HGR) maximal correlation problems, from which we characterize that for how many samples, utilizing the low-rank approximation in these problems are beneficial. Finally, numerical simulations are provided to verify our results. Xiangxiang Xu 0001, Weida Wang, Shao-Lun Huang |
ISIT | 3 |
| 2020 | A Local Characterization for Wyner Common InformationabstractWhile the Hirschfeld-Gebelein-Rényi (HGR) maximal correlation and the Wyner common information share similar information processing purposes of extracting common knowledge structures between random variables, the relationships between these approaches are generally unclear. In this paper, we demonstrate such relationships by considering the Wyner common information in the weakly dependent regime, called ε-common information. We show that the HGR maximal correlation functions coincide with the relative likelihood functions of estimating the auxiliary random variables in ε-common information, which establishes the fundamental connections these approaches. Moreover, we extend the ε-common information to multiple random variables, and derive a novel algorithm for extracting feature functions of data variables regarding their common information. Our approach is validated by the MNIST problem, and can potentially be useful in multi-modal data analyses. Shao-Lun Huang, Xiangxiang Xu 0001, Lizhong Zheng, Gregory W. Wornell |
ISIT | 1 |
| 2020 | Information Theoretic Counterfactual Learning from Missing-Not-At-Random FeedbackabstractCounterfactual learning for dealing with missing-not-at-random data (MNAR) is an intriguing topic in the recommendation literature, since MNAR data are ubiquitous in modern recommender systems. Instead, missing-at-random (MAR) data, namely randomized controlled trials (RCTs), are usually required by most previous counterfactual learning methods. However, the execution of RCTs is extraordinarily expensive in practice. To circumvent the use of RCTs, we build an information theoretic counterfactual variational information bottleneck (CVIB), as an alternative for debiasing learning without RCTs. By separating the task-aware mutual information term in the original information bottleneck Lagrangian into factual and counterfactual parts, we derive a contrastive information loss and an additional output confidence penalty, which facilitates balanced learning between the factual and counterfactual domains. Empirical evaluation on real-world datasets shows that our CVIB significantly enhances both shallow and deep models, which sheds light on counterfactual learning in recommendation that goes beyond RCTs. Zifeng Wang 0008, Xi Chen 0003, Rui Wen 0001, Shao-Lun Huang, Ercan E. Kuruoglu, Yefeng Zheng 0001 |
NeurIPS | 4 |
| 2020 | Reproducing Scientific Experiment with Cloud DevOpsabstractThe reproducibility of scientific experiment is vital for the advancement of disciplines based on previous work. To achieve this goal, many researchers focus on complex methodology and self-invented tools which have difficulty in practical usage. In this article, we introduce the Cloud DevOps infrastructure from software engineering community and shows how it can be used effectively for heterogeneous agents to reproduce experiments for computer science related disciplines. DevOps can be enabled using freely available cloud computing machines for medium-sized experiment and self-hosted computing engines for large-scale computing, thus powering researchers to share their experiment result with others in a more reliable way. Xingzhi Niu, Shao-Lun Huang, Lin Zhang 0001 |
SERVICES | 3 |
| 2020 | Mining Regional Mobility Patterns for Urban Dynamic Analytics
Jing Lian 0003, Yang Li 0104, Weixi Gu, Shao-Lun Huang, Lin Zhang 0001 |
Mob. Networks Appl. | 4 |
| 2019 | An Efficient Approach to Informative Feature Extraction from Multimodal DataabstractOne primary focus in multimodal feature extraction is to find the representations of individual modalities that are maximally correlated. As a well-known measure of dependence, the Hirschfeld-Gebelein-Rényi (HGR) maximal correlation be-´ comes an appealing objective because of its operational meaning and desirable properties. However, the strict whitening constraints formalized in the HGR maximal correlation limit its application. To address this problem, this paper proposes Soft-HGR, a novel framework to extract informative features from multiple data modalities. Specifically, our framework prevents the “hard” whitening constraints, while simultaneously preserving the same feature geometry as in the HGR maximal correlation. The objective of Soft-HGR is straightforward, only involving two inner products, which guarantees the efficiency and stability in optimization. We further generalize the framework to handle more than two modalities and missing modalities. When labels are partially available, we enhance the discriminative power of the feature representations by making a semi-supervised adaptation. Empirical evaluation implies that our approach learns more informative feature mappings and is more efficient to optimize. Lichen Wang, Jiaxiang Wu 0001, Shao-Lun Huang, Lizhong Zheng, Xiangxiang Xu 0001, Lin Zhang 0001, Junzhou Huang |
AAAI | 3 |
| 2019 | An Information-Theoretic Approach to Transferability in Task Transfer LearningabstractTask transfer learning is a popular technique in image processing applications that uses pre-trained models to reduce the supervision cost of related tasks. An important question is to determine task transferability, i.e. given a common input domain, estimating to what extent representations learned from a source task can help in learning a target task. Typically, transferability is either measured experimentally or inferred through task relatedness, which is often defined without a clear operational meaning. In this paper, we present a novel metric, H-score, an easily-computable evaluation function that estimates the performance of transferred representations from one task to another in classification problems using statistical and information theoretic principles. Experiments on real image data show that our metric is not only consistent with the empirical transferability measurement, but also useful to practitioners in applications such as source model selection and task transfer curriculum learning. Yajie Bao, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001, Lizhong Zheng, Amir Zamir, Leonidas J. Guibas |
ICIP | 3 |
| 2019 | Maximal Correlation Embedding Network for Multilabel Learning with Missing LabelsabstractMultilabel learning, the problem of mapping each data instance to a subset of labels, appears frequently in many real-world applications. However, obtaining complete label annotation for every instance requires tremendous efforts, especially when the label set is large. As a result, multilabel learning with missing labels remains as a common challenge. Existing works either cannot handle missing labels or lack nonlinear expressiveness and scalability to large label set. In this paper, we present a novel end-to-end solution for multilabel learning with missing labels. Our algorithm, Maximal Correlation Embedding Network learns a low dimensional label embedding using an encoder-decoder architecture. It exploits label similarity through a maximal correlation regularization in the embedded label space to reduce the classification bias due to missing labels. A series of experiments on popular multilabel datasets demonstrate that our approach outperforms state of the art, both in complete data and partially observed data. Yang Li 0104, Xiangxiang Xu 0001, Shao-Lun Huang, Lin Zhang 0001 |
ICME | 4 |
| 2019 | An End-to-End Learning Approach for Multimodal Emotion Recognition: Extracting Common and Private InformationabstractMultimodal emotion recognition is important for facilitating efficient interaction between humans and machines. To better detect emotional states from multimodal data, we need to effectively extract both the common information that captures dependencies among different modalities, and the private information that characterizes variations in each modality. However, existing works are mostly designed to pursue either one of these objectives but not both. In our work, we propose an end-to-end learning approach to simultaneously extract the common and private information for multimodal emotion recognition. Specifically, we use a correlation loss based on Hirschfeld-Gebelein-Renyi (HGR) maximal correlation and a reconstruction loss based on autoencoders to preserve the common and private information, respectively. Experimental results on eNTERFACE'05 database and RML database demonstrate the effectiveness of our proposed approach. Fei Ma 0006, Wei Zhang 0185, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001 |
ICME | 4 |
| 2019 | Info-Detection: An Information-Theoretic Approach to Detect Outlier
Fei Ma 0006, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001 |
ICONIP (5) | 4 |
| 2019 | Unsupervised anomaly detection via generative adversarial networks: poster abstractabstractUnsupervised anomaly detection is a fundamental problem in various research areas and application domains, namely the discrimination of abnormal samples from normal samples where training data are only composed of one class (normal) while testing data contains both among which the majority are normal samples. However, previous works can not effectively fit the distribution of high dimensional data and suffers from low AUC scores which measures the classification performance of imbalanced data. To solve these problems, we propose an unsupervised anomaly detection model based on GAN, i.e., UAD-GAN. Specifically, we adopt transfer learning to extract visual features with pre-trained Inception-v3 model and use the discriminator to detect anomalies. UAD-GAN can fit the data distribution and detect anomalies efficiently. Extensive experiments show that UAD-GAN achieves state-of-the-art performance compared to other approaches. Hanling Wang, Fei Ma 0006, Shao-Lun Huang, Lin Zhang 0001 |
IPSN | 4 |
| 2019 | On the Robustness of Noisy ACE Algorithm and Multi-Layer Residual LearningabstractIn this paper, we address the issue of computing the maximal correlation functions for jointly distributed high-dimensional random variables. In such cases, the operations in the alternative conditional expectation (ACE) algorithm can only be implemented by an approximated manner, which is modeled as a variational ACE algorithm with noise. We study the computational behaviors of this algorithm, where the optimal tradeoff between the learning rate, computation accuracy, and convergence rate is investigated. In addition, we establish a connection between the variational ACE algorithm and the residual learning architecture. Our results illustrate interesting interpretations of how multi-layer residual structure benefits function learning. Shao-Lun Huang, Xiangxiang Xu 0001 |
ISIT | 1 |
| 2019 | An Information Theoretic Interpretation to Deep Neural NetworksabstractIt is commonly believed that the hidden layers of deep neural networks (DNNs) attempt to extract informative features for learning tasks. In this paper, we formalize this intuition by showing that the features extracted by DNN coincide with the result of an optimization problem, which we call the "universal feature selection" problem, in a local analysis regime. We interpret the weights training in DNN as the projection of feature functions between feature spaces, specified by the network structure. Our formulation has direct operational meaning in terms of the performance for inference tasks, and gives interpretations to the internal computation results of DNNs. Results of numerical experiments are provided to support the analysis. Shao-Lun Huang, Xiangxiang Xu 0001, Lizhong Zheng, Gregory W. Wornell |
ISIT | 1 |
| 2019 | On The Sample Complexity of HGR Maximal Correlation FunctionsabstractThe Hirschfeld-Gebelein-Rényi (HGR) maximal correlation has been shown useful in many machine learning scenarios. In this paper, we investigate the sample complexity problem of estimating the HGR maximal correlation functions by the alternative conditional expectation (ACE) algorithm from a sequence of training data in the asymptotic regime. Specifically, we develop a mathematical framework to characterize the eigen-decomposition of perturbed matrices, and then establish the error exponent of the learning error for the computed HGR maximal correlation functions. Our result essentially indicates the number of training samples required for estimating the HGR maximal correlation functions to a targeted accuracy by the ACE algorithm. Shao-Lun Huang, Xiangxiang Xu 0001 |
ITW | 1 |
| 2019 | Anomaly detection in surface mount technology process using multi-modal data: poster abstractabstractAnomaly detection is an important area for both research and real-world applications. In the surface mounting technology (SMT) process, the defectives of solder paste printing need to be detected immediately or it may cause great effort for recycling and slow down the whole process. In this paper, we propose a novel model, MM-DNN, for anomaly detection with multi-modal data. We collect a multi-modal dataset from different sensors in the factory. Our method efficiently extracts both predictive features for classification and correlative features between multi-modal data to achieve a higher detection rate. As shown in the experiment, our method can further reduce 77% false alarm rate of the detection result in the factory while keeping 95% of real defectives be correctly detected. Hanling Wang, Yue Zhang 0044, Shao-Lun Huang, Lin Zhang 0001 |
SenSys | 4 |
| 2018 | The Geometric Structure of Generalized Softmax LearningabstractIn this paper, we formulate the generalized softmax learning (GSL) problem, as a symmetric extension of the softmax regression problem. We further study the geometric structure of GSL and demonstrate the equivalence of GSL and the original softmax regression problem. Besides, this geometric structure indicates the symmetry between a neural network and its reverse network, and the symmetric roles of the weights and feature in a neural network. Finally, we present a numerical simulation to verify these symmetry properties in neural networks. Xiangxiang Xu 0001, Shao-Lun Huang, Lizhong Zheng, Lin Zhang 0001 |
ITW | 2 |
| 2018 | Gaussian Universal Features, Canonical Correlations, and Common InformationabstractWe address the problem of optimal feature selection for a Gaussian vector pair in the weak dependence regime, when the inference task is not known in advance. In particular, we show that multiple formulations all yield the same solution, and correspond to the singular value decomposition (SVD) of the canonical correlation matrix. Our results reveal key connections between canonical correlation analysis (CCA), principal component analysis (PCA), the Gaussian information bottleneck, Wyner's common information, and the Ky Fan (nuclear) norms. Shao-Lun Huang, Gregory W. Wornell, Lizhong Zheng |
ITW | 1 |
| 2018 | Joint Mobility Pattern Mining with Urban Region PartitionsabstractMobility pattern mining answers the fundamental question of where people are likely to go from a given location. It plays an important role in city planning, public transport management and location-based mobile applications. Among these applications, many concern the mobility pattern over contiguous spatial regions as a whole. Traditional ways of mobility pattern mining either result in trip clusters with overlapped origin and destination regions, or require an extra step to partition the city into discrete regions, which may not be optimal for mobility pattern extraction. In this paper, we present a region-aware mobility pattern mining framework to jointly extract trip clusters while maintaining non-overlapping partitions of trip origins and destinations. We developed kernelized ACE, a novel extension to a classic algorithm in statistics to compute the optimal mobility clusters under spatial constraints. Experimental results using Beijing taxi trip data show that our approach outperforms other methods with only ~ 0.3% spatial overlap and 86.43% origin-destination correlation. Our case studies on New York City's and Beijing's taxi datasets also yield insightful findings that reveal city-scale mobility patterns and propose potential improvement for public transportation. Jing Lian 0003, Yang Li 0104, Weixi Gu, Shao-Lun Huang, Lin Zhang 0001 |
MobiQuitous | 4 |
| 2018 | Attention-based LSTM-CNNs For Time-series ClassificationabstractTime series classification is a critical problem in the machine learning field, which spawns numerous research works on it. In this work, we propose AttLSTM-CNNs, an attention-based LSTM network and convolution network that jointly extracts the underlying pattern among the time-series for the classification. The attention-based LSTM automatically captures the long-term temporal dependency among the series, and the CNN describes the spatial sparsity and heterogeneity in the data. The extensive experiments show that the proposed model outperforms the other methods for time-series classification. Qianjin Du, Weixi Gu, Lin Zhang 0001, Shao-Lun Huang |
SenSys | 4 |
| 2018 | Speech Emotion Recognition via Attention-based DNN from Multi-Task LearningabstractSpeech unlocks the huge potentials in emotion recognition. High accurate and real-time understanding of human emotion via speech assists Human-Computer Interaction. Previous works are often limited in either coarse-grained emotion learning tasks or the low precisions on the emotion recognition. To solve these problems, we construct a real-world large-scale corpus composed of 4 common emotions (i.e., anger, happiness, neutral and sadness). We also propose a multi-task attention-based DNN model (i.e., MT-A-DNN) on the emotion learning. MT-A-DNN efficiently learns the high-order dependency and non-linear correlations underlying in the audio data. Extensive experiments show that MT-A-DNN outperforms conventional methods on the emotion recognition. It could take one step further on the real-time acoustic emotion recognition in many smart audio-devices. Fei Ma 0006, Weixi Gu, Wei Zhang 0185, Shiguang Ni, Shao-Lun Huang, Lin Zhang 0001 |
SenSys | 5 |
| 2018 | Multimodal Emotion Recognition by extracting common and modality-specific informationabstractEmotion recognition technologies have been widely used in numerous areas including advertising, healthcare and online education. Previous works usually recognize the emotion from either the acoustic or the visual signal, yielding unsatisfied performances and limited applications. To improve the inference capability, we present a multimodal emotion recognition model, EMOdal. Apart from learning the audio and visual data respectively, EMOdal efficiently learns the common and modality-specific information underlying the two kinds of signals, and therefore improves the inference ability. The model has been evaluated on our large-scale emotional data set. The comprehensive evaluations demonstrate that our model outperforms traditional approaches. Wei Zhang 0185, Weixi Gu, Fei Ma 0006, Shiguang Ni, Lin Zhang 0001, Shao-Lun Huang |
SenSys | 6 |
| 2017 | Analysis and evaluation of driving behavior recognition based on a 3-axis accelerometer using a random forest approach: poster abstractabstractUnderstanding human drivers' behavior is critical for the self-driving cars, and has been intensively studied in the past decade. We exploit the widely available camera and motion sensor data from car recorders, and propose a hybrid method of recognizing driving events based on the random forest approach. The classification results are analyzed by comparing different features, classifiers and filters. A high accuracy of 98.1% on driving behavior classification is obtained and the robustness is verified on a dataset including 2400 driving events. Wangjing Cao, Kai Zhang 0012, Yuhan Dong, Shao-Lun Huang, Lin Zhang 0001 |
IPSN | 5 |
| 2017 | Zoning by mobility: evaluating city administrative regions by taxi data: poster abstractabstractThe accelerating urbanization procedure is putting increasing pressure on the management of cities. The administrative zones by which a city is managed are setup based on historical or political reasons, while the dynamics of people is hardly considered in the context. We exploit the widely available mobility data to divide the urban areas into zones by the joint K-mean clustering in origin and destination spaces. The method is evaluated with the New York City and Shenzhen taxi data, and the created zones are compared with the current static zoning plans of the city to evaluate the effectiveness. Liandong Zhou, Shao-Lun Huang, Lin Zhang 0001 |
IPSN | 2 |
| 2017 | An information-theoretic approach to universal feature selection in high-dimensional inferenceabstractWe develop an information theoretic framework for addressing feature selection in applications where the inference task is not specified in advance and the data is from a large alphabet. We introduce a natural notion of universality for such problems, and show that locally optimal solutions are straight forward to obtain, admit natural interpretations via information geometry, have computationally efficient implementations, and represent a practically useful learning methodology. Our development also reveals the key role of Hirschfeld-Gebelein-Renyi maximal correlation and the alternating conditional expectations (ACE) algorithm in such problems. Shao-Lun Huang, Anuran Makur, Lizhong Zheng, Gregory W. Wornell |
ISIT | 1 |
| 2017 | An information-theoretic approach to unsupervised feature selection for high-dimensional dataabstractIn this paper, we model the unsupervised learning of a sequence of observed data vector as a problem of extracting joint patterns among random variables. In particular, we formulate an information-theoretic problem to extract common features of random variables by measuring the loss of total correlation given the feature. This problem can be solved by a local geometric approach, where the solutions can be represented as singular vectors of some matrices related to the pairwise distributions of the data. In addition, we illustrate how these solutions can be transferred to feature functions in machine learning, which can be computed by efficient algorithms from data vectors. Moreover, we present a generalization of the HGR maximal correlation based on these feature functions, which can be viewed as a nonlinear generalization to linear PCA. Finally, the simulation result shows that our extracted feature functions have great performance in real-world problems. Shao-Lun Huang, Lin Zhang 0001, Lizhong Zheng |
ITW | 1 |
| 2016 | Communication theoretic inference on heterogeneous dataabstractStatistical learning has attracted considerable recent research interest due to the wide-ranging demands of big data analytics. The recent introduction of communication theory and information coupling theory into this area suggests a new perspective on statistical learning and inference for data analytics. This paper investigates inference of one data variable from heterogeneous data variables, a problem that plays an increasingly important role in the emerging applications of big data analytics. To generalize the existing conceptual approach, information coupling filtering under hidden data structure or unknown knowledge of interactions among data variables is developed. A least-mean-squares (LMS) filtering approach for non-stationary data similar to an equalizer is suggested, while the training data gives the depth of the filter analogously to model selection in learning theory. The information combining in diversity communication is extended to fuse more data variables for even greater precision of inference. Extending from multiuser detection, an algorithm based on Multiple Signal Classification (MUSIC) is demonstrated to identify useful data variables for inference, as a novel solution to knowledge discovery. A series of examples illustrate the effectiveness of this framework, suggesting that statistical communication theory and statistical signal processing can substantially contribute to statistical learning theory. Kwang-Cheng Chen, Baturalp Mankir, Shao-Lun Huang, Lizhong Zheng, H. Vincent Poor |
ICC | 3 |
| 2016 | Data extraction via histogram and arithmetic mean queries: Fundamental limits and algorithmsabstractThe problems of extracting information from a data set via histogram queries or arithmetic mean queries are considered. We first show that the fundamental limit on the number of histogram queries, m, so that the entire data set of size n can be extracted losslessly, is m = Θ(n/log n), sub-linear in the size of the data set. For proving the lower bound (converse), we use standard arguments based on simple counting. For proving the upper bound (achievability), we proposed two query mechanisms. The first mechanism is random sampling, where in each query, the items to be included in the queried subset are uniformly randomly selected. With random sampling, it is shown that the entire data set can be extracted with vanishing error probability using Ω(n/log n) queries. The second one is a non-adaptive deterministic algorithm. With this algorithm, it is shown that the entire data set can be extracted exactly (no error) using Ω(n/log n) queries. We then extend the results to arithmetic mean queries, and show that for data sets taking values in a real-valued finite arithmetic progression, the fundamental limit on the number of arithmetic mean queries to extract the entire data set is also Θ(n/log n). I-Hsiang Wang, Shao-Lun Huang, Kuan-Yun Lee, Kwang-Cheng Chen |
ISIT | 2 |
| 2015 | Information cascades in social networks via dynamic system analysesabstractSystematically analyzing the dynamic behaviors of social networks is one of the central topic in understanding the structure of large networks. In particular, the information cascade [1] introduced by Banerjee provides great insights in characterizing the opinion exchanging between network agents. Traditionally studies of information cascades focus on the Bayesian models, which are often difficult to model real world situations. In this paper, we attempt to study the information cascades from a non-Bayesian point of view. In particular, we consider a sequential decision model but with an arbitrary decision rule. We show that the fraction of agents in a network making any specific decision will converge. Thus, the agents in the network reach a sort of consensus with high probability, which allows us to predict the herd behaviors. In addition, we also apply our non-Bayesian model to different network structures, such as ER model and network with communities, in which the affect of information cascades are quantified. Finally, we simulate the decision process for multiple communities, which justifies our proposed model to comprehend real world complex user behaviors and dynamics. Shao-Lun Huang, Kwang-Cheng Chen |
ICC | 1 |
| 2015 | On locally decodable source codingabstractWith the boom of big data, traditional source coding techniques face the common obstacle to decode only a small portion of information efficiently. In this paper, we aim to resolve this difficulty by introducing a specific type of source coding scheme called locally decodable source coding (LDSC). Rigorously, LDSC is capable of recovering an arbitrary bit of the unencoded message from its encoded version, by only feeding a small number of the encoded message to the decoder, and we call the decoder t-local if only t encoded symbols are required.We consider both almost lossless (block error) and lossy (bit error) cases for LDSC. First, we show that using linear encoder and a decoder with bounded locality, the reliable compress rate can not be less than one. More importantly, we show that even with a general encoder and 2-local decoders (t = 2), the rate of LDSC is still one. On the contrary, the achievability bounds for almost lossless and lossy compressions with excess distortion suggest that optimal compression rate is achievable when O(log n) encoded symbols is queried by the decoder with block-length n. We also show that, rate distortion is achievable when the number of queries is scaled over n with a bound on the rate in finite-length regime. Although the achievability bounds are simply based on the concatenation of code blocks, they outperform the existing bounds in succinct data structures literature. Ali Makhdoumi, Shao-Lun Huang, Muriel Médard, Yury Polyanskiy |
ICC | 2 |
| 2015 | A spectrum decomposition to the feature spaces and the application to big data analyticsabstractIn this paper, we investigate how to efficiently extract informative features of high-dimensional data through noisy channels. Specifically, we decompose the feature space of the data into a sequence of score functions with decreasing information volumes, such that different scores are uncorrelated. From this decomposition, the features of the data become a sequence of score functions such that the most informative lowdimensional feature can be selected as the first few scores. This greatly simplifies the feature selection problem. In addition, we apply this spectrum decomposition to data with high-dimensional structures, i.e., the hidden Markov model (HMM). We show that in HMM, it is desirable to consider a particular class of score functions called as the node scores, which allows us to efficiently extract informative features of the hidden variables by applying the spectrum decomposition approach. Finally, we develop efficient algorithms to extract such features from node scores, and present an example to illustrate the performance of the node scores. Shao-Lun Huang, Lizhong Zheng |
ISIT | 1 |
| 2015 | Communication Theoretic Data AnalyticsabstractWidespread use of the Internet and social networks invokes the generation of big data, which is proving to be useful in a number of applications. To deal with explosively growing amounts of data, data analytics has emerged as a critical technology related to computing, signal processing, and information networking. In this paper, a formalism is considered in which data are modeled as a generalized social network and communication theory and information theory are thereby extended to data analytics. First, the creation of an equalizer to optimize information transfer between two data variables is considered, and financial data are used to demonstrate the advantages of this approach. Then, an information coupling approach based on information geometry is applied for dimensionality reduction, with a pattern recognition example to illustrate the effectiveness of this formalism. These initial trials suggest the potential of communication theoretic data analytics for a wide range of applications. Kwang-Cheng Chen, Shao-Lun Huang, Lizhong Zheng, H. Vincent Poor |
IEEE J. Sel. Areas Commun. | 2 |
| 2015 | Euclidean Information Theory of NetworksabstractIn this paper, we extend the information theoretic framework that was developed in earlier works to multi-hop network settings. For a given network, we construct a novel deterministic model that quantifies the ability of the network in transmitting private and common messages across users. Based on this model, we formulate a linear optimization problem that explores the throughput of a multi-layer network, thereby offering the optimal strategy as to what kind of common messages should be generated in the network to maximize the throughput. With this deterministic model, we also investigate the role of feedback for multi-layer networks, from which we identify a variety of scenarios in which feedback can improve transmission efficiency. Our results provide fundamental guidelines as to how to coordinate cooperation between users to enable efficient information exchanges across them. Shao-Lun Huang, Changho Suh, Lizhong Zheng |
IEEE Trans. Inf. Theory | 1 |
| 2014 | Green Traffic Compression in Wireless Sensor NetworksabstractEmerging multi-hop machine-to-machine (M2M) communications that likely support a large number of wireless devices create new challenges for spectrum scarcity and energy efficiency. In parallel to pursuing physical layer transmission efficiency, traffic compression to reduce required wireless transmissions suggests a new paradigm of wireless networks. Utilizing the natures of broadcasting and information collection in wireless sensor or machine networks, cognitive traffic compression can be facilitated by our proposed optimal fusion rules and topology compression algorithm. Therefore, only the necessary and connected sensors/machines in M2M networks are required to transmit, to achieve the desirable distortion of information collection (i.e. detection/estimation error). In other words, given the desirable distortion, the number of sensors to transmit, or equivalently the total energy consumption, serves our purpose of energy efficiency for end-to-end networking. Numerical results show successful compression of total network traffic to significantly enhance networking energy efficiency. Kang-Hao Peng, Kwang-Cheng Chen, Shao-Lun Huang, Shao-Chou Hung, Xinhao Cheng |
VTC Spring | 3 |
| 2013 | Euclidean information theory of networksabstractIn this paper, we extend the information theoretical framework that was developed in [1] to multi-hop communication networks. For a given network, we construct a deterministic model that models the ability of the channels in transmitting private and common messages between users in this network. Based on this model, we formulate a linear optimization problem to study the network throughput, where the solution indicates what kind of common messages should be generated in a network to optimize the throughput. Our results provide fundamental guidelines of how users in a network should cooperate with each other to communicate efficiently. Shao-Lun Huang, Changho Suh, Lizhong Zheng |
ISIT | 1 |
| 2013 | Match and Replace: A Functional ECO Engine for Multierror Circuit RectificationabstractFunctional engineering change order (ECO) is a popular technique for rectifying design errors after synthesis and placement stages. We present a new approach to generating the patch circuits for multierror circuit rectification. In this paper, we propose a two-phase approach of: 1) discovering the functional matches in two circuits followed by 2) determining the final patch circuits from the matches. The ECO engine in this paper discovers functional and structural matches in two circuits by coordinating the SAT-sweeping and the cut-matching algorithms. Then, the patch selection is conducted by the combinational equivalence checking technique and a linear-time selection heuristic. The experimental results on public benchmark and industrial circuits demonstrate that this ECO engine outperforms state-of-the-art interpolation-based engines. Shao-Lun Huang, Wei-Hsun Lin, Po-Kai Huang, Chung-Yang Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Proof of the Outage Probability Conjecture for MISO ChannelsabstractIt is conjectured that the covariance matrices minimizing the outage probability under a power constraint for multiple-input multiple-output channels with Gaussian fading are diagonal with either zeros or constant values on the diagonal. In the multiple-input single-output (MISO) setting, this is equivalent to conjecture that the Gaussian quadratic forms having largest tail probability correspond to such diagonal matrices. This paper provides a proof of the conjecture in this MISO setting. Emmanuel Abbe, Shao-Lun Huang, Emre Telatar |
IEEE Trans. Inf. Theory | 2 |
| 2012 | Linear information coupling problemsabstractMany network information theory problems face the similar difficulty of single letterization. We argue that this is due to the lack of a geometric structure on the space of probability distribution. In this paper, we develop such a structure by assuming that the distributions of interest are close to each other. Under this assumption, the K-L divergence is reduced to the squared Euclidean metric in an Euclidean space. Moreover, we construct the notion of coordinate and inner product, which will facilitate solving communication problems. We will also present the application of this approach to the point-to-point channel and the general broadcast channel, which demonstrates how our technique simplifies information theory problems. Shao-Lun Huang, Lizhong Zheng |
ISIT | 1 |
| 2012 | Proof of the outage probability conjecture for MISO channelsabstractIt is conjectured in [6] that the covariance matrices minimizing the outage probability under a power constraint for MIMO channels with Gaussian fading are diagonal with either zeros or constant values on the diagonal. In the MISO setting, this is equivalent to conjecture that the Gaussian quadratic forms having largest tail probability correspond to such diagonal matrices. This paper provides a proof of the conjecture in this MISO setting. Emmanuel Abbe, Shao-Lun Huang, Emre Telatar |
ITW | 2 |
| 2011 | A robust ECO engine by resource-constraint-aware technology mapping and incremental routing optimizationabstractECO re-mapping is a key step in functional ECO tools. It implements a given patch function on a layout database with a limited spare cell resource. Previous ECO re-mapping algorithms are based on existing technology mappers. However, these mappers are not designed to consider the resource limitation and thus the corresponding ECO results are generally not good enough, or even become much worse when the spare cells are sparse. In this paper, we proposed a new solution for ECO remapping. It includes a robust resource-constraint-aware technology mapper and a fast incremental router for wire-length optimization. Moreover, we adopt a Pseudo-Boolean solver to search feasible solutions when the spare cells are sparse. Our experimental results show that our ECO engine can outperform the previous tool in both runtime and routing costs. We also demonstrate the robustness of our tool by performing ECOs on various spare cell limitations. Shao-Lun Huang, Chi-An Wu, Kai-Fu Tang, Chang-Hong Hsu, Chung-Yang Huang |
ASP-DAC | 1 |
| 2011 | Match and replace - A functional ECO engine for multi-error circuit rectificationabstractFunctional ECO has been an indispensible technique in modern VLSI design flow. This paper proposes an ECO engine in a two-phase approach: a matching phase for rectification pair identification and a replacement phase for pair selection and patch minimization. The rectification pair identification algorithm explores rectification pairs between the original circuit and the golden circuit. The rectification pair selector determines final patches through a linear-time heuristic. A gate-recycle process performs patch minimization for final refinement. The experiments show that this ECO engine outperforms a state-of-the-art interpolation-based engine in both patch quality and runtime. Shao-Lun Huang, Wei-Hsun Lin, Chung-Yang Huang |
ICCAD | 1 |
| 2009 | SAT-controlled redundancy addition and removal: a novel circuit restructuring techniqueabstractWe proposed a novel Boolean Satisfiability (SAT)-controlled redundancy addition and removal (RAR) algorithm to resolve the performance and quality problems of the previous RAR approaches. With the introduction of modern SAT techniques, such as efficient Boolean constraint propagation (BCP), conflict-driven learning, and flexible decision procedure, our RAR engine can identify 10x more alternative wires/gates while achieving 70% reduction in runtime. Chi-An Wu, Ting-Hao Lin, Shao-Lun Huang, Chung-Yang Huang |
ASP-DAC | 3 |
| 2009 | Interpolant generation without constructing resolution graphabstractIn this paper, we proposed a novel interpolant generation algorithm without constructing the resolution graph of the unsatisfiability proof. Our algorithm generates the interpolant by building sub-interpolants from conflict analyses and then merges them based on the last decision conflict. The experimental results show that our algorithm has the advantages over the prior interpolant generation techniques in both memory usage and interpolation circuit size. Chih-Jen Hsu, Shao-Lun Huang, Chi-An Wu, Chung-Yang Huang |
ICCAD | 2 |