VLDB 2026 Research / reviewers in the wild / expert
Zhitang Chen
dblp:06/10875
· DBLP profile ↗
48ranked-venue papers
8as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Systems, architecture and hardware · 7 · 7 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 1 since 2021Computer networks · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time AdaptationabstractDiffusion-based Neural Combinatorial Optimization (NCO) has demonstrated effectiveness in solving NP-complete (NPC) problems by learning discrete diffusion models for solution generation, eliminating hand-crafted domain knowledge. Despite their success, existing NCO methods face significant challenges in both cross-scale and cross-problem generalization, and high training costs compared to traditional solvers. While recent studies on diffusion models have introduced training-free guidance approaches that leverage pre-defined guidance functions for conditional generation, such methodologies have not been extensively explored in combinatorial optimization. To bridge this gap, we propose a training-free inference time adaptation framework (DIFU-Ada) that enables both the zero-shot cross-problem transfer and cross-scale generalization capabilities of diffusion-based NCO solvers without requiring additional training. We provide theoretical analysis that helps understanding the cross-problem transfer capability. Our experimental results demonstrate that a diffusion solver, trained exclusively on the Traveling Salesman Problem (TSP), can achieve competitive zero-shot transfer performance across different problem scales on TSP variants, such as Prize Collecting TSP (PCTSP) and the Orienteering Problem (OP), through inference time adaptation. Haoyu Lei, Kaiwen Zhou 0001, Yinchuan Li, Zhitang Chen, Farzan Farnia |
AAAI | 4 |
| 2026 | A²Flow: Automating Agentic Workflow Generation via Self-Adaptive Abstraction OperatorsabstractLarge language models (LLMs) have shown strong potential in automating the design of agentic workflows. However, existing methods still rely heavily on manually predefined operators, limiting generalization and scalability. To address this issue, we propose A²Flow, a fully automated framework for agentic workflow generation based on self-adaptive abstraction operators. A²Flow employs a three-stage operator extraction process: 1) Case-based Initial Operator Generation: leveraging expert demonstrations and LLM reasoning to generate case-specific operators; 2) Operator Clustering and Preliminary Abstraction: grouping similar operators across tasks to form preliminary abstractions; and 3) Deep Extraction for Abstract Execution Operators: applying long chain-of-thought prompting and multi-path reasoning to derive compact and generalizable execution operators. These operators serve as reusable building blocks for workflow construction without manual predefinition. Furthermore, we enhance node-level workflow search with an operator memory mechanism, which retains historical outputs to enrich context and improve decision-making. Experiments on general and embodied benchmarks show that A²Flow achieves a 2.4% and 19.3% average performance improvement and reduces resource usage by 37% over state-of-the-art baselines. Xiaokang Wei, Yuanqi Shao, Kaiwen Zhou 0001, Siwei Rao, Junhui Zhan, Zhitang Chen |
AAAI | 8 |
| 2025 | Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPOabstractDirect alignment methods typically train large language models (LLMs) by contrasting the likelihoods of preferred and dispreferred responses. While effective at capturing relative preferences, these methods are widely observed to suppress the absolute likelihoods of example responses. As a result, aligned models can deviate from expected patterns, exhibiting reward‑hacking effect even without an explicit reward model. This fundamental limitation of contrastive alignment, termed likelihood underdetermination, motivates us to revisit direct preference optimization (DPO)—the seminal direct alignment method. Interestingly, we show that the DPO loss admits a principled decomposition. The reformulated loss not only extends naturally to a broader range of feedback types, but also unveils the root cause of likelihood underdetermination. Specifically, we identify that standard DPO implicitly oversimplifies a regularizer in the reformulated loss; restoring this full term effectively resolves the underdetermination. Building on these insights, we introduce PRoximalized PReference Optimization (PRO), a unified alignment method that accommodates diverse feedback types while eliminating likelihood underdetermination through an efficient approximation of the full regularizer. Empirical evaluations demonstrate the consistent superiority of PRO over existing methods across pairwise, binary and scalar feedback. Kaiyang Guo, Yinchuan Li, Zhitang Chen |
NeurIPS | 3 |
| 2024 | Sampling is as easy as keeping the consistency: convergence guarantee for Consistency ModelsabstractWe provide the first convergence guarantee for the Consistency Models (CMs), a newly emerging type of one-step generative models that is capable of generating comparable samples to those sampled from state-of-the-art Diffusion Models. Our main result is that, under the basic assumptions on score-matching errors, consistency errors, and smoothness of the data distribution, CMs can efficiently generate samples in one step with small $W_2$ error to any real data distribution. Our results (1) hold for $L^2$-accurate assumptions on both score and consistency functions (rather than $L^\infty$-accurate assumptions); (2) do not require strong assumptions on the data distribution such as log-Sobelev conditions; (3) scale polynomially in all parameters; and (4) match the state-of-the-art convergence guarantee for score-based generative models. We also show that the Multi-step Consistency Sampling procedure can further reduce the error comparing to one step sampling, which supports the original statement from Song Yang’s work. Our result can be generalized to arbitrary bounded data distributions that may be supported on some low-dimensional sub-manifolds. Our results further imply TV error guarantees when making some Langevin-based modifications to the output distributions. Junlong Lyu, Zhitang Chen, Shoubo Feng |
ICML | 2 |
| 2024 | On Low-Rank Directed Acyclic Graphs and Causal Structure LearningabstractDespite several advances in recent years, learning causal structures represented by directed acyclic graphs (DAGs) remains a challenging task in high-dimensional settings when the graphs to be learned are not sparse. In this article, we propose to exploit a low-rank assumption regarding the (weighted) adjacency matrix of a DAG causal model to help address this problem. We utilize existing low-rank techniques to adapt causal structure learning methods to take advantage of this assumption and establish several useful results relating interpretable graphical conditions to the low-rank assumption. Specifically, we show that the maximum rank is highly related to hubs, suggesting that scale-free (SF) networks, which are frequently encountered in practice, tend to be low rank. Our experiments demonstrate the utility of the low-rank adaptations for a variety of data models, especially with relatively large and dense graphs. Moreover, with a validation procedure, the adaptations maintain a superior or comparable performance even when graphs are not restricted to be low rank. Zhuangyan Fang, Shengyu Zhu 0001, Jiji Zhang, Zhitang Chen, Yangbo He |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Neighbor Auto-Grouping Graph Neural Networks for Handover Parameter Configuration in Cellular NetworkabstractThe mobile communication enabled by cellular networks is the one of the main foundations of our modern society. Optimizing the performance of cellular networks and providing massive connectivity with improved coverage and user experience has a considerable social and economic impact on our daily life. This performance relies heavily on the configuration of the network parameters. However, with the massive increase in both the size and complexity of cellular networks, network management, especially parameter configuration, is becoming complicated. The current practice, which relies largely on experts' prior knowledge, is not adequate and will require lots of domain experts and high maintenance costs. In this work, we propose a learning-based framework for handover parameter configuration. The key challenge, in this case, is to tackle the complicated dependencies between neighboring cells and jointly optimize the whole network. Our framework addresses this challenge in two ways. First, we introduce a novel approach to imitate how the network responds to different network states and parameter values, called auto-grouping graph convolutional network (AG-GCN). During the parameter configuration stage, instead of solving the global optimization problem, we design a local multi-objective optimization strategy where each cell considers several local performance metrics to balance its own performance and its neighbors. We evaluate our proposed algorithm via a simulator constructed using real network data. We demonstrate that the handover parameters our model can find, achieve better average network throughput compared to those recommended by experts as well as alternative baselines, which can bring better network quality and stability. It has the potential to massively reduce costs arising from human expert intervention and maintenance. Mehrtash Mehrabi, Walid Masoudimansour, Yingxue Zhang 0001, Jie Chuai, Zhitang Chen, Mark Coates, Jianye Hao, Yanhui Geng |
AAAI | 5 |
| 2023 | A Novel Extrapolation Technique to Accelerate WMMSEabstractPrecoding design is essential for massive multi-user multiple-input multiple-output (MU-MIMO) systems, which aims at maximizing the weighted sum-rate (WSR). This problem is known to be NP-hard, and iterative algorithms are typically used to approximately solve it. The weighted minimum mean-squared error (WMMSE) algorithm is a popular solver for WSR maximization, which efficiently finds a local maxima of WSR. In this work, we introduce a novel extrapolation technique to further accelerate WMMSE. This technique is inspired by the momentum technique in convex optimization, and can be interpreted as an accelerated second-order method. The merits of the proposed extrapolation technique are (i) lightweight, as it almost does not increase the iteration complexity, (ii) generic, since it works in various settings such as the sum power constraint or per-antenna power constraint cases and coordinated multi-point joint transmission networks, and (iii) effective, that our simulation results show it significantly accelerates the convergence of WMMSE in the high channel correlation regime. Kaiwen Zhou 0001, Guochen Liu, Zhitang Chen |
ICASSP | 4 |
| 2023 | FastGR: Global Routing on CPU-GPU with Heterogeneous Task Graph Scheduler (Extended Abstract)abstractRunning time is a key metric across the standard physical design flow stages. However, with the rapid growth in design sizes, routing runtime has become the runtime bottleneck in the physical design flow. To improve the effectiveness of the modern global router, we propose a global routing framework with GPU-accelerated routing algorithms and a heterogeneous task graph scheduler, called FastGR. Its runtime-oriented version FastGRL achieves 2.489× speedup compared with the state-of-the-art global router. Furthermore, the GPU-accelerated L-shape pattern routing used in FastGRL can contribute to 9.324× speedup over the sequential algorithm on CPU. Its quality-oriented version FastGRH offers further quality improvement over FastGRL with similar acceleration. Siting Liu 0002, Yuan Pu 0001, Peiyu Liao, Hongzhong Wu, Rui Zhang 0040, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001 |
IJCAI | 6 |
| 2023 | Efficient Robust Bayesian Optimization for Arbitrary Uncertain inputsabstractBayesian Optimization (BO) is a sample-efficient optimization algorithm widely employed across various applications. In some challenging BO tasks, input uncertainty arises due to the inevitable randomness in the optimization process, such as machining errors, execution noise, or contextual variability. This uncertainty deviates the input from the intended value before evaluation, resulting in significant performance fluctuations in the final result. In this paper, we introduce a novel robust Bayesian Optimization algorithm, AIRBO, which can effectively identify a robust optimum that performs consistently well under arbitrary input uncertainty. Our method directly models the uncertain inputs of arbitrary distributions by empowering the Gaussian Process with the Maximum Mean Discrepancy (MMD) and further accelerates the posterior inference via Nystrom approximation. Rigorous theoretical regret bound is established under MMD estimation error and extensive experiments on synthetic functions and real problems demonstrate that our approach can handle various input uncertainties and achieve a state-of-the-art performance. Junlong Lyu, Wenlong Lyu, Zhitang Chen |
NeurIPS | 4 |
| 2023 | A Unified Framework for Layout Pattern Analysis With Deep Causal EstimationabstractThe decrease of feature size and the growing complexity of the fabrication process lead to more failures in manufacturing semiconductor devices. Therefore, identifying the root cause layout patterns of failures becomes increasingly crucial for yield improvement. In this article, a novel layout-aware diagnosis-based layout pattern analysis framework is proposed to identify the root cause efficiently. At the first stage of the framework, an encoder network trained using contrastive learning is used to extract representations of layout snippets that are invariant to trivial transformations, including shift, rotation, and mirroring, which are then clustered to form layout patterns. At the second stage, we model the causal relationship between any potential root cause layout patterns and the systematic defects by a structural causal model, which is then used to estimate the average causal effect (ACE) of candidate layout patterns on the systematic defect to identify the true root cause. Experimental results on real industrial cases demonstrate that our framework outperforms a commercial tool with higher accuracies and around$\times 8.4$speedup on average. Ran Chen 0001, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Bei Yu 0001, Pengyun Li, Yu Huang 0005, Jianye Hao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | FastGR: Global Routing on CPU-GPU With Heterogeneous Task Graph SchedulerabstractRunning time is a key metric across the standard physical design flow stages. However, with the rapid growth in design sizes, routing runtime has become the runtime bottleneck in the physical design flow. As a result, speeding routing becomes a critical and pressing task for IC design automation. Aside from the running time, we need to evaluate the quality of the global routing solution since a poor global routing engine degrades the solution performance after the entire routing stage. This work takes both of them into consideration. We propose a global routing framework with GPU-accelerated routing algorithms and a heterogeneous task graph scheduler, called FastGR, to accelerate the procedure of the modern global router and improve its effectiveness. Its runtime-oriented version$\text {FastGR}^{\text {L}}$achieves$2.489\times $speedup compared with the state-of-the-art global router. Furthermore, the GPU-accelerated L-shape pattern routing algorithm used in$\text {FastGR}^{\text {L}}$can contribute to$9.324\times $speedup over the sequential algorithm on CPU. Its quality-oriented version$\text {FastGR}^{\text {H}}$offers a 27.855% improvement of the number of shorts over the runtime-oriented version and still gets$1.970\times $faster than the most advanced global router. Siting Liu 0002, Yuan Pu 0001, Peiyu Liao, Hongzhong Wu, Rui Zhang 0040, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Contrastive-ACE: Domain Generalization Through Alignment of Causal MechanismsabstractDomain generalization aims to learn knowledge invariant across different distributions while semantically meaningful for downstream tasks from multiple source domains, to improve the model's generalization ability on unseen target domains. The fundamental objective is to understand the underlying "invariance" behind these observational distributions and such invariance has been shown to have a close connection to causality. While many existing approaches make use of the property that causal features are invariant across domains, we consider the invariance of the average causal effect of the features to the labels. This invariance regularizes our training approach in which interventions are performed on features to enforce stability of the causal prediction by the classifier across domains. Our work thus sheds some light on the domain generalization problem by introducing invariance of the mechanisms into the learning process. Experiments on several benchmark datasets demonstrate the performance of the proposed method against SOTAs. The codes are available at: https://github.com/lithostark/Contrastive-ACE. Furui Liu, Zhitang Chen, Yik-Chung Wu, Jianye Hao, Guangyong Chen, Pheng-Ann Heng |
IEEE Trans. Image Process. | 3 |
| 2022 | Out-of-distribution Generalization with Causal Invariant TransformationsabstractIn real-world applications, it is important and desirable to learn a model that performs well on out-of-distribution (OOD) data. Recently, causality has become a powerful tool to tackle the OOD generalization problem, with the idea resting on the causal mechanism that is invariant across domains of interest. To leverage the generally unknown causal mechanism, existing works assume a linear form of causal feature or require sufficiently many and diverse training domains, which are usually restrictive in practice. In this work, we obviate these assumptions and tackle the OOD problem without explicitly recovering the causal feature. Our approach is based on transformations that modify the non-causal feature but leave the causal part unchanged, which can be either obtained from prior knowledge or learned from the training data in the multi-domain scenario. Under the setting of invariant causal mechanism, we theoretically show that if all such transformations are available, then we can learn a minimax optimal model across the domains using only single domain data. Noticing that knowing a complete set of these causal invariant transformations may be impractical, we further show that it suffices to know only a subset of these transformations. Based on the theoretical findings, a regularized training procedure is proposed to improve the OOD generalization capability. Extensive experimental results on both synthetic and real datasets verify the effectiveness of the proposed algorithm, even with only a few causal invariant transformations. Ruoyu Wang 0016, Mingyang Yi, Zhitang Chen, Shengyu Zhu 0001 |
CVPR | 3 |
| 2022 | DREAMPlace 4.0: Timing-driven Global Placement with Momentum-based Net WeightingabstractTiming optimization is critical to integrated circuit (IC) design closure. Existing global placement algorithms mostly focus on wirelength optimization without considering timing. In this paper, we propose a timing-driven global placement algorithm leveraging a momentum-based net weighting strategy. Besides, we improve the preconditioner to incorporate our net weighting scheme. Experimental results on ICCAD 2015 contest benchmarks demonstrate that our algorithm can significantly improve total negative slack (TNS) and meanwhile be beneficial to worse negative slack (WNS). Peiyu Liao, Siting Liu 0002, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001 |
DATE | 3 |
| 2022 | FastGR: Global Routing on CPU-GPU with Heterogeneous Task Graph SchedulerabstractRouting is an essential step to integrated circuits (IC) design closure. With the rapid increase of design scales, routing has become the runtime bottleneck in the physical design flow. Thus, accelerating routing becomes a vital and urgent task for IC design automation. This paper proposes a global routing framework running on hybrid CPU-GPU platforms with a heterogeneous task scheduler and a GPU-accelerated pattern routing algorithm. We demonstrate that the task scheduler can lead to 2.307 × speedup compared with the widely-adopted batch-based parallelization strategy on CPU and the GPU-accelerated pattern routing algorithm can contribute to 10.877 × speedup over the sequential algorithm on CPU. Finally, the combined techniques can achieve 2.426 × speedup without quality degradation compared with the state-of-the-art global router. Siting Liu 0002, Peiyu Liao, Rui Zhang 0040, Zhitang Chen, Wenlong Lv, Yibo Lin, Bei Yu 0001 |
DATE | 4 |
| 2022 | Batch Sequential Black-Box Optimization with Embedding Alignment Cells for Logic SynthesisabstractDuring the logic synthesis flow of EDA, a sequence of graph transformation operators are applied to the circuits so that the Quality of Results (QoR) of the circuits highly depends on the chosen operators and their specific parameters in the sequence, making the search space operator-dependent and increasingly exponential. In this paper, we formulate the logic synthesis design space exploration as a conditional sequence optimization problem, where at each transformation step, an optimization operator is selected and its corresponding parameters are decided. To solve this problem, we propose a novel sequential black-box optimization approach without human intervention: 1) Due to the conditional and sequential structure of operator sequence with variable length, we build an embedding alignment cells based recurrent neural network as a surrogate model to estimate the QoR of the logic synthesis flow with historical data. 2) With the surrogate model, we construct acquisition function to balance exploration and exploitation with respect to each metric of the QoR. 3) We use multi-objective optimization algorithm to find the Pareto front of the acquisition functions, along which a batch of sequences, consisting of parameterized operators, are (randomly) selected to users for evaluation under the budget of computing resource. We repeat the above three steps until convergence or time limit. Experimental results on public EPFL benchmarks demonstrate the superiority of our approach over the expert-crafted optimization flows and other machine learning based methods. Compared to resyn2, we achieve 11.8% LUT-6 count descent improvements without sacrificing level values. Chang Feng, Wenlong Lyu, Zhitang Chen, Junjie Ye 0002, Mingxuan Yuan, Jianye Hao |
ICCAD | 3 |
| 2022 | RCANet: Root Cause Analysis via Latent Variable Interaction Modeling for Yield ImprovementabstractIdentifying root causes of systematic defects is a crucial step in yield enhancement process of integrated circuit (IC) manufacturing. With increasing complexity of fabrication processes and decreasing sizes of pattern features, more systematic defects occur at advanced technology nodes, and traditional methods are unfeasible to directly identify failure causes, due to expensive time and labor costs. Root cause analysis (RCA) technology is thus studied to automatically identify common root causes in a short time. In this paper, we develop RCANet, an end-to-end unsupervised learning-based RCA framework, which analyses diagnosis reports of failing dies within a wafer and identifies both layout-aware and cell-internal root causes efficiently. Experimental results on designs with different technologies demonstrate that RCANet outperforms both a commercial tool and the state-of-the-art method. Xiaopeng Zhang 0009, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Evangeline F. Y. Young, Pengyun Li, Yu Huang 0005, Jianye Hao |
ITC | 3 |
| 2022 | Para-CFlows: $C^k$-universal diffeomorphism approximators as superior neural surrogatesabstractInvertible neural networks based on Coupling Flows (CFlows) have various applications such as image synthesis and data compression. The approximation universality for CFlows is of paramount importance to ensure the model expressiveness. In this paper, we prove that CFlows}can approximate any diffeomorphism in $C^k$-norm if its layers can approximate certain single-coordinate transforms. Specifically, we derive that a composition of affine coupling layers and invertible linear transforms achieves this universality. Furthermore, in parametric cases where the diffeomorphism depends on some extra parameters, we prove the corresponding approximation theorems for parametric coupling flows named Para-CFlows. In practice, we apply Para-CFlows as a neural surrogate model in contextual Bayesian optimization tasks, to demonstrate its superiority over other neural surrogate models in terms of optimization performance and gradient approximations. Junlong Lyu, Zhitang Chen, Chang Feng, Wenjing Cun, Shengyu Zhu 0001, Yanhui Geng, Zhijie Xu, Chen Yongwei |
NeurIPS | 2 |
| 2022 | Masked Gradient-Based Causal Structure LearningabstractThis paper studies the problem of learning causal structures from observational data. We reformulate the Structural Equation Model (SEM) with additive noises in a form parameterized by binary graph adjacency matrix and show that, if the original SEM is identifiable, then the binary adjacency matrix can be identified up to super-graphs of the true causal graph under mild conditions. We then utilize the reformulated SEM to develop a causal structure learning method that can be efficiently trained using gradient-based optimization, by leveraging a smooth characterization on acyclicity and the Gumbel-Softmax approach to approximate the binary adjacency matrix. It is found that the obtained entries are typically near zero or one and can be easily thresholded to identify the edges. We conduct experiments on synthetic and real datasets to validate the effectiveness of the proposed method, and show that it readily includes different smooth model functions and achieves a much improved performance on most datasets considered. Ignavier Ng, Shengyu Zhu 0001, Zhuangyan Fang, Haoyang Li 0002, Zhitang Chen, Jun Wang 0012 |
SDM | 5 |
| 2022 | Reframed GES with a neural conditional dependence measureabstractIn a nonparametric setting, the causal structure is often identifiable only up to Markov equivalence, and for the purpose of causal inference, it is useful to learn a graphical representation of the Markov equivalence class (MEC). In this paper, we revisit the Greedy Equivalence Search (GES) algorithm, which is widely cited as a score-based algorithm for learning the MEC of the underlying causal structure. We observe that in order to make the GES algorithm consistent in a nonparametric setting, it is not necessary to design a scoring metric that evaluates graphs. Instead, it suffices to plug in a consistent estimator of a measure of conditional dependence to guide the search. We therefore present a reframing of the GES algorithm, which is more flexible than the standard score-based version and readily lends itself to the nonparametric setting with a general measure of conditional dependence. In addition, we propose a neural conditional dependence (NCD) measure, which utilizes the expressive power of deep neural networks to characterize conditional independence in a nonparametric manner. We establish the optimality of the reframed GES algorithm under standard assumptions and the consistency of using our NCD estimator to decide conditional independence. Together these results justify the proposed approach. Experimental results demonstrate the effectiveness of our method in causal discovery, as well as the advantages of using our NCD measure over kernel-based measures. Xinwei Shen 0002, Shengyu Zhu 0001, Jiji Zhang, Shoubo Hu, Zhitang Chen |
UAI | 5 |
| 2022 | Weakly Supervised Disentangled Generative Causal Representation LearningabstractThis paper proposes a Disentangled gEnerative cAusal Representation (DEAR) learning method under appropriate supervised information. Unlike existing disentanglement methods that enforce independence of the latent variables, we consider the general case where the underlying factors of interests can be causally related. We show that previous methods with independent priors fail to disentangle causally related factors even under supervision. Motivated by this finding, we propose a new disentangled learning method called DEAR that enables causal controllable generation and causal representation learning. The key ingredient of this new formulation is to use a structural causal model (SCM) as the prior distribution for a bidirectional generative model. The prior is then trained jointly with a generator and an encoder using a suitable GAN algorithm incorporated with supervised information on the ground-truth factors and their underlying causal structure. We provide theoretical justification on the identifiability and asymptotic convergence of the proposed method. We conduct extensive experiments on both synthesized and real data sets to demonstrate the effectiveness of DEAR in causal controllable generation, and the benefits of the learned representations for downstream tasks in terms of sample efficiency and distributional robustness. Xinwei Shen 0002, Furui Liu, Hanze Dong, Qing Lian, Zhitang Chen, Tong Zhang 0001 |
J. Mach. Learn. Res. | 5 |
| 2021 | CausalVAE: Disentangled Representation Learning via Neural Structural Causal ModelsabstractLearning disentanglement aims at finding a low dimensional representation which consists of multiple explanatory and generative factors of the observational data. The framework of variational autoencoder (VAE) is commonly used to disentangle independent factors from observations. However, in real scenarios, factors with semantics are not necessarily independent. Instead, there might be an underlying causal structure which renders these factors dependent. We thus propose a new VAE based framework named CausalVAE, which includes a Causal Layer to transform independent exogenous factors into causal endogenous ones that correspond to causally related concepts in data. We further analyze the model identifiabitily, showing that the proposed model learned from observations recovers the true one up to a certain degree. Experiments are conducted on various datasets, including synthetic and real word benchmark CelebA. Results show that the causal representations learned by CausalVAE are semantically interpretable, and their causal relationship as a Directed Acyclic Graph (DAG) is identified with good accuracy. Furthermore, we demonstrate that the proposed CausalVAE model is able to generate counterfactual data through "do-operation" to the causal factors. Mengyue Yang, Furui Liu, Zhitang Chen, Xinwei Shen 0002, Jianye Hao, Jun Wang 0012 |
CVPR | 3 |
| 2021 | A Unified Framework for Layout Pattern Analysis with Deep Causal EstimationabstractThe decrease of feature size and the growing complexity of the fabrication process lead to more failures in manufacturing semiconductor devices. Therefore, identifying the root cause layout patterns of failures becomes increasingly crucial for yield improvement. In this paper, a novel layout-aware diagnosis-based layout pattern analysis framework is proposed to identify the root cause efficiently. At the first stage of the framework, an encoder network trained using contrastive learning is used to extract representations of layout snippets that are invariant to trivial transformations including shift, rotation, and mirroring, which are then clustered to form layout patterns. At the second stage, we model the causal relationship between any potential root cause layout patterns and the systematic defects by a structural causal model, which is then used to estimate the Average Causal Effect (ACE) of candidate layout patterns on the systematic defect to identify the true root cause. Experimental results on real industrial cases demonstrate that our framework outperforms a commercial tool with higher accuracies and around x8.4 speedup on average. Ran Chen 0001, Shoubo Hu, Zhitang Chen, Shengyu Zhu 0001, Bei Yu 0001, Pengyun Li, Yu Huang 0005, Jianye Hao |
ICCAD | 3 |
| 2021 | Ordering-Based Causal Discovery with Reinforcement LearningabstractIt is a long-standing question to discover causal relations among a set of variables in many empirical sciences. Recently, Reinforcement Learning (RL) has achieved promising results in causal discovery from observational data. However, searching the space of directed graphs and enforcing acyclicity by implicit penalties tend to be inefficient and restrict the existing RL-based method to small scale problems. In this work, we propose a novel RL-based approach for causal discovery, by incorporating RL into the ordering-based paradigm. Specifically, we formulate the ordering search problem as a multi-step Markov decision process, implement the ordering generating process with an encoder-decoder architecture, and finally use RL to optimize the proposed model based on the reward mechanisms designed for each ordering. A generated ordering would then be processed using variable selection to obtain the final causal graph. We analyze the consistency and computational complexity of the proposed method, and empirically show that a pretrained model can be exploited to accelerate training. Experimental results on both synthetic and real data sets shows that the proposed method achieves a much improved performance over existing RL-based method. Xiaoqiang Wang 0003, Yali Du 0001, Shengyu Zhu 0001, Liangjun Ke, Zhitang Chen, Jianye Hao, Jun Wang 0012 |
IJCAI | 5 |
| 2021 | Asymptotically Optimal One- and Two-Sample Testing With KernelsabstractWe characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test achieves the maximum rate subject to a constant level constraint on the type-I error probability. With Sanov's theorem, we derive a sufficient condition for one-sample tests to achieve the optimal error exponent in the universal setting, i.e., for any distribution defining the alternative hypothesis. We then show that two classes of Maximum Mean Discrepancy (MMD) based tests attain the optimal type-II error exponent on \mathbb Rd, while the quadratic-time Kernel Stein Discrepancy (KSD) based tests achieve this optimality with an asymptotic level constraint. For general two-sample testing, however, Sanov's theorem is insufficient to obtain a similar sufficient condition. We proceed to establish an extended version of Sanov's theorem and derive an exact error exponent for the quadratic-time MMD based two-sample tests. The obtained error exponent is further shown to be optimal among all two-sample tests satisfying a given level constraint. Our work hence provides an achievability result for optimal nonparametric one- and two-sample testing in the universal setting. Application to off-line change detection and related issues are also discussed. Shengyu Zhu 0001, Biao Chen 0001, Zhitang Chen, Pengfei Yang 0003 |
IEEE Trans. Inf. Theory | 3 |
| 2020 | Causal Discovery with Reinforcement Learning
Shengyu Zhu 0001, Ignavier Ng, Zhitang Chen |
ICLR | 3 |
| 2020 | Stable Learning via Differentiated Variable DecorrelationabstractRecently, as the applications of artificial intelligence gradually seeping into some risk-sensitive areas such as justice, healthcare and autonomous driving, an upsurge of research interest on model stability and robustness has arisen in the field of machine learning. Rather than purely fitting the observed training data, stable learning tries to learn a model with uniformly good performance under non-stationary and agnostic testing data. The key challenge of stable learning in practice is that we do not have any knowledge about the true model and test data distribution as a priori. Under such condition, we cannot expect a faithful estimation of model parameters and its stability over wild changing environments. Previous methods resort to a reweighting scheme to remove the correlations between all the variables through a set of new sample weights. However, we argue that such aggressive decorrelation between all the variables may cause the over-reduced sample size, which leads to the variance inflation and possible underperformance. In this paper, we incorporate the unlabled data from multiple environments into the variable decorrelation framework and propose a Differentiated Variable Decorrelation (DVD) algorithm based on the clustering of variables. Specifically, the variables are clustered according to the stability of their correlations and the variable decorrelation module learns a set of sample weights to remove the correlations merely between the variables of different clusters. Empirical studies on both synthetic and real world datasets clearly demonstrate the efficacy of our DVD algorithm on improving the model parameter estimation and the prediction stability over changing distributions. Zheyan Shen, Peng Cui 0001, Tong Zhang 0001, Bo Li 0064, Zhitang Chen |
KDD | 6 |
| 2020 | Discriminative training of feed-forward and recurrent sum-product networks by extended Baum-Welch
Haonan Duan 0002, Abdullah Rashwan, Pascal Poupart, Zhitang Chen |
Int. J. Approx. Reason. | 4 |
| 2019 | Universal Hypothesis Testing with Kernels: Asymptotically Optimal Tests for Goodness of FitabstractWe characterize the asymptotic performance of nonparametric goodness of fit testing. The exponential decay rate of the type-II error probability is used as the asymptotic performance metric, and a test is optimal if it achieves the maximum rate subject to a constant level constraint on the type-I error probability. We show that two classes of Maximum Mean Discrepancy (MMD) based tests attain this optimality on $\mathbb R^d$, while the quadratic-time Kernel Stein Discrepancy (KSD) based tests achieve the maximum exponential decay rate under a relaxed level constraint. Under the same performance metric, we proceed to show that the quadratic-time MMD based two-sample tests are also optimal for general two-sample problems, provided that kernels are bounded continuous and characteristic. Key to our approach are Sanov’s theorem from large deviation theory and the weak metrizable properties of the MMD and KSD. Shengyu Zhu 0001, Biao Chen 0001, Pengfei Yang 0003, Zhitang Chen |
AISTATS | 4 |
| 2019 | Kernel-based Multi-Task Contextual Bandits in Cellular Network ConfigurationabstractCellular network configuration plays a critical role n network performance. In current practice, network configuration depends heavily on field experience of engineers and often remains static for a long period of time. This practice is far from optimal. To address this limitation, online-learning-based approaches have great potentials to automate and optimize network configuration. Learning-based approaches face the challenges of learning a highly complex function for each base station and balancing the fundamental exploration-exploitation tradeoff while minimizing the exploration cost. Fortunately, in cellular networks, base stations (BSs) often have similarities even though they are not identical. To leverage such similarities, we propose kernel-based multi-BS contextual bandit algorithm based on multi-task learning. In the algorithm, we leverage the similarity among different BSs defined by conditional kernel embedding. We present theoretical analysis of the proposed algorithm in terms of regret and multi-task-learning efficiency. We evaluate the effectiveness of our algorithm based on a simulator built by real traces. Xiaoxiao Wang 0002, Xueying Guo, Jie Chuai, Zhitang Chen, Xin Liu 0002 |
IEEE BigData | 4 |
| 2019 | A Collaborative Learning Based Approach for Parameter Configuration of Cellular NetworksabstractCellular network performance depends heavily on the configuration of its network parameters. Current practice of parameter configuration relies largely on expert experience, which is often suboptimal, time-consuming, and error-prone. Therefore, it is desirable to automate this process to improve the accuracy and efficiency via learning-based approaches. However, such approaches need to address several challenges in real operational networks: the lack of diverse historical data, a limited amount of experiment budget set by network operators, and highly complex and unknown network performance functions. To address those challenges, we propose a collaborative learning approach to leverage data from different cells to boost the learning efficiency and to improve network performance. Specifically, we formulate the problem as a transferable contextual bandit problem, and prove that by transfer learning, one could significantly reduce the regret bound. Based on the theoretical result, we further develop a practical algorithm that decomposes a cell's policy into a common homogeneous policy learned using all cells' data and a cell-specific policy that captures each individual cell's heterogeneous behavior. We evaluate our proposed algorithm via a simulator constructed using real network data and demonstrates faster convergence compared to baselines. More importantly, a live field test is also conducted on a real metropolitan cellular network consisting 1700+ cells to optimize five parameters for two weeks. Our proposed algorithm shows a significant performance improvement of 20%. Jie Chuai, Zhitang Chen, Guochen Liu, Xueying Guo, Xiaoxiao Wang 0002, Xin Liu 0002, Chongming Zhu, Feiyi Shen |
INFOCOM | 2 |
| 2019 | Domain Generalization via Multidomain Discriminant Analysis
Shoubo Hu, Kun Zhang 0001, Zhitang Chen, Lai-Wan Chan |
UAI | 3 |
| 2019 | Model-free inference of diffusion networks using RKHS embeddings
Shoubo Hu, Bogdan Cautis, Zhitang Chen, Lai-Wan Chan, Yanhui Geng, Xiuqiang He 0001 |
Data Min. Knowl. Discov. | 3 |
| 2018 | Discovering and Removing Exogenous State Variables and Rewards for Reinforcement LearningabstractExogenous state variables and rewards can slow down reinforcement learning by injecting uncontrolled variation into the reward signal. We formalize exogenous state variables and rewards and identify conditions under which an MDP with exogenous state can be decomposed into an exogenous Markov Reward Process involving only the exogenous state+reward and an endogenous Markov Decision Process defined with respect to only the endogenous rewards. We also derive a variance-covariance condition under which Monte Carlo policy evaluation on the endogenous MDP is accelerated compared to using the full MDP. Similar speedups are likely to carry over to all RL algorithms. We develop two algorithms for discovering the exogenous variables and test them on several MDPs. Results show that the algorithms are practical and can significantly speed up reinforcement learning. Thomas G. Dietterich, George Trimponias, Zhitang Chen |
ICML | 3 |
| 2018 | Causal Inference and Mechanism Clustering of A Mixture of Additive Noise ModelsabstractThe inference of the causal relationship between a pair of observed variables is a fundamental problem in science, and most existing approaches are based on one single causal model. In practice, however, observations are often collected from multiple sources with heterogeneous causal models due to certain uncontrollable factors, which renders causal analysis results obtained by a single model skeptical. In this paper, we generalize the Additive Noise Model (ANM) to a mixture model, which consists of a finite number of ANMs, and provide the condition of its causal identifiability. To conduct model estimation, we propose Gaussian Process Partially Observable Model (GPPOM), and incorporate independence enforcement into it to learn latent parameter associated with each observation. Causal inference and clustering according to the underlying generating mechanisms of the mixture model are addressed in this work. Experiments on synthetic and real data demonstrate the effectiveness of our proposed approach. Shoubo Hu, Zhitang Chen, Vahid Partovi Nia, Lai-Wan Chan, Yanhui Geng |
NeurIPS | 2 |
| 2018 | Learning-Based Joint Configuration for Cellular NetworksabstractCellular network configuration is critical for network performance. Current practice is mostly based on field experience and manual adjustment. The process is labor-intensive, error-prone, and far from optimal. To automate and optimize cellular network configuration, in this paper, we propose an online-learning-based joint-optimization approach that addresses a few specific challenges: limited data availability, convoluted sample data, highly complex optimization due to interactions among neighboring cells, and the need to adapt to network dynamics. In our approach, to learn an appropriate utility function for a cell, we develop a neural-network-based model that addresses the convoluted sample data issue and achieves good accuracy based on data aggregation. Based on the utility function learned, we formulate a global network configuration optimization problem. To solve this high-dimensional nonconcave maximization problem, we design a Gibbs-sampling-based algorithm that converges to an optimal solution when a technical parameter is small enough. Furthermore, we design an online scheme that updates the learned utility function and solves the corresponding maximization problem efficiently to adapt to network dynamics. To illustrate the idea, we use the case study of pilot power configuration. Numerical results illustrate the effectiveness of the proposed approach. Xueying Guo, George Trimponias, Xiaoxiao Wang 0002, Zhitang Chen, Yanhui Geng, Xin Liu 0002 |
IEEE Internet Things J. | 4 |
| 2018 | A Kernel Embedding-Based Approach for Nonstationary Causal Model InferenceabstractAlthough nonstationary data are more common in the real world, most existing causal discovery methods do not take nonstationarity into consideration. In this letter, we propose a kernel embedding-based approach, ENCI, for nonstationary causal model inference where data are collected from multiple domains with varying distributions. In ENCI, we transform the complicated relation of a cause-effect pair into a linear model of variables of which observations correspond to the kernel embeddings of the cause-and-effect distributions in different domains. In this way, we are able to estimate the causal direction by exploiting the causal asymmetry of the transformed linear model. Furthermore, we extend ENCI to causal graph discovery for multiple variables by transforming the relations among them into a linear nongaussian acyclic model. We show that by exploiting the nonstationarity of distributions, both cause-effect pairs and two kinds of causal graphs are identifiable under mild conditions. Experiments on synthetic and real-world data are conducted to justify the efficacy of ENCI over major existing methods. Shoubo Hu, Zhitang Chen, Lai-Wan Chan |
Neural Comput. | 2 |
| 2017 | Seq2Img: A sequence-to-image based approach towards IP traffic classification using convolutional neural networksabstractIP traffic classification has been a vitally important topic that attracts persistent interest in the networking and machine learning communities for past decades. While there exist quite a number of works applying machine learning techniques to realize IP traffic classification, most works suffer from limitations like either heavily depending on handcrafted features or be only able to handle offline traffic classification. To get rid of the aforementioned weakness, in this paper, we propose our online Convolutional Neural Networks (CNNs) based traffic classification framework named Seq2Img. The basic idea is to employ a compact nonparametric kernel embedding based method to convert early flow sequences into images which fully capture the static and dynamic behaviors of different applications and avoid using handcrafted features that might cause loss of information. A CNN is then applied on the generated images to obtain traffic classification results. Experiments on real network traffic are conducted and encouraging results justify the efficacy of our proposed approach. Zhitang Chen, Yanhui Geng |
IEEE BigData | 1 |
| 2017 | Cellular network configuration via online learning and joint optimizationabstractCellular network configuration is critical for network performance. Current practice is labor-intensive, error-prone, and far from optimal. To automate efficient cellular network configuration, in this work, we propose an online-learning-based joint-optimization approach that addresses a few specific challenges: limited data availability, convoluted sample data, highly complex optimization due to interactions among neighboring cells, and the need to adapt to network dynamics. In our approach, to learn an appropriate utility function for a cell, we develop a neural-network-based model that addresses the convoluted sample data issue and achieves good accuracy based on data aggregation. Based on the utility function learned, we formulate a global network configuration optimization problem. To solve this high-dimensional non-concave maximization problem, we design a Gibbs-sampling-based algorithm that converges to an optimal solution when a technical parameter is small enough. Furthermore, we design an online scheme that updates the learned utility function and solves the corresponding maximization problem efficiently to adapt to network dynamics. To illustrate the idea, we use the case study of pilot power configuration. Numerical results illustrate the effectiveness of the proposed approach. Xueying Guo, George Trimponias, Xiaoxiao Wang 0002, Zhitang Chen, Yanhui Geng, Xin Liu 0002 |
IEEE BigData | 4 |
| 2017 | Online Bayesian Transfer Learning for Sequential Data Modeling
Priyank Jaini, Zhitang Chen, Pablo Carbajal, Edith Law, Laura Middleton, Kayla Regan, Mike Schaekermann, George Trimponias, James Tung, Pascal Poupart |
ICLR (Poster) | 2 |
| 2016 | Online Relative Entropy Policy Search using Reproducing Kernel Hilbert Space EmbeddingsabstractKernel methods have been successfully applied to reinforcement learning problems to address some challenges such as high dimensional and continuous states, value function approximation and state transition probability modeling. In this paper, we develop an online policy search algorithm based on a recent state-of-the-art algorithm REPS-RKHS that uses conditional kernel embeddings. Our online algorithm inherits the advantages of REPS-RKHS, including the ability to learn non-parametric control policies for infinite horizon continuous MDPs with high- dimensional sensory representations. Different from the original REPS-RKHS algorithm which is based on batch learning, the proposed online algorithm updates the model in an online fashion and thus is able to capture and respond to rapid changes in the system dynamics. In addition, the online update operation takes constant time (i.e., independent of the sample size n), which is much more efficient computationally and allows the policy to be continuously revised. Experiments on different domains are conducted and results show that our online algorithm outperforms the original algorithm. Zhitang Chen, Pascal Poupart, Yanhui Geng |
AISTATS | 1 |
| 2016 | Predicting future traffic using Hidden Markov ModelsabstractNetwork traffic volume estimation and prediction is an important research topic that attracts persistent attention from the networking community and the machine learning community. Although there has been extensive work on estimating or predicting the traffic matrix using time series models, low rank matrix decomposition et. al, to the best of our knowledge, there is few work investigating the problem whether we are able to estimate and predict the traffic volume based on some statistics of the traffic which are much less costly to collect, for example, the flow counts. In this paper, we propose to model the relationship between the traffic volume and simple statistics about flows using a Hidden Markov Model based on which we can avoid direct measurement of the traffic volume but instead we estimate and predict the hidden traffic volume based on those simple flow statistics which are collected by some sketch techniques. We demonstrate the feasibility and effectiveness of our proposed method using some semi-simulation and real data experimental results. Zhitang Chen, Jiayao Wen, Yanhui Geng |
ICNP | 1 |
| 2016 | Online flow size prediction for improved network routingabstractWe describe an emerging application of data mining in the context of computer networks. This application concerns the problem of predicting the size of a flow and detecting elephant flows (very large flows). Flow size is a very important statistic that can be used to improve routing, load balancing and scheduling in computer networks. Flow size prediction is particularly challenging since flow patterns continuously change and predictions must be done in real time (milliseconds) to avoid delays. We describe how to formulate the problem as an online machine learning task to continuously adjust to changes in flow traffic. We evaluate the predictive nature of a set of features and the accuracy of three online predictors based on neural networks, Gaussian process regression and online Bayesian Moment Matching on three datasets of real traffic. We also demonstrate how to use such online predictors to improve routing (i.e., reduced flow completion time) in a network simulation. Pascal Poupart, Zhitang Chen, Priyank Jaini, Fred Fung, Hengky Susanto, Yanhui Geng, Li Chen 0008, Kai Chen 0005 |
ICNP | 2 |
| 2014 | Causal Discovery via Reproducing Kernel Hilbert Space EmbeddingsabstractCausal discovery via the asymmetry between the cause and the effect has proved to be a promising way to infer the causal direction from observations. The basic idea is to assume that the mechanism generating the cause distribution p(x) and that generating the conditional distribution p(y|x) correspond to two independent natural processes and thus p(x) and p(y|x) fulfill some sort of independence condition. However, in many situations, the independence condition does not hold for the anticausal direction; if we consider p(x, y) as generated via p(y)p(x|y), then there are usually some contrived mutual adjustments between p(y) and p(x|y). This kind of asymmetry can be exploited to identify the causal direction. Based on this postulate, in this letter, we define an uncorrelatedness criterion between p(x) and p(y|x) and, based on this uncorrelatedness, show asymmetry between the cause and the effect in terms that a certain complexity metric on p(x) and p(y|x) is less than the complexity metric on p(y) and p(x|y). We propose a Hilbert space embedding-based method EMD (an abbreviation for EMbeDding) to calculate the complexity metric and show that this method preserves the relative magnitude of the complexity metric. Based on the complexity metric, we propose an efficient kernel-based algorithm for causal discovery. The contribution of this letter is threefold. It allows a general transformation from the cause to the effect involving the noise effect and is applicable to both one-dimensional and high-dimensional data. Furthermore it can be used to infer the causal ordering for multiple variables. Extensive experiments on simulated and real-world data are conducted to show the effectiveness of the proposed method. Zhitang Chen, Kun Zhang 0001, Lai-Wan Chan, Bernhard Schölkopf |
Neural Comput. | 1 |
| 2013 | Nonlinear Causal Discovery for High Dimensional Data: A Kernelized Trace MethodabstractCausal discovery for high-dimensional observations is a useful tool in many fields such as climate analysis and financial market analysis. A linear Trace method has been proposed to identify the causal direction between two linearly coupled high-dimensional observations X and Y. However, in reality, the relations between X and Y are usually nonlinear and consequently the linear Trace method may fail. In this paper, we propose a method to infer the nonlinear causal relations for two high-dimensional observations X and Y. The idea is to map the observations to high dimensional Reproducing Kernel Hilbert Space (RKHS) such that the nonlinear relations become simple linear ones. We show that the linear Trace condition holds for the causal direction but it is violated for the anti-causal direction in RKHS. Based on this theoretical result, we develop a simple algorithm to infer the causal direction for nonlinearly coupled causal pairs. Synthetic data and real world data experiments are conducted to show the effectiveness of our proposed method. Zhitang Chen, Kun Zhang 0001, Lai-Wan Chan |
ICDM | 1 |
| 2013 | Causality in Linear Nongaussian Acyclic Models in the Presence of Latent Gaussian ConfoundersabstractLiNGAM has been successfully applied to some real-world causal discovery problems. Nevertheless, causal sufficiency is assumed; that is, there is no latent confounder of the observations, which may be unrealistic for real-world problems. Taking into the consideration latent confounders will improve the reliability and accuracy of estimations of the real causal structures. In this letter, we investigate a model called linear nongaussian acyclic models in the presence of latent gaussian confounders (LiNGAM-GC) which can be seen as a specific case of lvLiNGAM. This model includes the latent confounders, which are assumed to be independent gaussian distributed and statistically independent of the disturbances. To tackle the causal discovery problem of this model, first we propose a pairwise cumulant-based measure of causal directions for cause-effect pairs. We prove that in spite of the presence of latent gaussian confounders, the causal direction of the observed cause-effect pair can be identified under the mild condition that the disturbances are simultaneously supergaussian or subgaussian. We propose a simple and efficient method to detect the violation of this condition. We extend our work to multivariate causal network discovery problems. Specifically we propose algorithms to estimate the causal network structure, including causal ordering and causal strengths, using an iterative root finding-removing scheme based on pairwise measure. To address the redundant edge problem due to the finite sample size effect, we develop an efficient bootstrapping-based pruning algorithm. Experiments on synthetic data and real-world data have been conducted to show the applicability of our model and the effectiveness of our proposed algorithms. Zhitang Chen, Lai-Wan Chan |
Neural Comput. | 1 |
| 2012 | Causal discovery with scale-mixture model for spatiotemporal variance dependenciesabstractIn conventional causal discovery, structural equation models (SEM) are directly applied to the observed variables, meaning that the causal effect can be represented as a function of the direct causes themselves. However, in many real world problems, there are significant dependencies in the variances or energies, which indicates that causality may possibly take place at the level of variances or energies. In this paper, we propose a probabilistic causal scale-mixture model with spatiotemporal variance dependencies to represent a specific type of generating mechanism of the observations. In particular, the causal mechanism including contemporaneous and temporal causal relations in variances or energies is represented by a Structural Vector AutoRegressive model (SVAR). We prove the identifiability of this model under the non-Gaussian assumption on the innovation processes. We also propose algorithms to estimate the involved parameters and discover the contemporaneous causal structure. Experiments on synthesis and real world data are conducted to show the applicability of the proposed model and algorithms. Zhitang Chen, Kun Zhang 0001, Lai-Wan Chan |
NIPS | 1 |
| 2011 | New approaches for solving permutation indeterminacy and scaling ambiguity in frequency domain separation of convolved mixturesabstractPermutation indeterminacy and scaling ambiguity occur in ICA and they are particularly problematic in time-frequency domain separation of convolutive mixtures. The quality of separation is severely degraded if these two problems are not well addressed. In this paper, we propose new approaches to solve the permutation indeterminacy and scaling ambiguity in the separation of convolutive mixture in frequency domain. We first apply Short Time Fourier Transform to the observed signals in order to transform the convolutive mixing in time domain to instantaneous mixing in time-frequency domain. A fixed-point algorithm with test of saddle point is adopted to derive the separated components in each frequency bin. To solve the permutation problem,we propose a new matching algorithm for this purpose. First we use discrete Haar Wavelet Transform to extract the feature vectors from the magnitude waveforms of the separated components and use Singular Value Decomposition to achieve dimension reduction. The permutation problem is solved by clustering the feature vectors using the new matching algorithm which is a combination of basic K-means and Hungarian algorithm. To solve the scaling ambiguity problem, we treat it as an overcomplete problem and realize it by maximizing the posterior of the scaling factor. Finally, experiments are conducted using benchmark data to present the effectiveness and performance of our proposed algorithms. Zhitang Chen, Lai-Wan Chan |
IJCNN | 1 |