EDBT 2026 Demo / reviewers in the wild / expert
Xiao Fu 0001
dblp:60/4601-1
· DBLP profile ↗
75ranked-venue papers
16as first author
40since 2021 · last 2026
0000-0003-4847-9586ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 10 first-author · 24 since 2021Artificial intelligence and machine learning · 27 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Coupled Tensor Analysis for Hyperspectral Superresolution: Recoverable Modeling under Endmember VariabilityabstractAbstract. This work revisits the hyperspectral superresolution (HSR) problem, i.e., fusing a pair of spatially co-registered hyperspectral (HSI) and multispectral (MSI) images to recover a superresolution image (SRI) that enhances the spatial resolution of the HSI. Coupled tensor decomposition–based methods have gained traction in this domain, offering recoverability guarantees under various assumptions. Existing models such as canonical polyadic decomposition (CPD) and Tucker decomposition provide strong expressive power but lack physical interpretability. The block-term decomposition model with rank-[Formula: see text] terms (the LL1 model) yields interpretable factors under the linear mixture model (LMM) of spectral images, but LMM assumptions are often violated in practice—primarily due to nonlinear effects such as endmember variability (EV). To address this issue, we propose representing spectral images using a more flexible block-term tensor model with rank-[Formula: see text] terms (the LMN model). This modeling choice retains interpretability, subsumes CPD, Tucker, and LL1 as special cases, and robustly accounts for non-ideal effects such as EV, offering a balanced trade-off between expressiveness and interpretability for HSR. Importantly, under the LMN model for HSI and MSI, recoverability of the SRI can still be established under proper conditions—providing strong theoretical support. Extensive experiments on synthetic and real datasets demonstrate the effectiveness and robustness of the proposed method. Xiao Fu 0001 |
SIAM J. Imaging Sci. | 2 |
| 2025 | Integrated Interpolation and Matrix Completion for Radio Map Estimation: A Convex Optimization ApproachabstractRadio map estimation (RME) is crucial for effective planning and optimization of wireless networks. Traditional approaches such as interpolation excel at capturing local smoothness in densely populated data but struggle with sparse or irregular data. Conversely, matrix completion (MC) approaches utilize global structures but require huge number of samples and may produce non-smooth estimates. To integrate these strengths, we propose a convex optimization approach for RME (IIMC-RME) that merges interpolation with MC. This approach formulates the RME task as a low-rank MC problem constrained by interpolated results. Additionally, we have developed a convergent algorithm utilizing the alternating direction method of multipliers (ADMM) to efficiently solve the IIMC-RME problem. Experimental evaluations on both synthetic and real-world datasets have shown that IIMC-RME surpasses existing approaches, thereby achieving superior accuracy in RME. Hongcheng Dong, Wenqiang Pu, Rui Zhou 0016, Xiao Fu 0001, Feng Yin 0001 |
ICASSP | 4 |
| 2025 | Chinese Speech Processing via Chinese Character FeatureabstractThis paper focuses on the basic structure of Chinese characters: semantic-phonetic compound characters. This paper takes advantage of this feature of Chinese characters and innovatively proposes a Chinese speech-processing method based on character shape. We use the association between the character shape and pronunciation of Chinese characters to construct a new character stroke-based dataset. We use two neural network structures, RNN and Transformer, to verify our proposed Chinese speech processing method.It is proved through experiments that the method improves the performance of Mandarin ASR(Automatic Speech Recognition) by about 2% and AEC(ASR Error Correction) by about 1%. Theoretically, this method applies to all Chinese speech-processing algorithms based on the attention mechanism. Wei Xi 0003, Xiao Fu 0001, Jizhong Zhao |
ICASSP | 4 |
| 2025 | Under-Counted Matrix Completion Without Detection FeaturesabstractUnder-counted matrix completion (UC-MC) has many important applications, especially in epidemiology and ecology where the observed data are often smaller than the actual numbers. Existing works model the under-counting effects using entry-wise miss detection probabilities, which are usually formulated as functions of detection-related side information or features (e.g., weather conditions for observing a certain species). However, such features are not always available. This work proposes a model for UC-MC that circumvents using such side information. By assuming that the detection probabilities for a large proportion of entries are similar, the under-counting probabilities are approximated by a specially structured (i.e., rank-one plus sparse) matrix. This way, a structure-regularized UC-MC formulation is attained, and no detection-related side information is used. A re-parameterization-based implementation is proposed, allowing us to employ off-the-shelf gradient-based optimizers to tackle the loss function of interest. Simulations are used to demonstrate the effectiveness of the proposed approach. Tri Nguyen 0004, Shahana Ibrahim, Rebecca A. Hutchinson, Xiao Fu 0001 |
ICASSP | 4 |
| 2025 | Radio Map Estimation via Latent-Domain Plug-and-Play DenoisersabstractRadio map estimation (RME) aims to construct a map of radio strength across multiple domains (e.g., space and frequency) from limited measurements. Data-driven deep neural model-based RME showed promising performance, yet requiring excessive training resources. This work puts forth an RME approach that can effectively incorporate learned information without training over radio map data. Our idea is to employ the plug-and-play (PnP) denoising scheme from computational imaging. The PnP framework allows incorporating denoisers trained over natural images to handle other types of data, e.g., ocean sound fields and medical images, due to the similarity of their denoising processes. Conventional PnP methods mostly use the learned denoisers in the data domain. Instead, the proposed approach applies PnP in the latent domain through a tailored algorithm design and spatial-spectral factorization of radio maps. This way, the proposed RME method exhibits enhanced scalability and noise robustness. Simulations are used to illustrate the effectiveness of the proposed approach. Lei Cheng 0003, Wenqiang Pu, Xiao Fu 0001 |
ICASSP | 5 |
| 2025 | Open-Modality Latent Modality Interaction Maximization for Audio-Visual LearningabstractThe utilization of multimodal cues enhances the effectiveness of specific cognitive tasks in audio-visual learning. However, on the one hand, designing a unified model for multimodal learning poses challenges due to the presence of information redundancy and modality noise. On the other hand, existing multimodal models face limitations in handling the modality-missing inference. In this work, we propose a Latent Modality Interaction with mutual information Maximization (LMIM) model architecture for multimodal learning, which effectively integrates multimodal cues by learning essential modality information and reducing the redundant information. We employ a group of latent tokens as pivots to filter out noise and redundancy across different modalities. Simultaneously, mutual information maximization and distribution alignment are utilized to preserve task-related information through multimodal fusion. Furthermore, a random modality masking training strategy is employed to mitigate potential over-reliance on dominant modality. Extensive experiments demonstrate that our model achieves significant improvement over current competitive baselines on two datasets, including UCF51 and Kinetics-Sounds datasets. Xiao Fu 0001, Wei Xi 0003, Jizhong Zhao |
ICASSP | 3 |
| 2025 | Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech RecognitionabstractAudio-visual speech recognition (AVSR) aims to enhance the robustness of an automatic speech recognition (ASR) systems by incorporating visual information from lip movements, especially in challenging noisy environments. Nevertheless, most current approaches either involve training from scratch or fully finetuning a pre-trained model, both of which incur significant computational costs and are often impractical for large-scale speech foundation models. This gap highlights the need for more efficient methods to leverage visual and acoustic information in AVSR tasks. To address this challenge, we propose AVWhisper, a parameter-efficient model that integrates visual and acoustic representations by injecting visual features from the AV-HuBERT encoder into the pre-trained Whisper model. Our approach leverages the existing attention mechanisms in Whisper to facilitate cross-modal interaction and integrates auxiliary visual information through lightweight adapters based on Low-Rank Adaptation (LoRA) and prompt-based techniques. Furthermore, a two-phase training strategy is adopted to effectively handle cross-domain differences and visual information injection problems respectively. Extensive experiments on the LRS3-TED dataset demonstrate that AVWhisper consistently outperforms state-of-the-art methods across various noise conditions, offering a more efficient and scalable solution for audio-visual speech recognition. Yue Heng Yeo, Xiao Fu 0001, Weiguang Chen, Wei Xi 0003, Jizhong Zhao |
ICASSP | 4 |
| 2025 | Content-Style Learning from Unaligned Domains: Identifiability under Unknown Latent DimensionsabstractUnderstanding identifiability of latent content and style variables from unaligned multi-domain data is essential for tasks such as
domain translation and data generation. Existing works on content-style identification were often developed under somewhat stringent conditions, e.g., that all latent components are mutually independent and that the dimensions of the content and style variables are known. We introduce a new analytical framework via cross-domain *latent distribution matching* (LDM), which establishes content-style identifiability under substantially more relaxed conditions. Specifically, we show that restrictive assumptions such as component-wise independence of the latent variables can be removed. Most notably, we prove that prior knowledge of the content and style dimensions is not necessary for ensuring identifiability, if sparsity constraints are properly imposed onto the learned latent representations. Bypassing the knowledge of the exact latent dimension has been a longstanding aspiration in unsupervised representation learning---our analysis is the first to underpin its theoretical and practical viability. On the implementation side, we recast the LDM formulation into a regularized multi-domain GAN loss with coupled latent variables. We show that the reformulation is equivalent to LDM under mild conditions---yet requiring considerably less computational resource. Experiments corroborate with our theoretical claims. Sagar Shrestha, Xiao Fu 0001 |
ICLR | 2 |
| 2025 | Diversified Flow Matching with Translation IdentifiabilityabstractDiversified distribution matching (DDM) finds a unified translation function mapping a diverse collection of conditional source distributions to their target counterparts. DDM was proposed to resolve content misalignment issues in unpaired domain translation, achieving translation identifiability. However, DDM has only been implemented using GANs due to its constraints on the translation function. GANs are often unstable to train and do not provide the transport trajectory information---yet such trajectories are useful in applications such as single-cell evolution analysis and robot route planning. This work introduces *diversified flow matching* (DFM), an ODE-based framework for DDM. Adapting flow matching (FM) to enforce a unified translation function as in DDM is challenging, as FM learns the translation function's velocity rather than the translation function itself. A custom bilevel optimization-based training loss, a nonlinear interpolant, and a structural reformulation are proposed to address these challenges, offering a tangible implementation. To our knowledge, DFM is the first ODE-based approach guaranteeing translation identifiability. Experiments on synthetic and real-world datasets validate the proposed method. Sagar Shrestha, Xiao Fu 0001 |
ICML | 2 |
| 2025 | Enhancing Multimodal Model Robustness Under Missing Modalities via Memory-Driven Prompt LearningabstractExisting multimodal models typically assume the availability of all modalities, leading to significant performance degradation when certain modalities are missing. Recent methods have introduced prompt learning to adapt pretrained models to incomplete data, achieving remarkable performance when the missing cases are consistent during training and inference. However, these methods rely heavily on distribution consistency and fail to compensate for missing modalities, limiting their ability to generalize to unseen missing cases. To address this issue, we propose Memory-Driven Prompt Learning, a framework that adaptively compensates for missing modalities through prompt learning. The compensation strategies are achieved by two types of prompts: generative prompts and shared prompts. Generative prompts retrieve semantically similar samples from a predefined prompt memory that stores modality-specific semantic information, while shared prompts leverage available modalities to provide cross-modal compensation. Extensive experiments demonstrate the effectiveness of the proposed model, achieving significant improvements across diverse missing-modality scenarios, with average performance increasing from 34.76% to 40.40% on MM-IMDb, 62.71% to 77.06% on Food101, and 60.40% to 62.77% on Hateful Memes. The code is available at https://github.com/zhao-yh20/MemPrompt. Yihan Zhao, Wei Xi 0003, Xiao Fu 0001, Jizhong Zhao |
IJCAI | 3 |
| 2025 | Visually-Adaptive Guided Robust Speech Recognition with Parameter-Efficient Adaptation
Yue Heng Yeo, Xiao Fu 0001, Wei Xi 0003, Jizhong Zhao |
INTERSPEECH | 4 |
| 2025 | LES-CLIP: A Lightweight Emotion-Sensitive Adaptation of CLIP for Precise Similar Emotion DiscriminationabstractCLIP has been widely adopted in affective computing for its strong vision-language representation capabilities. However, it fails to accurately distinguish visually similar yet label-distinct facial expressions. This limitation is rooted in CLIP's encoding paradigm and large-scale contrastive pretraining, which bias the model toward focusing primarily on globally salient visual features and aligning them with broad semantic concepts. Such alignment overlooks subtle facial variations and induces representational shortcuts, where emotionally distinct categories are projected into overlapping regions of the shared semantic space. This semantic entanglement severely compromises the model's ability to preserve emotional separability. We propose LES-CLIP, a Lightweight and Emotion-Sensitive framework that adapts CLIP for precise discrimination of similar emotions. LES-CLIP achieves fine-grained emotional sensitivity using only simple text prompts and facial images. It introduces three novel components: 1) an Emotion-Sensitive Adaptive Mixture-of-Experts, which pre-adapts representations for subtle expression discrimination; 2) a Prompt-Guided Emotion Discrimination module that activates CLIP's visual sensitivity to fine-grained facial cues; and 3) a LES hybrid loss that guides contrastive learning toward accurate emotion-label alignment. Extensive experiments demonstrate that LES-CLIP achieves state-of-the-art performance, reaching 70.18% on the 8-class AffectNet dataset. Moreover, it converges faster and requires significantly fewer parameters. Xiao Fu 0001, Wei Xi 0003, Kun Zhao 0002, Jiadong Feng, Jizhong Zhao |
ACM Multimedia | 1 |
| 2025 | Diverse Influence Component Analysis: A Geometric Approach to Nonlinear Mixture IdentifiabilityabstractLatent component identification from unknown *nonlinear* mixtures is a foundational challenge in machine learning, with applications in tasks such as self-supervised learning and causal representation learning. Prior work in *nonlinear independent component analysis* (nICA) has shown that auxiliary signals---such as weak supervision---can support *identifiability* of conditionally independent latent components. More recent approaches explore structural assumptions, like sparsity in the Jacobian of the mixing function, to relax such requirements. In this work, we introduce *Diverse Influence Component Analysis* (DICA), a framework that exploits the convex geometry of the mixing function’s Jacobian. We propose a *Jacobian Volume Maximization* (J-VolMax) criterion, which enables latent component identification by encouraging diversity in their influence on the observed variables. Under suitable conditions, this approach achieves identifiability without relying on auxiliary information, latent component independence, or Jacobian sparsity assumptions. These results extend the scope of identifiability analysis and offer a complementary perspective to existing methods. Hoang-Son Nguyen, Xiao Fu 0001 |
NeurIPS | 2 |
| 2025 | Domain-Factored Untrained Deep Prior for Spectrum CartographyabstractSpectrum cartography(SC) aims to estimate the radio power map of multiple emitters over space and frequency using limited sensor data. Recent advances leverage learneddeep generative models(DGMs) as structural priors, achieving state-of-the-art performance by capturing complex spatial-spectral patterns. However, DGMs require large training datasets and may suffer under distribution shifts. To address these limitations, we propose atraining-freeSC approach based onuntrained neural networks(UNNs), which encode structural priors through architectural design. Our custom UNN exploits a spatio-spectral factorization model rooted in the physical structure of radio maps, enabling low sample complexity. Experiments show that our method matches the performance of DGM-based SC without any training data. Subash Timilsina, Sagar Shrestha, Lei Cheng 0003, Xiao Fu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | Probabilistic Simplex Component Analysis via Variational Auto-EncodingabstractSimplex component analysis (SCA) aims to estimate the vertices of the convex hull where data samples reside in. SCA finds various applications in signal processing, e.g., hyperspectral unmixing and noisy label learning. Recent works proposed to tackle SCA from a probabilistic viewpoint using variational inference (VI) tools, which fends against noise more effectively relative to the deterministic counterparts. However, the computational efficiency of VI for SCA hinges on the use of the Dirichlet variational posterior. Such variational posterior appears to lack expressiveness—making the SCA performance limited if the true posterior is complex. This work proposes to employ a logistic-normal variational posterior, which exhibits enhanced expressive power. To circumvent the computational bottleneck, a neural representation-based inference algorithm is proposed—which exploits a connection between the logistic-normal distribution and variational auto-encoding. Numerical experiments using simulated and semi-real data are conducted to showcase the effectiveness of our algorithm design. Yuening Li, Xiao Fu 0001, Wing-Kin Ma |
ICASSP | 2 |
| 2024 | A Smoothed Bregman Proximal Gradient Algorithm for Decentralized Nonconvex OptimizationabstractDecentralized computation has received considerable research interest lately, due to its wide applications in information processing systems. However, one key requirement to establish convergence for almost all decentralized algorithms, for convex and non-convex problems alike, is that the loss function has Lipschitz-continuous gradient (LipGrad). This is a strong assumption, which does not hold for many practical problems, such as matrix/tensor factorization, neural network training, etc. On the contrary, in the centralized setting, one can utilize techniques such as the Bregman proximal gradient (BPG) method to deal with the lack of LipGrad. This work fills the gap between centralized and decentralized cases by developing a novel smoothed decentralized BPG algorithm to deal with a class of nonconvex decentralized problem, where the local problems do not have LipGrad objective functions. By leveraging the recent notion of relative smoothness and primal-dual error bounds, we show that the proposed algorithm achieves a certain ε-stationary solution by using $\mathcal{O}\left( {{\varepsilon ^{ - 2}}} \right)$ iterations, matching the rate of the centralized Bregman proximal gradient method. To our knowledge, this is the first decentralized algorithm that matches the centralized convergence rate bounds under the class of considered problems. Our numerical results on the decentralized quadratic regression example demonstrate the effectiveness of proposed algorithm. Wenqiang Pu, Jiawei Zhang 0007, Rui Zhou 0016, Xiao Fu 0001, Mingyi Hong 0001 |
ICASSP | 4 |
| 2024 | Towards Identifiable Unsupervised Domain Translation: A Diversified Distribution Matching ApproachabstractUnsupervised domain translation (UDT) aims to find functions that convert samples from one domain (e.g., sketches) to another domain (e.g., photos) without changing the high-level semantic meaning (also referred to as "content"). The translation functions are often sought by probability distribution matching of the transformed source domain and target domain. CycleGAN stands as arguably the most representative approach among this line of work. However, it was noticed in the literature that CycleGAN and variants could fail to identify the desired translation functions and produce content-misaligned translations.
This limitation arises due to the presence of multiple translation functions---referred to as ``measure-preserving automorphism" (MPA)---in the solution space of the learning criteria. Despite awareness of such identifiability issues, solutions have remained elusive. This study delves into the core identifiability inquiry and introduces an MPA elimination theory. Our analysis shows that MPA is unlikely to exist, if multiple pairs of diverse cross-domain conditional distributions are matched by the learning function.
Our theory leads to a UDT learner using distribution matching over auxiliary variable-induced subsets of the domains---other than over the entire data domains as in the classical approaches. The proposed framework is the first to rigorously establish translation identifiability under reasonable UDT settings, to our best knowledge.
Experiments corroborate with our theoretical claims. Sagar Shrestha, Xiao Fu 0001 |
ICLR | 2 |
| 2024 | MRFER: Multi-Channel Robust Feature Enhanced Fusion for Multi-Modal Emotion RecognitionabstractIn multi-modal emotion recognition, previous studies focus on obtaining more distinguishable unimodal features and expanding complementary information across modalities. However, a considerable amount of latent emotional information is neglected. It leads to insufficient intra-modal representations and a one-sided perspective on inter-modal relationship learning. To address these challenges, we propose a novel framework named MRFER, which explores strategies to reduce the loss of emotional information. It models robust unimodal features through a multi-path feature extractor and captures more comprehensive inter-modal relationships through a text-guided dual attention fusion module. Systematic evaluation covers generalization and overall performance, showcasing MRFER’s advancement beyond existing state-of-the-art approaches. Xiao Fu 0001, Wei Xi 0003, Dianwen Ng, Jizhong Zhao |
ICME | 1 |
| 2024 | Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd WisdomabstractThe generation of label noise is often modeled as a process involving a probability transition matrix (also interpreted as the _annotator confusion matrix_) imposed onto the label distribution. Under this model, learning the ``ground-truth classifier''---i.e., the classifier that can be learned if no noise was present---and the confusion matrix boils down to a model identification problem. Prior works along this line demonstrated appealing empirical performance, yet identifiability of the model was mostly established by assuming an instance-invariant confusion matrix. Having an (occasionally) instance-dependent confusion matrix across data samples is apparently more realistic, but inevitably introduces outliers to the model. Our interest lies in confusion matrix-based noisy label learning with such outliers taken into consideration. We begin with pointing out that under the model of interest, using labels produced by only one annotator is fundamentally insufficient to detect the outliers or identify the ground-truth classifier. Then, we prove that by employing a crowdsourcing strategy involving multiple annotators, a carefully designed loss function can establish the desired model identifiability under reasonable conditions. Our development builds upon a link between the noisy label model and a column-corrupted matrix factorization mode---based on which we show that crowdsourced annotations distinguish nominal data and instance-dependent outliers using a low-dimensional subspace. Experiments show that our learning scheme substantially improves outlier detection and the classifier's testing accuracy. Tri Nguyen 0004, Shahana Ibrahim, Xiao Fu 0001 |
NeurIPS | 3 |
| 2024 | Identifiable Shared Component Analysis of Unpaired Multimodal MixturesabstractA core task in multi-modal learning is to integrate information from multiple feature spaces (e.g., text and audio), offering modality-invariant essential representations of data. Recent research showed that, classical tools such as canonical correlation analysis (CCA) provably identify the shared components up to minor ambiguities, when samples in each modality are generated from a linear mixture of shared and private components. Such identifiability results were obtained under the condition that the cross-modality samples are aligned/paired according to their shared information. This work takes a step further, investigating shared component identifiability from multi-modal linear mixtures where cross-modality samples are unaligned. A distribution divergence minimization-based loss is proposed, under which a suite of sufficient conditions ensuring identifiability of the shared components are derived. Our conditions are based on cross-modality distribution discrepancy characterization and density-preserving transform removal, which are much milder than existing studies relying on independent component analysis. More relaxed conditions are also provided via adding reasonable structural constraints, motivated by available side information in various applications. The identifiability claims are thoroughly validated using synthetic and real-world data. Subash Timilsina, Sagar Shrestha, Xiao Fu 0001 |
NeurIPS | 3 |
| 2023 | Towards Efficient and Optimal Joint Beamforming and Antenna Selection: A Machine Learning ApproachabstractThis work revisits the joint transmit beamforming and antenna selection problem. Existing approaches find approximate solutions to this NP-hard problem via various heuristics, e.g., convex/nonconvex relaxation, greedy method, and (deep) supervised learning. However, optimality (or even feasibility) of these heuristics is not guaranteed. To avoid sub-optimal solutions, an effective branch and bound (B&B) algorithm is proposed. B&B algorithms are ensured to return optimal solutions, but have scalability challenges. In order to enhance efficiency, a graph neural network (GNN)-based classfier is trained with imitation learning to accelerate the B&B algorithm— where the GNN is carefully designed to suit the dynamic nature of wireless communication scenarios. The GNN-based acceleration is shown to provably retain the optimality of B&B with high probability, while substantially reducing the computational burden, under reasonable conditions. Numerical experiments show that our GNN-based method always finds near-optimal and feasible solutions with significantly reduced complexity relative to the plain-vanilla B&B. Sagar Shrestha, Xiao Fu 0001, Mingyi Hong 0001 |
ICASSP | 2 |
| 2023 | Deep Spectrum Cartography Using Quantized MeasurementsabstractSpectrum cartography (SC) techniques craft multi-domain (e.g., space and frequency) radio maps from limited measurements, which is an ill-posed inverse problem. Recent works used low-dimensional priors such as a low tensor rank structure and a deep generative model to assist radio map estimation—with provable guarantees. However, a premise of these approaches is that the sensors are able to send real-valued feedback to a fusion center for SC—yet practical communication systems often use (heavy) quantization for signaling. This work puts forth a limited feedback-based SC framework. Similar to a prior work, a generative adversarial network (GAN)-based deep prior is used in our framework for fending against heavy shadowing. However, instead of using real-valued feedback, a random quantization strategy is adopted and a maximum likelihood estimation (MLE) criterion is proposed. Analysis shows that the MLE provably recovers the radio map, under reasonable conditions. Simulations are conducted to showcase the effectiveness of the proposed approach. Subash Timilsina, Sagar Shrestha, Xiao Fu 0001 |
ICASSP | 3 |
| 2023 | Deep Learning From Crowdsourced Labels: Coupled Cross-Entropy Minimization, Identifiability, and Regularization
Shahana Ibrahim, Tri Nguyen 0004, Xiao Fu 0001 |
ICLR | 3 |
| 2023 | Under-Counted Tensor Completion with Neural Incorporation of AttributesabstractSystematic under-counting effects are observed in data collected across many disciplines, e.g., epidemiology and ecology. Under-counted tensor completion (UC-TC) is well-motivated for many data analytics tasks, e.g., inferring the case numbers of infectious diseases at unobserved locations from under-counted case numbers in neighboring regions. However, existing methods for similar problems often lack supports in theory, making it hard to understand the underlying principles and conditions beyond empirical successes. In this work, a low-rank Poisson tensor model with an expressive unknown nonlinear side information extractor is proposed for under-counted multi-aspect data. A joint low-rank tensor completion and neural network learning algorithm is designed to recover the model. Moreover, the UC-TC formulation is supported by theoretical analysis showing that the fully counted entries of the tensor and each entry's under-counting probability can be provably recovered from partial observations---under reasonable conditions. To our best knowledge, the result is the first to offer theoretical supports for under-counted multi-aspect data completion. Simulations and real-data experiments corroborate the theoretical claims. Shahana Ibrahim, Xiao Fu 0001, Rebecca A. Hutchinson, Eugene Seo 0001 |
ICML | 2 |
| 2023 | Deep Clustering with Incomplete Noisy Pairwise Annotations: A Geometric Regularization ApproachabstractThe recent integration of deep learning and pairwise similarity annotation-based constrained clustering—i.e., deep constrained clustering (DCC)—has proven effective for incorporating weak supervision into massive data clustering: Less than 1% of pair similarity annotations can often substantially enhance the clustering accuracy. However, beyond empirical successes, there is a lack of understanding of DCC. In addition, many DCC paradigms are sensitive to annotation noise, but performance-guaranteed noisy DCC methods have been largely elusive. This work first takes a deep look into a recently emerged logistic loss function of DCC, and characterizes its theoretical properties. Our result shows that the logistic DCC loss ensures the identifiability of data membership under reasonable conditions, which may shed light on its effectiveness in practice. Building upon this understanding, a new loss function based on geometric factor analysis is proposed to fend against noisy annotations. It is shown that even under unknown annotation confusions, the data membership can still be provably identified under our proposed learning criterion. The proposed approach is tested over multiple datasets to validate our claims. Tri Nguyen 0004, Shahana Ibrahim, Xiao Fu 0001 |
ICML | 3 |
| 2023 | Dual Acoustic Linguistic Self-supervised Representation Learning for Cross-Domain Speech Recognition
Dianwen Ng, Chong Zhang 0003, Xiao Fu 0001, Wei Xi 0003, Chongjia Ni, Chng Eng Siong, Bin Ma 0001, Jizhong Zhao |
INTERSPEECH | 4 |
| 2023 | Finite-Sample Analysis of Deep CCA-Based Unsupervised Post-Nonlinear Multimodal LearningabstractCanonical correlation analysis (CCA) has been essential in unsupervised multimodal/multiview latent representation learning and data fusion. Classic CCA extracts shared information from multiple modalities of data using linear transformations. In recent years, deep neural networks-based nonlinear feature extractors were combined with CCA to come up with new variants, namely the ``DeepCCA'' line of work. These approaches were shown to have enhanced performance in many applications. However, theoretical supports of DeepCCA are often lacking. To address this challenge, the recent work of Lyu and Fu (2020) showed that, under a reasonable postnonlinear generative model, a carefully designed DeepCCA criterion provably removes unknown distortions in data generation and identifies the shared information across modalities. Nonetheless, a critical assumption used by Lyu and Fu (2020) for identifiability analysis was that unlimited data is available, which is unrealistic. This brief paper puts forth a finite-sample analysis of the DeepCCA method by Lyu and Fu (2020). The main result is that the finite-sample version of the method can still estimate the shared information with a guaranteed accuracy when the number of samples is sufficiently large. Our analytical approach is a nontrivial integration of statistical learning, numerical differentiation, and robust system identification, which may be of interest beyond the scope of DeepCCA and benefit other unsupervised learning paradigms. Qi Lyu, Xiao Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Communication-Efficient Distributed MAX-VAR Generalized CCA via Error Feedback-Assisted QuantizationabstractGeneralized canonical correlation analysis (GCCA) aims to learn common low-dimensional representations from multiple "views" of the data (e.g., audio and video of the same event). In the era of big data, GCCA computation encounters many new challenges. In particular, distributed optimization for GCCA—which is well-motivated in applications like internet of things and parallel computing—may incur prohibitively high communication costs. To address this challenge, this work proposes a communication-efficient distributed GCCA algorithm under the popular MAX-VAR GCCA paradigm. A quantization strategy for information exchange among the computing agents is employed in the proposed algorithm. It is observed that our design, leveraging the idea of error feedback-based quantization, can reduce communication cost by at least 90% while maintaining essentially the same GCCA performance as the unquantized version. Furthermore, the proposed method is guar-anteed to converge to a neighborhood of the optimal solution in a geometric rate—even under aggressive quantization. The effectiveness of our method is demonstrated using both synthetic and real data experiments. Sagar Shrestha, Xiao Fu 0001 |
ICASSP | 2 |
| 2022 | Understanding Latent Correlation-Based Multiview Learning and Self-Supervision: An Identifiability Perspective
Qi Lyu, Xiao Fu 0001, Songtao Lu |
ICLR | 2 |
| 2022 | On Finite-Sample Identifiability of Contrastive Learning-Based Nonlinear Independent Component AnalysisabstractNonlinear independent component analysis (nICA) aims at recovering statistically independent latent components that are mixed by unknown nonlinear functions. Central to nICA is the identifiability of the latent components, which had been elusive until very recently. Specifically, Hyvärinen et al. have shown that the nonlinearly mixed latent components are identifiable (up to often inconsequential ambiguities) under a generalized contrastive learning (GCL) formulation, given that the latent components are independent conditioned on a certain auxiliary variable. The GCL-based identifiability of nICA is elegant, and establishes interesting connections between nICA and popular unsupervised/self-supervised learning paradigms in representation learning, causal learning, and factor disentanglement. However, existing identifiability analyses of nICA all build upon an unlimited sample assumption and the use of ideal universal function learners—which creates a non-negligible gap between theory and practice. Closing the gap is a nontrivial challenge, as there is a lack of established “textbook” routine for finite sample analysis of such unsupervised problems. This work puts forth a finite-sample identifiability analysis of GCL-based nICA. Our analytical framework judiciously combines the properties of the GCL loss function, statistical generalization analysis, and numerical differentiation. Our framework also takes the learning function’s approximation error into consideration, and reveals an intuitive trade-off between the complexity and expressiveness of the employed function learner. Numerical experiments are used to validate the theorems. Qi Lyu, Xiao Fu 0001 |
ICML | 2 |
| 2022 | Provable Subspace Identification Under Post-Nonlinear MixturesabstractUnsupervised mixture learning (UML) aims at identifying linearly or nonlinearly mixed latent components in a blind manner. UML is known to be challenging: Even learning linear mixtures requires highly nontrivial analytical tools, e.g., independent component analysis or nonnegative matrix factorization. In this work, the post-nonlinear (PNL) mixture model---where {\it unknown} element-wise nonlinear functions are imposed onto a linear mixture---is revisited. The PNL model is widely employed in different fields ranging from brain signal classification, speech separation, remote sensing, to causal discovery. To identify and remove the unknown nonlinear functions, existing works often assume different properties on the latent components (e.g., statistical independence or probability-simplex structures). This work shows that under a carefully designed UML criterion, the existence of a nontrivial {\it null space} associated with the underlying mixing system suffices to guarantee identification/removal of the unknown nonlinearity. Compared to prior works, our finding largely relaxes the conditions of attaining PNL identifiability, and thus may benefit applications where no strong structural information on the latent components is known. A finite-sample analysis is offered to characterize the performance of the proposed approach under realistic settings. To implement the proposed learning criterion, a block coordinate descent algorithm is proposed. A series of numerical experiments corroborate our theoretical claims. Qi Lyu, Xiao Fu 0001 |
NeurIPS | 2 |
| 2022 | Hyperspectral Denoising Using Unsupervised Disentangled Spatiospectral Deep PriorsabstractImage denoising is often empowered by accurate prior information. In recent years, data-driven neural network priors have shown promising performance for RGB natural image denoising. Compared to classic handcrafted priors (e.g., sparsity and total variation), the “deep priors” are learned using a large number of training samples, which can accurately model the complex image generating process. However, data-driven priors are hard to acquire for hyperspectral images (HSIs) due to the lack of training data. A remedy is to use the so-called unsupervised deep image prior (DIP). Under the unsupervised DIP framework, it is hypothesized and empirically demonstrated that proper neural network structures are reasonable priors of certain types of images, and the network weights can be learned without training data. Nonetheless, the most effective unsupervised DIP structures were proposed for natural images instead of HSIs. The performance of unsupervised DIP-based HSI denoising is limited by a couple of serious challenges, namely network structure design and network complexity. This work puts forth an unsupervised DIP framework that is based on the classic spatiospectral decomposition of HSIs. Utilizing the so-called linear mixture model of HSIs, two types of unsupervised DIPs, that is, U-Net-like network and fully connected networks, are employed to model the abundance maps and endmembers contained in the HSIs, respectively. This way, empirically validated unsupervised DIP structures for natural images can be easily incorporated for HSI denoising. Besides, the decomposition also substantially reduces network complexity. An efficient alternating optimization algorithm is proposed to handle the formulated denoising problem. Simulated and real data experiments are employed to showcase the effectiveness of the proposed approach. Yu-Chun Miao, Xi-Le Zhao, Xiao Fu 0001, Jian-Li Wang, Yu-Bang Zheng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | StatEcoNet: Statistical Ecology Neural Networks for Species Distribution ModelingabstractThis paper focuses on a core task in computational sustainability and statistical ecology: species distribution modeling (SDM). In SDM, the occurrence pattern of a species on a landscape is predicted by environmental features based on observations at a set of locations. At first, SDM may appear to be a binary classification problem, and one might be inclined to employ classic tools (e.g., logistic regression, support vector machines, neural networks) to tackle it. However, wildlife surveys introduce structured noise (especially under-counting) in the species observations. If unaccounted for, these observation errors systematically bias SDMs. To address the unique challenges of SDM, this paper proposes a framework called StatEcoNet. Specifically, this work employs a graphical generative model in statistical ecology to serve as the skeleton of the proposed computational framework and carefully integrates neural networks under the framework. The advantages of StatEcoNet over related approaches are demonstrated on simulated datasets as well as bird species data. Since SDMs are critical tools for ecological science and natural resource management, StatEcoNet may offer boosted computational and analytical powers to a wide range of applications that have significant social impacts, e.g., the study and conservation of threatened species. Eugene Seo 0001, Rebecca A. Hutchinson, Xiao Fu 0001, Chelsea Li, Tyler A. Hallman, John Kilbride, W. Douglas Robinson |
AAAI | 3 |
| 2021 | Learning Mixed Membership from Adjacency Graph Via Systematic Edge Query: Identifiability and AlgorithmabstractGraph clustering is a core technique for network analysis problems, e.g., community detection. This work puts forth a node clustering approach for largely incomplete adjacency graphs. Under the considered scenario, instead of having access to the complete graph, only a small amount of queries about the graph edges can be made for node clustering. This task is well-motivated in many large-scale network analysis problems, where complete graph acquisition is prohibitively costly. Prior work tackles this problem under the setting that the nodes only admit single membership and the clusters are disjoint, yet multiple membership nodes and overlapping clusters often arise in practice. Existing approaches also rely on random edge query patterns and convex optimization-based formulations, which give rise to a number of implementation and scalability challenges. This work offers a framework that provably learns the mixed membership of nodes from overlapping clusters using limited edge information. Our method is equipped with a systematic edge query pattern, which is arguably easier to implement relative to the random counterparts in certain applications, e.g., field survey based graph analysis. A lightweight scalable algorithm is proposed, and its performance characterizations are presented. Numerical experiments are used to showcase the effectiveness of our method. Shahana Ibrahim, Xiao Fu 0001 |
ICASSP | 2 |
| 2021 | Fiber-Sampled Stochastic Mirror Descent for Tensor Decomposition with β-DivergenceabstractCanonical polyadic decomposition (CPD) has been a workhorse for multimodal data analytics. This work puts forth a stochastic algorithmic framework for CPD under β-divergence, which is well-motivated in statistical learning—where the Euclidean distance is typically not preferred. Despite the existence of a series of prior works addressing this topic, pressing computational and theoretical challenges, e.g., scalability and convergence issues, still remain. In this paper, a unified stochastic mirror descent framework is developed for large-scale β-divergence CPD. Our key contribution is the integrated design of a tensor fiber sampling strategy and a flexible stochastic Bregman divergence-based mirror descent iterative procedure, which significantly reduces the computation and memory cost per iteration for various β. Leveraging the fiber sampling scheme and the multilinear algebraic structure of low-rank tensors, the proposed lightweight algorithm also ensures global convergence to a stationary point under mild conditions. Numerical results on synthetic and real data show that our framework attains significant computational saving compared with state-of-the-art methods. Wenqiang Pu, Shahana Ibrahim, Xiao Fu 0001, Mingyi Hong 0001 |
ICASSP | 3 |
| 2021 | Deep Generative Model Learning For Blind Spectrum Cartography with NMF-Based Radio Map DisaggregationabstractSpectrum cartography (SC) aims at estimating the multi-aspect (e.g., space, frequency, and time) interference level caused by multiple emitters from limited measurements. Early SC approaches rely on model assumptions about the radio map, e.g., sparsity and smoothness, which may be grossly violated under critical scenarios, e.g., in the presence of severe shadowing. More recent data-driven methods train deep generative networks to distill parsimonious representations of complex scenarios, in order to enhance performance of SC. The challenge is that the state space of this learning problem is extremely large—induced by different combinations of key problem constituents, e.g., the number of emitters, the emitters’ carrier frequencies, and the emitter locations. Learning over such a huge space can be costly in terms of sample complexity and training time; it also frequently leads to generalization problems. Our method integrates the favorable traits of model and data-driven approaches, which substantially ‘shrinks’ the state space. Specifically, the proposed learning paradigm only needs to learn a generative model for the radio map of a single emitter (as opposed to numerous combinations of multiple emitters), leveraging a nonnegative matrix factorization (NMF)-based emitter disaggregation process. Numerical evidence shows that the proposed method outperforms state-of-the-art purely model-driven and purely data-driven approaches. Sagar Shrestha, Xiao Fu 0001, Mingyi Hong 0001 |
ICASSP | 2 |
| 2021 | Learning to Continuously Optimize Wireless Resource in Episodically Dynamic EnvironmentabstractThere has been a growing interest in developing data-driven, in particular deep neural network (DNN) based methods for modern communication tasks. For a few popular tasks such as power control, beamforming, and MIMO detection, these methods achieve state-of-the-art performance while requiring less computational efforts, less channel state information (CSI), etc. However, it is often challenging for these approaches to learn in a dynamic environment where parameters such as CSIs keep changing.This work develops a methodology that enables data-driven methods to continuously learn and optimize in a dynamic environment. Specifically, we consider an "episodically dynamic" setting where the environment changes in "episodes", and in each episode the environment is stationary. We propose a continual learning (CL) framework for wireless systems, which can incrementally adapt the learning models to the new episodes, without forgetting models learned from the previous episodes. Our design is based on a novel min-max formulation which ensures certain "fairness" across different episodes. Finally, we demonstrate the effectiveness of the CL approach by customizing it to a popular DNN based model for power control, and testing using both synthetic and real data. Wenqiang Pu, Minghe Zhu, Xiao Fu 0001, Tsung-Hui Chang, Mingyi Hong 0001 |
ICASSP | 4 |
| 2021 | Crowdsourcing via Annotator Co-occurrence Imputation and Provable Symmetric Nonnegative Matrix FactorizationabstractUnsupervised learning of the Dawid-Skene (D&S) model from noisy, incomplete and crowdsourced annotations has been a long-standing challenge, and is a critical step towards reliably labeling massive data. A recent work takes a coupled nonnegative matrix factorization (CNMF) perspective, and shows appealing features: It ensures the identifiability of the D&S model and enjoys low sample complexity, as only the estimates of the co-occurrences of annotator labels are involved. However, the identifiability holds only when certain somewhat restrictive conditions are met in the context of crowdsourcing. Optimizing the CNMF criterion is also costly—and convergence assurances are elusive. This work recasts the pairwise co-occurrence based D&S model learning problem as a symmetric NMF (SymNMF) problem—which offers enhanced identifiability relative to CNMF. In practice, the SymNMF model is often (largely) incomplete, due to the lack of co-labeled items by some annotators. Two lightweight algorithms are proposed for co-occurrence imputation. Then, a low-complexity shifted rectified linear unit (ReLU)-empowered SymNMF algorithm is proposed to identify the D&S model. Various performance characterizations (e.g., missing co-occurrence recoverability, stability, and convergence) and evaluations are also presented. Shahana Ibrahim, Xiao Fu 0001 |
ICML | 2 |
| 2021 | Multi-User Adaptive Video Delivery Over Wireless Networks: A Physical Layer Resource-Aware Deep Reinforcement Learning ApproachabstractIn this paper, we investigate the adaptive video delivery for multiple users over time-varying and mutually interfering multi-cell wireless networks. The key research challenge is to jointly design the physical-layer resource allocation scheme and application-layer rate adaptation logic, such that the users' long-term fair quality of experience (QoE) can be maximized. Due to the timescale mismatch between these two layers and the asynchrony of user requests, however, it is difficult to directly model the cross-layer stochastic control problem by using a reinforcement learning framework. To address this difficulty, we propose a novel two-level decision framework where an optimization-based beamforming scheme (performed at the base stations) and a deep reinforcement learning (DRL)-based rate adaptation scheme (performed at the user terminals) are, respectively, developed, such that a highly complex long-term multi-user QoE fairness problem is decomposed into some relatively simple problems and solved effectively. Our strategy represents a significant departure from the existing schemes with consideration of either a short-term multi-user QoE maximization or a long-term single-user point-to-point QoE maximization. Extensive simulations demonstrate that the proposed cross-layer design is effective and promising. Kexin Tang, Nuowen Kan, Junni Zou, Xiao Fu 0001, Mingyi Hong 0001, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Link Prediction Under Imperfect Detection: Collaborative Filtering for Ecological NetworksabstractMatrix completion based collaborative filtering is considered scalable and effective for online service link prediction (e.g., movie recommendation) but does not meet the challenges of link prediction in ecological networks. A unique challenge of ecological networks is that the observed data are subject to systematic imperfect detection, due to the difficulty of accurate field sampling. In this work, we propose a new framework customized for ecological bipartite network link prediction. Our approach starts with incorporating the Poisson N-mixture model, a widely used framework in statistical ecology for modeling imperfect detection of a single species in field sampling. Despite its extensive use for single species analysis, this model has never been considered for link prediction between different species, perhaps because of the complex nature of both link prediction and N-mixture model inference. By judiciously combining the Poisson N-mixture model with a probabilistic nonnegative matrix factorization (NMF) model in latent space, we propose an intuitive statistical model for the problem of interest. We also offer a scalable and convergence-guaranteed optimization algorithm to handle the associated maximum likelihood identification problem. Experimental results on synthetic data and two real-world ecological networks data are employed to validate our proposed approach. Xiao Fu 0001, Eugene Seo 0001, Justin Clarke, Rebecca A. Hutchinson |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | Low-Complexity Levenberg-Marquardt Algorithm for Tensor Canonical Polyadic DecompositionabstractIn this paper, we propose CPD-fLM++, a fast implementation of the Levenberg-Marquardt (LM) algorithm for the tensor canonical polyadic decomposition. The overall algorithmic framework follows exactly the LM approach, which enjoys locally a super-linear convergence rate and has been observed to be able to avoid the “swamp” effect commonly seen in other CPD algorithms. However, unlike the common wisdom that LM requires very high per-iteration complexity to execute, we show each iteration of CPD-fLM++ only requires O(NK6) flops to perform matrix inversions, where N is the number of modes and K is the target CPD rank. This is done by carefully exploiting the structures in the Jacobian Gramian matrix, and is by far the most efficient implementation of LM for CPD using direct methods. Experiments on synthetic data confirm the good performance of CPD-fLM++ for large-scale higher-order tensors. Kejun Huang, Xiao Fu 0001 |
ICASSP | 2 |
| 2020 | On Recoverability of Randomly Compressed Tensors With Low CP RankabstractOur interest lies in the recoverability properties of compressed tensors under the canonical polyadic decomposition (CPD) model. The considered problem is well-motivated in many applications, e.g., hyperspectral image and video compression. Prior work studied this problem under a variety of assumptions, e.g., that the latent factors of the tensor are sparse and that the compressing matrix follows a joint absolutely continuous distribution. These results leverage analytical tools such as CPD uniqueness and algebraic geometry-which are elegant. In this work, we offer an alternative result: We show that if the tensor is compressed by a subgaussian linear mapping, then the tensor is recoverable if the number of measurements is on the same order of magnitude as that of the model parameters. Unlike existing results, our proof is based on deriving a restricted isometry property (R.I.P.) under the CPD model via set covering techniques, and thus exhibits a flavor of classic compressive sensing. The new recoverability result enriches the understanding to the compressed CP tensor recovery problem. It offers theoretical guarantees for recovering tensors whose elements are not necessarily sparse; the compressing matrix is also not restricted to continuous matrices under our framework. The newly derived covering number for tensors with low CP rank may also benefit future research, e.g., sketching based tensor compression for reducing computational burden. Shahana Ibrahim, Xiao Fu 0001, Xingguo Li |
IEEE Signal Process. Lett. | 2 |
| 2020 | Hyperspectral Super-Resolution via Global-Local Low-Rank Matrix EstimationabstractHyperspectral super-resolution (HSR) is a problem that aims to estimate an image of high spectral and spatial resolutions from a pair of coregistered multispectral (MS) and hyperspectral (HS) images, which have coarser spectral and spatial resolutions, respectively. In this article, we pursue a lowrank matrix estimation approach for HSR. We assume that the spectral-spatial matrices associated with the whole image and the local areas of the image have low-rank structures. The local low-rank assumption, in particular, has the aim of providing a more flexible model for accounting for local variation effects due to endmember variability. We formulate the HSR problem as a global-local rank-regularized least-squares problem. By leveraging on the recent advances in nonconvex large-scale optimization, namely the smooth Schatten-p approximation and the accelerated majorization-minimization method, we develop an efficient algorithm for the global-local low-rank problem. Numerical experiments on synthetic, semi-real, and real data show that the proposed algorithm outperforms a number of benchmark algorithms in terms of recovery performance. Ruiyuan Wu, Wing-Kin Ma, Xiao Fu 0001, Qiang Li 0017 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Block-randomized Stochastic Proximal Gradient for Constrained Low-rank Tensor FactorizationabstractThis work focuses on canonical polyadic decomposition (CPD) for large-scale tensors. Many prior works rely on data sparsity to develop scalable CPD algorithms, which are not suitable for handling dense tensor, while dense tensors often arise in applications such as image and video processing. As an alternative, stochastic algorithms utilize data sampling to reduce per-iteration complexity and thus are very scalable, even when handling dense tensors. However, existing stochastic CPD algorithms are facing some challenges. For example, some algorithms are based on randomly sampled tensor entries, and thus each iteration can only updates a small portion of the latent factors. This may result in slow improvement of the estimation accuracy of the latent factors. In addition, the convergence properties of many stochastic CPD algorithms are unclear, perhaps because CPD poses a hard nonconvex problem and is challenging for analysis under stochastic settings. In this work, we propose a stochastic optimization strategy that can effectively circumvent the above challenges. The proposed algorithm updates a whole latent factor at each iteration using sampled fibers of a tensor, which can quickly increase the estimation accuracy. The algorithm is flexible-many commonly used regularizers and constraints can be easily incorporated in the computational framework. The algorithm is also backed by a rigorous convergence theory. Simulations on large-scale dense tensors are employed to showcase the effectiveness of the algorithm. Xiao Fu 0001, Hoi-To Wai, Kejun Huang |
ICASSP | 1 |
| 2019 | Regular Sampling of Tensor Signals: Theory and Application to FMRIabstractSampling lies at the heart of signal processing. The celebrated Shan-non - Nyquist theorem states that in order to reconstruct a continuous or discrete time signal from uniform samples one must sample at a rate twice the highest frequency present in the signal. Numerous signals and images of interest, however, are not even approximately bandlimited. While much progress has happened in recent years, reconstruction from sub-Nyquist samples still hinges on the use of random / incoherent (aggregate) sampling patterns, instead of uniform or regular sampling, which is far more simple, practical, and natural in many applications. In this work, we study regular sampling and reconstruction of three- or higher-dimensional signals (tensors). We prove that exact tensor reconstruction from regular samples is feasible under mild conditions on the rank of the tensor. Furthermore we cast the functional magnetic resonance imaging (fMRI) acceleration task as a regular tensor sampling problem and provide an algorithmic framework that effectively handles the reconstruction task. Experiments based on synthetic data and real fMRI data showcase the effectiveness of our approach. Charilaos I. Kanatsoulis, Nicholas D. Sidiropoulos, Mehmet Akçakaya, Xiao Fu 0001 |
ICASSP | 4 |
| 2019 | Detecting Overlapping and Correlated Communities without Pure Nodes: Identifiability and AlgorithmabstractMany machine learning problems come in the form of networks with relational data between entities, and one of the key unsupervised learning tasks is to detect communities in such a network. We adopt the mixed-membership stochastic blockmodel as the underlying probabilistic model, and give conditions under which the memberships of a subset of nodes can be uniquely identified. Our method starts by constructing a second-order graph moment, which can be shown to converge to a specific product of the true parameters as the size of the network increases. To correctly recover the true membership parameters, we formulate an optimization problem using insights from convex geometry. We show that if the true memberships satisfy a so-called sufficiently scattered condition, then solving the proposed problem correctly identifies the ground truth. We also propose an efficient algorithm for detecting communities, which is significantly faster than prior work and with better convergence properties. Experiments on synthetic and real data justify the validity of the proposed learning framework for network data. Kejun Huang, Xiao Fu 0001 |
ICML | 2 |
| 2019 | Crowdsourcing via Pairwise Co-occurrences: Identifiability and AlgorithmsabstractThe data deluge comes with high demands for data labeling. Crowdsourcing (or, more generally, ensemble learning) techniques aim to produce accurate labels via integrating noisy, non-expert labeling from annotators. The classic Dawid-Skene estimator and its accompanying expectation maximization (EM) algorithm have been widely used, but the theoretical properties are not fully understood. Tensor methods were proposed to guarantee identification of the Dawid-Skene model, but the sample complexity is a hurdle for applying such approaches---since the tensor methods hinge on the availability of third-order statistics that are hard to reliably estimate given limited data. In this paper, we propose a framework using pairwise co-occurrences of the annotator responses, which naturally admits lower sample complexity. We show that the approach can identify the Dawid-Skene model under realistic conditions. We propose an algebraic algorithm reminiscent of convex geometry-based structured matrix factorization to solve the model identification problem efficiently, and an identifiability-enhanced algorithm for handling more challenging and critical scenarios. Experiments show that the proposed algorithms outperform the state-of-art algorithms under a variety of scenarios. Shahana Ibrahim, Xiao Fu 0001, Nikos Kargas, Kejun Huang |
NeurIPS | 2 |
| 2019 | Multiuser Video Streaming Rate Adaptation: A Physical Layer Resource-Aware Deep Reinforcement Learning ApproachabstractIn this paper, we propose a cross-layer decision framework for multiuser adaptive video delivery over time-varying and mutually interfering wireless cellular network. The key idea is to synthetically design the physical-layer optimization-based beamforming scheme (performed at the base stations) and the application-layer deep reinforcement learning (DRL)-based rate adaptation scheme (performed at the user terminals), so that a very complex multi-user overall fair long-term quality of experience (QoE) maximization problem can be decomposed to two layers and solved effectively. Extensive simulations show that the proposed cross-layer design is effective and promising. Kexin Tang, Nuowen Kan, Junni Zou, Xiao Fu 0001, Mingyi Hong 0001, Hongkai Xiong |
VCIP | 4 |
| 2019 | Anchor-Free Correlated Topic ModelingabstractIn topic modeling, identifiability of the topics is an essential issue. Many topic modeling approaches have been developed under the premise that each topic has a characteristic anchor word that only appears in that topic. The anchor-word assumption is fragile in practice, because words and terms have multiple uses; yet it is commonly adopted because it enables identifiability guarantees. Remedies in the literature include using three- or higher-order word co-occurence statistics to come up with tensor factorization models, but such statistics need many more samples to obtain reliable estimates, and identifiability still hinges on additional assumptions, such as consecutive words being persistently drawn from the same topic. In this work, we propose a new topic identification criterion using second order statistics of the words. The criterion is theoretically guaranteed to identify the underlying topics even when the anchor-word assumption is grossly violated. An algorithm based on alternating optimization, and an efficient primal-dual algorithm are proposed to handle the resulting identification problem. The former exhibits high performance and is completely parameter-free; the latter affords up to 200 times speedup relative to the former, but requires step-size tuning and a slight sacrifice in accuracy. A variety of real text copora are employed to showcase the effectiveness of the approach, where the proposed anchor-free method demonstrates substantial improvements compared to a number of anchor-word based approaches under various evaluation metrics. Xiao Fu 0001, Kejun Huang, Nicholas D. Sidiropoulos, Qingjiang Shi, Mingyi Hong 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Amplitude Retrieval for Channel Estimation of MIMO Systems With One-Bit ADCsabstractThis letter revisits the channel estimation problem for MIMO systems with one-bit analog-to-digital converters (ADCs) through a novel algorithm-Amplitude Retrieval (AR). Unlike the state-of-the-art methods such as those based on one-bit compressive sensing, AR takes a different approach. It accounts for the lost amplitudes of the one-bit quantized measurements, and performs channel estimation and amplitude completion jointly. This way, the direction information of the propagation paths can be estimated via accurate direction finding algorithms in array processing, e.g., maximum likelihood. The upsot is that AR is able to handle off-grid angles and provide more accurate channel estimates. Simulation results are included to showcase the advantages of AR. Cheng Qian 0001, Xiao Fu 0001, Nicholas D. Sidiropoulos |
IEEE Signal Process. Lett. | 2 |
| 2019 | Efficient and Distributed Generalized Canonical Correlation Analysis for Big Multiview DataabstractGeneralized canonical correlation analysis (GCCA) integrates information from data samples that are acquired at multiple feature spaces (or `views') to produce low-dimensional representations-which is an extension of classical two-view CCA. Since the 1960s, (G)CCA has attracted much attention in statistics, machine learning, and data mining because of its importance in data analytics. Despite these efforts, the existing GCCA algorithms have serious complexity issues. The memory and computational complexities of the existing algorithms usually grow as a quadratic and cubic function of the problem dimension (the number of samples / features), respectively-e.g., handling views with ≈1,000 features using such algorithms already occupies ≈106memory and the periteration complexity is ≈109flops-which makes it hard to push these methods much further. To circumvent such difficulties, we first propose a GCCA algorithm whose memory and computational costs scale linearly in the problem dimension and the number of nonzero data elements, respectively. Consequently, the proposed algorithm can easily handle very large sparse views whose sample and feature dimensions both exceed 100,000. Our second contribution lies in proposing two distributed algorithms for GCCA, which compute the canonical components of different views in parallel and thus can further reduce the runtime significantly if multiple computing agents are available. We provide detailed convergence analyses of the proposed algorithms and show that all the largescale GCCA algorithms converge to a Karush-Kuhn-Tucker (KKT) point at least sublinearly. Judiciously designed synthetic and realdata experiments are employed to showcase the effectiveness of the proposed algorithms. Xiao Fu 0001, Kejun Huang, Evangelos E. Papalexakis, Hyun Ah Song, Partha P. Talukdar, Nicholas D. Sidiropoulos, Christos Faloutsos, Tom M. Mitchell |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | On Convergence of Epanechnikov Mean ShiftabstractEpanechnikov Mean Shift is a simple yet empirically very effective algorithm for clustering. It localizes the centroids of data clusters via estimating modes of the probability distribution that generates the data points, using the "optimal" Epanechnikov kernel density estimator. However, since the procedure involves non-smooth kernel density functions,the convergence behavior of Epanechnikov mean shift lacks theoretical support as of this writing---most of the existing analyses are based on smooth functions and thus cannot be applied to Epanechnikov Mean Shift. In this work, we first show that the original Epanechnikov Mean Shift may indeed terminate at a non-critical point, due to the non-smoothness nature. Based on our analysis, we propose a simple remedy to fix it. The modified Epanechnikov Mean Shift is guaranteed to terminate at a local maximum of the estimated density, which corresponds to a cluster centroid, within a inite number of iterations. We also propose a way to avoid running the Mean Shift iterates from every data point, while maintaining good clustering accuracies under non-overlapping spherical Gaussian mixture models. This further pushes Epanechnikov Mean Shift to handle very large and high-dimensional data sets. Experiments show surprisingly good performance compared to the Lloyd's K-means algorithm and the EM algorithm. Kejun Huang, Xiao Fu 0001, Nicholas D. Sidiropoulos |
AAAI | 2 |
| 2018 | Tensor-Based Parameter Estimation of Double Directional Massive Mimo Channel with Dual-Polarized AntennasabstractThe 3GPP suggests to combine dual polarized (DP) antenna arrays with the double directional (DD) channel model for downlink channel estimation. This combination strikes a good balance between high-capacity communications and parsimonious channel modeling, and also brings limited feedback schemes for downlink channel estimation within reach. However, most existing channel estimation work under the DD model has not considered DP arrays, perhaps because of the complex array manifold and the resulting difficulty in algorithm design. In this paper, we first reveal that the DD channel with DP arrays at the transmitter and receiver can be naturally modeled as a low-rank four-way tensor, and thus the parameters can be effectively estimated via tensor decomposition algorithms. To reduce computational complexity, we show that the problem can be recast as a four-snapshot three-dimensional harmonic retrieval problem, which can be solved using computationally efficient subspace methods. On the theory side, we show that the DD channel with DP arrays is identifiable under very mild conditions, leveraging identifiability of low-rank tensors. Numerical simulations are employed to showcase the effectiveness of our methods. Cheng Qian 0001, Xiao Fu 0001, Nicholas D. Sidiropoulos |
ICASSP | 2 |
| 2018 | Shared Human-Machine Control for Self-Aware ProsthesesabstractThis paper presents a framework for shared, human-machine control of a prosthetic arm. The method employs electromyogram and peripheral neural signals to decode motor intent, and incorporates a higher-level goal in the controller to augment human effort. The controller derivation employs Markov Decision Processes. The system is trained using a gradient ascent approach in which the policy is parameterized using a Kalman Filter and the goal is incorporated by adapting the Kalman filter output online. Results of experimental performance analysis of the shared controller when the goal information is imperfect are presented in the paper. These results, obtained from an amputee subject and a subject with intact arms, demonstrate that a system controlled by the human user and the machine together exhibit better performance than systems employing machine-only or human-only control. Henrique Dantas, Jacob Nieven, Tyler S. Davis, Xiao Fu 0001, Gregory A. Clark, David J. Warren, V. John Mathews |
ICASSP | 4 |
| 2018 | Large-Scale Regularized Sumcor GCCA via Penalty-Dual DecompositionabstractThe sum-of-correlations (SUMCOR) generalized canonical correlation analysis (GCCA) aims at producing low-dimensional representations of multiview data via enforcing pairwise similarity of the reduced-dimension views. SUMCOR has been applied to a large variety of applications including blind separation, multilingual word embedding, and cross-modality retrieval. Despite the NP-hardness of SUMCOR, recent work has proposed effective algorithms for handling it at very large scale. However, the existing scalable algorithms are not easy to extend to incorporate structural regularization and prior information - which are critical for real-world applications where outliers and modeling mismatches are present. In this work, we propose a new computational framework for large-scale SUMCOR GCCA. The algorithm can easily incorporate a suite of structural regularizers which are frequently used in data analytics, has lightweight updates and low memory complexity, and can be easily implemented in a parallel fashion. The proposed algorithm is also guaranteed to converge to a Karush-Kuhn-Tucker (KKT) point of the regularized SUMCOR problem. Carefully designed simulations are employed to demonstrate the effectiveness of the proposed algorithm. Charilaos I. Kanatsoulis, Xiao Fu 0001, Nicholas D. Sidiropoulos, Mingyi Hong 0001 |
ICASSP | 2 |
| 2018 | Hyperspectral Super-Resolution Via Coupled Tensor Factorization: Identifiability and AlgorithmsabstractThis work focuses on the problem of fusing a hyperspectral image (HSI) and a multispectral image (MSI) to produce a super-resolution image that admits high spatial and spectral resolutions. Existing algorithms are mostly based on joint low-rank factorization of the ma-tricized HSI and MSI. This framework is effective to some extent, but several challenges remain. First, it is unclear whether or not the super-resolution image is identifiable in theory under this framework, while identifiability usually plays an essential role in such estimation problems. Second, most algorithms assume that the degradation operators from the super-resolution image to the HSI and MSI are known or can be easily estimated - which is hardly true in practice. In this work, we propose a novel coupled tensor decomposition method that can effectively circumvent these issues. The proposed approach guarantees the identifiability of the super-resolution image under realistic conditions. The method can work even without knowing the spatial degradation operator, which could be hard to accurately estimate in practice. Simulations using AVIRIS Cuprite data are employed to demonstrate the effectiveness of the proposed approach. Charilaos I. Kanatsoulis, Xiao Fu 0001, Nicholas D. Sidiropoulos, Wing-Kin Ma |
ICASSP | 2 |
| 2018 | Hi, Bcd! Hybrid Inexact Block Coordinate Descent for Hyperspectral Super-ResolutionabstractHyperspectral super-resolution (HSR) is a problem of recovering a high-spectral-spatial-resolution image from a multispectral measurement and a hyperspectral measurement, which have low spectral and spatial resolutions, respectively. We consider a low-rank structured matrix factorization formulation for HSR, which is a non-convex large-scale optimization problem. Our contributions contain both computational and theoretical aspects. On the computational side, we develop three inexact block coordinate descent (BCD) schemes that are empirically found to run many times faster than a state-of-the-art method, which uses exact BCD. We achieve this by applying concepts in the proximal gradient (PG) and Frank-Wolfe (FW) methods and by exploiting the HSR problem structures. On the theoretical side, we show that these inexact BCD schemes guarantee convergence to a stationary point. In particular, the convergence result for a hybrid PG- FW inexact BCD scheme is new. Ruiyuan Wu, Chun-Hei Chan, Hoi-To Wai, Wing-Kin Ma, Xiao Fu 0001 |
ICASSP | 5 |
| 2018 | Hyperspectral Super-Resolution: Combining Low Rank Tensor and Matrix StructureabstractHyperspectral super-resolution refers to the task of fusing a hyperspectral image (HSI) and a multispectral image (MSI) in order to produce a super-resolution image (SRI) that has high spatial and spectral resolution. Popular methods leverage matrix factorization that models each spectral pixel as a convex combination of spectral signatures belonging to a few endmembers. These methods are considered state-of-the-art, but several challenges remain. First, multiband images are naturally three dimensional (3-d) signals, while matrix methods usually ignore the 3-d structure, which is prone to information losses. Second, these methods do not provide identifiability guarantees under which the reconstruction task is feasible. Third, a tacit assumption is that the degradation operators from SRI to MSI and HSI are known - which is hardly the case in practice. Recently [1], [2] proposed a coupled tensor factorization approach to handle these issues. In this work we propose a hybrid model that combines the benefits of tensor and matrix factorization approaches. We also develop a new algorithm that is mathematically simple, enjoys identifiability under relaxed conditions and is completely agnostic of the spatial degradation operator. Experimental results with real hyperspectral data showcase the effectiveness of the proposed approach. Charilaos I. Kanatsoulis, Xiao Fu 0001, Nicholas D. Sidiropoulos, Wing-Kin Ma |
ICIP | 2 |
| 2018 | Learning Hidden Markov Models from Pairwise Co-occurrences with Application to Topic ModelingabstractWe present a new algorithm for identifying the transition and emission probabilities of a hidden Markov model (HMM) from the emitted data. Expectation-maximization becomes computationally prohibitive for long observation records, which are often required for identification. The new algorithm is particularly suitable for cases where the available sample size is large enough to accurately estimate second-order output probabilities, but not higher-order ones. We show that if one is only able to obtain a reliable estimate of the pairwise co-occurrence probabilities of the emissions, it is still possible to uniquely identify the HMM if the emission probability is sufficiently scattered. We apply our method to hidden topic Markov modeling, and demonstrate that we can learn topics with higher quality if documents are modeled as observations of HMMs sharing the same emission (topic) probability, compared to the simple but widely used bag-of-words model. Kejun Huang, Xiao Fu 0001, Nicholas D. Sidiropoulos |
ICML | 2 |
| 2018 | On Identifiability of Nonnegative Matrix FactorizationabstractIn this letter, we propose a new identification criterion that guarantees the recovery of the low-rank latent factors in the nonnegative matrix factorization (NMF) generative model, under mild conditions. Specifically, using the proposed criterion, it suffices to identify the latent factors if the rows of one factor are sufficiently scattered over the nonnegative orthant, while no structural assumption is imposed on the other factor except being full-rank. This is by far the mildest condition under which the latent factors are provably identifiable from the NMF model. Xiao Fu 0001, Kejun Huang, Nicholas D. Sidiropoulos |
IEEE Signal Process. Lett. | 1 |
| 2017 | Scalable and flexible Max-Var generalized canonical correlation analysis via alternating optimizationabstractUnlike dimensionality reduction (DR) tools for single-view data, e.g., principal component analysis (PCA), canonical correlation analysis (CCA) and generalized CCA (GCCA) are able to integrate information from multiple feature spaces of data. This is critical in multi-modal data fusion and analytics, where samples from a single view may not be enough for meaningful DR. In this work, we focus on a popular formulation of GCCA, namely, MAX-VAR GCCA. The classic MAX-VAR problem is optimally solvable via eigen-decomposition, but this solution has serious scalability issues. In addition, how to impose regularizers on the sought canonical components was unclear - while structure-promoting regularizers are often desired in practice. We propose an algorithm that can easily handle datasets whose sample and feature dimensions are both large by exploiting data sparsity. The algorithm is also flexible in incorporating regularizers on the canonical components. Convergence properties of the proposed algorithm are carefully analyzed. Numerical experiments are presented to showcase its effectiveness. Xiao Fu 0001, Kejun Huang, Mingyi Hong 0001, Nicholas D. Sidiropoulos, Anthony Man-Cho So |
ICASSP | 1 |
| 2017 | A stochastic maximum-likelihood framework for simplex structured matrix factorizationabstractConsider a structured matrix factorizaton (SMF) whose coefficient vectors are constrained to lie in the unit simplex. This kind of simplex SMF (SSMF) has received growing attention and has found many applications such as hyperspectral unmixing in remote sensing, text mining in machine learning, and blind source separation in signal processing. The aim of this paper is to establish a maximum-likelihood (ML) estimation framework for SSMF in the presence of Gaussian noise and outliers, and to demonstrate its potential. Our ML formulation has the coefficient vectors marginalized in accordance with a prescribed probabilistic model, and this leads to a likelihood function that contains multi-dimensional integrals. Unfortunately these integrals do not appear to have analytically tractable solutions, and this makes the ML problem challenging. We tackle the problem by using sample average approximation in stochastic optimization and majorization-minimization. Simulation results show that the resulting ML algorithm significantly outperforms several existing methods when noise and outliers are present. Ruiyuan Wu, Wing-Kin Ma, Xiao Fu 0001 |
ICASSP | 3 |
| 2017 | Towards K-means-friendly Spaces: Simultaneous Deep Learning and ClusteringabstractMost learning approaches treat dimensionality reduction (DR) and clustering separately (i.e., sequentially), but recent research has shown that optimizing the two tasks jointly can substantially improve the performance of both. The premise behind the latter genre is that the data samples are obtained via linear transformation of latent representations that are easy to cluster; but in practice, the transformation from the latent space to the data can be more complicated. In this work, we assume that this transformation is an unknown and possibly nonlinear function. To recover the `clustering-friendly’ latent representations and to better cluster the data, we propose a joint DR and K-means clustering approach in which DR is accomplished via learning a deep neural network (DNN). The motivation is to keep the advantages of jointly optimizing the two tasks, while exploiting the deep neural network’s ability to approximate any nonlinear function. This way, the proposed approach can work well for a broad class of generative models. Towards this end, we carefully design the DNN structure and the associated joint optimization criterion, and propose an effective and scalable algorithm to handle the formulated optimization problem. Experiments using different real datasets are employed to showcase the effectiveness of the proposed approach. Bo Yang 0053, Xiao Fu 0001, Nicholas D. Sidiropoulos, Mingyi Hong 0001 |
ICML | 2 |
| 2017 | BrainZoom: High Resolution Reconstruction from Multi-modal Brain SignalsabstractHow close can we zoom in to observe brain activity? Our understanding is limited by the resolution of imaging modalities that exhibit good spatial but poor temporal resolution, or vice-versa. In this paper, we propose BrainZoom, an efficient imaging algorithm that cross-leverages multi-modal brain signals. BrainZoom (a) constructs high resolution brain images from multi-modal signals, (b) is scalable, and (c) is flexible in that it can easily incorporate various priors on the brain activities, such as sparsity, low rank, or smoothness. We carefully formulate the problem to tackle nonlinearity in the measurements (via variable splitting) and auto-scale between different modal signals, and judiciously design an inexact alternating optimization-based algorithmic framework to handle the problem with provable convergence guarantees. Our experiments using a popular realistic brain signal simulator to generate fMRI and MEG demonstrate that high spatio-temporal resolution brain imaging is possible from these two modalities. The experiments also suggest that smoothness seems to be the best prior, among several we tried. Xiao Fu 0001, Kejun Huang, Otilia Stretcu, Hyun Ah Song, Evangelos E. Papalexakis, Partha P. Talukdar, Tom M. Mitchell, Nicholas D. Sidiropoulos, Christos Faloutsos, Barnabás Póczos |
SDM | 1 |
| 2017 | Non-uniform directional dictionary-based limited feedback for massive MIMO systemsabstractThis work proposes a new limited feedback channel estimation framework. The proposed approach exploits a sparse representation of the double directional wireless channel model involving an over complete dictionary that accounts for the antenna directivity patterns at both base station (BS) and user equipment (UE). Under this sparse representation, a computationally efficient limited feedback algorithm that is based on single-bit compressive sensing is proposed to effectively estimate the downlink channel. The algorithm is lightweight in terms of computation, and suitable for real-time implementation in practical systems. More importantly, under our design, using a small number of feedback bits, very satisfactory channel estimation accuracy is achieved even when the number of BS antennas is very large, which makes the proposed scheme ideal for massive MIMO 5G cellular networks. Judiciously designed simulations reveal that the proposed algorithm outperforms a number of popular feedback schemes in terms of beam forming gain for subsequent downlink transmission, and reduces feedback overhead substantially when the BS has a large number of antennas. Panos N. Alevizos, Xiao Fu 0001, Nicholas D. Sidiropoulos, Aggelos Bletsas |
WiOpt | 2 |
| 2016 | Robust volume minimization-based matrix factorization via alternating optimizationabstractThis paper focuses on volume minimization (VolMin)-based structured matrix factorization (SMF), which factors a data matrix into a full-column rank basis and a coefficient matrix whose columns reside in the unit simplex. The VolMin criterion achieves this goal via finding a minimum-volume enclosing convex hull of the data. Recent works showed that VolMin guarantees the identifiability of the factor matrices under mild and realistic conditions, which suit many applications in signal processing and machine learning. However, the existing VolMin algorithms are sensitive to outliers or lack efficiency in dealing with volume-associated cost functions. In this work, we propose a new VolMin-based matrix factorization criterion and algorithm that take outliers into consideration. The proposed algorithm detects outliers and suppress them automatically, and it does so in an algorithmically very simple way. Simulations are used to showcase the effectiveness of the proposed algorithm. Xiao Fu 0001, Wing-Kin Ma, Kejun Huang, Nicholas D. Sidiropoulos |
ICASSP | 1 |
| 2016 | Efficient and Distributed Algorithms for Large-Scale Generalized Canonical Correlations AnalysisabstractGeneralized canonical correlation analysis (GCCA) aims at extracting common structure from multiple 'views', i.e., high-dimensional matrices representing the same objects in different feature domains – an extension of classical two-view CCA. Existing (G)CCA algorithms have serious scalability issues, since they involve square root factorization of the correlation matrices of the views. The memory and computational complexity associated with this step grow as a quadratic and cubic function of the problem dimension (the number of samples / features), respectively. To circumvent such difficulties, we propose a GCCA algorithm whose memory and computational costs scale linearly in the problem dimension and the number of nonzero data elements, respectively. Consequently, the proposed algorithm can easily handle very large sparse views whose sample and feature dimensions both exceed 100,000 – while the current approaches can only handle thousands of features / samples. Our second contribution is a distributed algorithm for GCCA, which computes the canonical components of different views in parallel and thus can further reduce the runtime significantly (by ≥ 30% in experiments) if multiple cores are available. Judiciously designed synthetic and real-data experiments using a multilingual dataset are employed to showcase the effectiveness of the proposed algorithms. Xiao Fu 0001, Kejun Huang, Evangelos E. Papalexakis, Hyun Ah Song, Partha P. Talukdar, Nicholas D. Sidiropoulos, Christos Faloutsos, Tom M. Mitchell |
ICDM | 1 |
| 2016 | Anchor-Free Correlated Topic Modeling: Identifiability and AlgorithmabstractIn topic modeling, many algorithms that guarantee identifiability of the topics have been developed under the premise that there exist anchor words -- i.e., words that only appear (with positive probability) in one topic. Follow-up work has resorted to three or higher-order statistics of the data corpus to relax the anchor word assumption. Reliable estimates of higher-order statistics are hard to obtain, however, and the identification of topics under those models hinges on uncorrelatedness of the topics, which can be unrealistic. This paper revisits topic modeling based on second-order moments, and proposes an anchor-free topic mining framework. The proposed approach guarantees the identification of the topics under a much milder condition compared to the anchor-word assumption, thereby exhibiting much better robustness in practice. The associated algorithm only involves one eigen-decomposition and a few small linear programs. This makes it easy to implement and scale up to very large problem instances. Experiments using the TDT2 and Reuters-21578 corpus demonstrate that the proposed anchor-free approach exhibits very favorable performance (measured using coherence, similarity count, and clustering accuracy metrics) compared to the prior art. Kejun Huang, Xiao Fu 0001, Nicholas D. Sidiropoulos |
NIPS | 2 |
| 2016 | Robustness Analysis of Structured Matrix Factorization via Self-Dictionary Mixed-Norm OptimizationabstractWe are interested in a low-rank matrix factorization problem where one of the matrix factors has a special structure; specifically, its columns live in the unit simplex. This problem finds applications in diverse areas such as hyperspectral unmixing, video summarization, spectrum sensing, and blind speech separation. Prior works showed that such a factorization problem can be formulated as a self-dictionary sparse optimization problem under some assumptions that are considered realistic in many applications, and convex mixed norms were employed as optimization surrogates to realize the factorization in practice. Numerical results have shown that the mixed-norm approach demonstrates promising performance. In this letter, we conduct performance analysis of the mixed-norm approach under noise perturbations. Our result shows that using a convex mixed norm can indeed yield provably good solutions. More importantly, we also show that using nonconvex mixed (quasi) norms is more advantageous in terms of robustness against noise. Xiao Fu 0001, Wing-Kin Ma |
IEEE Signal Process. Lett. | 1 |
| 2016 | Semiblind Hyperspectral Unmixing in the Presence of Spectral Library MismatchesabstractThe dictionary-aided sparse regression (SR) approach has recently emerged as a promising alternative to hyperspectral unmixing (HU) in remote sensing. By using an available spectral library as a dictionary, the SR approach identifies the underlying materials in a given hyperspectral image by selecting a small subset of spectral samples in the dictionary to represent the whole image. A drawback with the current SR developments is that an actual spectral signature in the scene is often assumed to have zero mismatch with its corresponding dictionary sample, and such an assumption is considered too ideal in practice. In this paper, we tackle the spectral signature mismatch problem by proposing a dictionary-adjusted nonconvex sparsity-encouraging regression (DANSER) framework. The main idea is to incorporate dictionary correcting variables in an SR formulation. A simple and low per-iteration complexity algorithm is tailor-designed for practical realization of DANSER. Using the same dictionary correcting idea, we also propose a robust subspace solution for dictionary pruning. Extensive simulations and real-data experiments show that the proposed method is effective in mitigating the undesirable spectral signature mismatch effects. Xiao Fu 0001, Wing-Kin Ma, José M. Bioucas-Dias, Tsung-Han Chan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Translation Invariant Word EmbeddingsabstractKejun Huang, Matt Gardner, Evangelos Papalexakis, Christos Faloutsos, Nikos Sidiropoulos, Tom Mitchell, Partha P. Talukdar, Xiao Fu. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Kejun Huang, Matt Gardner 0001, Evangelos E. Papalexakis, Christos Faloutsos, Nicholas D. Sidiropoulos, Tom M. Mitchell, Partha P. Talukdar, Xiao Fu 0001 |
EMNLP | 8 |
| 2014 | Blind spectra separation and direction finding for cognitive radio using temporal correlation-domain ESPRITabstractUnraveling power spectra mixtures and finding the directions of the constituent sources can enable effective spatial occupancy prediction by location-dependent combining of the recovered source power spectra; and it is also useful for primary interference avoidance. Such unmixing and direction-finding is a challenging leap beyond ordinary `aggregate' spectrum sensing. This paper presents a promising new method for blind (power) spectra separation and emitter direction finding using a network of cognitive radios. Each radio has a pair of antennas, and the baselines of different radios are aligned (e.g., using a compass), in a configuration reminiscent of classical spatial correlation-based ESPRIT. Unlike classical ESPRIT, array geometry is exploited here in the temporal correlation domain to come up with a simple and effective blind spectra separation and direction finding solution with guaranteed identifiability and robustness to noise. A notable feature is that the different radios need not be synchronized, as they do in spatial ESPRIT. Xiao Fu 0001, Nicholas D. Sidiropoulos, Wing-Kin Ma, John H. Tranter |
ICASSP | 1 |
| 2013 | Blind separation of convolutive mixtures of speech sources: Exploiting local sparsityabstractThis paper presents an efficient method for blind source separation of convolutively mixed speech signals. The method follows the popular frequency-domain approach, wherein researchers are faced with two main problems, namely, per-frequency mixing system estimation, and permutation alignment of source components at all frequencies. We adopt a novel concept, where we utilize local sparsity of speech sources in transformed domain, together with non-stationarity, to address the two problems. Such exploitation leads to a closed-form solution for per-frequency mixing system estimation and a numerically simple method for permutation alignment, both of which are efficient to implement. Simulations show that the proposed method yields comparable source recovery performance to that of a state-of-the-art method, while requires much less computation time. Xiao Fu 0001, Wing-Kin Ma |
ICASSP | 1 |
| 2013 | A Khatri-Rao subspace approach to blind identification of mixtures of quasi-stationary sources
Ka-Kit Lee, Wing-Kin Ma, Xiao Fu 0001, Tsung-Han Chan, Chong-Yung Chi |
Signal Process. | 3 |
| 2012 | A simple closed-form solution for overdetermined blind separation of locally sparse quasi-stationary sourcesabstractWe consider the scenario of an unknown overdetermined instantaneous mixture of quasi-stationary sources. Blind source separation (BSS) under this scenario has drawn much attention, motivated by applications such as speech and audio separation. The ideas in the existing BSS works often focus on exploiting the time-varying statistics characteristics of quasi-stationary sources, through various kinds of formulations and optimization methods. In this paper, we are interested in further assuming that the sources exhibit some form of local sparsity, which is generally satisfied in speech. By exploiting this additional assumption, we show that there is a simple closed-form solution for the BSS problem. Simulation results based on real speech show that the proposed closed-form algorithm is computationally much lower than some existing BSS algorithms, while delivering a promising mean-square-error performance. Xiao Fu 0001, Wing-Kin Ma |
ICASSP | 1 |