EDBT 2026 Demo / reviewers in the wild / expert
Andong Wang
dblp:190/5540
· DBLP profile ↗
26ranked-venue papers
13as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-view subspace tensorization with attentive clustering embedding
Yanghang Zheng, Haonan Huang, Yihao Luo, Yuning Qiu, Andong Wang, Guoxu Zhou, Qibin Zhao |
Neural Networks | 5 |
| 2026 | Robust Tensor Decomposition Under Multi-Mode Outlier Corruptions
Yuning Qiu, Guoxu Zhou, Andong Wang, Qibin Zhao, Shengli Xie 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Follow-Your-Click: Open-domain Regional Image Animation via Motion PromptsabstractDespite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may need to control the movement of different objects or regions. Additionally, current I2V methods require users not only to describe the target motion but also to provide redundant detailed descriptions of frame contents.These two issues hinder the practical utilization of current I2V tools. In this paper, we propose a practical framework, named Follow-Your-Click, to achieve image animation with a simple user click (for specifying what to move) and a motion prompt (for specifying how to move). Technically, we propose the first-frame masking strategy, which significantly improves the video generation quality, and a motion-augmented module equipped with a motion prompt dataset to improve the motion prompt following abilities of our model. To further control the motion speed, we propose flow-based motion magnitude control to control the speed of target movement more precisely. Extensive experiments compared with 7 baselines, including both commercial tools and research methods on 8 metrics, suggest the superiority of our approach. Yue Ma 0016, Yingqing He, Hongfa Wang, Andong Wang, Leqi Shen, Jixuan Ying, Chengfei Cai, Zhifeng Li 0001, Harry Shum, Wei Liu 0005, Qifeng Chen 0001 |
AAAI | 4 |
| 2025 | STEPS: Sequential Probability Tensor Estimation for Text-to-Image Hard Prompt SearchabstractRecent text-to-image (T2I) diffusion models have demonstrated remarkable capabilities in visual synthesis, yet their performance heavily relies on the quality of input prompts. However, optimizing discrete prompts remains challenging because the discrete nature of tokens prevents the direct application of the gradient descent method and the vast search space of possible token combinations. As a result, existing approaches either suffer from quantization errors when employing continuous optimization techniques or be- come trapped in local optima due to coordinate-wise greedy search. In this paper, we propose STEPS, a novel Sequential probability Tensor Estimation approach for hard Prompt Search. Our method reformulates discrete prompt optimization as a sequential probability tensor estimation problem, leveraging the inherent low-rank characteristics to address the curse of dimensionality. To further improve the computational efficiency, we develop a memory-bounded sampling approach that shrinks the prompt space without the iteration step dependency while preserving sequential optimization dynamics. Extensive experiments on various public datasets demonstrate that our method consistently outperforms existing approaches in T2I generation, cross-model prompt transferability, and harmful prompt optimization, validating the effectiveness of the proposed framework. Yuning Qiu, Andong Wang, Chao Li 0013, Haonan Huang, Guoxu Zhou, Qibin Zhao |
CVPR | 2 |
| 2025 | Towards Multiple Character Image Animation Through Enhancing Implicit DecouplingabstractControllable character image animation has a wide range of applications. Although existing studies have consistently improved performance, challenges persist in the field of character image animation, particularly concerning stability in complex backgrounds and tasks involving multiple characters. To address these challenges, we propose a novel multi-condition guided framework for character image animation, employing several well-designed input modules to enhance the implicit decoupling capability of the model. First, the optical flow guider calculates the background optical flow map as guidance information, which enables the model to implicitly learn to decouple the background motion into background constants and background momentum during training, and generate a stable background by setting zero background momentum during inference. Second, the depth order guider calculates the order map of the characters, which transforms the depth information into the positional information of multiple characters. This facilitates the implicit learning of decoupling different characters, especially in accurately separating the occluded body parts of multiple characters. Third, the reference pose map is input to enhance the ability to decouple character texture and pose information in the reference image. Furthermore, to fill the gap of fair evaluation of multi-character image animation, we propose a new benchmark comprising about 4,000 frames. Extensive qualitative and quantitative evaluations demonstrate that our method excels in generating high-quality character animations, especially in scenarios of complex backgrounds and multiple characters. Jingyun Xue, Hongfa Wang, Qi Tian 0003, Yue Ma 0016, Andong Wang, Zhiyuan Zhao 0002, Shaobo Min, Kaihao Zhang, Harry Shum, Wei Liu 0005, Mengyang Liu, Wenhan Luo |
ICLR | 5 |
| 2025 | Low-Rank Tensor Transitions (LoRT) for Transferable Tensor RegressionabstractTensor regression is a powerful tool for analyzing complex multi-dimensional data in fields such as neuroimaging and spatiotemporal analysis, but its effectiveness is often hindered by insufficient sample sizes. To overcome this limitation, we adopt a transfer learning strategy that leverages knowledge from related source tasks to improve performance in data-scarce target tasks. This approach, however, introduces additional challenges including model shifts, covariate shifts, and decentralized data management. We propose the Low-Rank Tensor Transitions (LoRT) framework, which incorporates a novel fusion regularizer and a two-step refinement to enable robust adaptation while preserving low-tubal-rank structure. To support decentralized scenarios, we extend LoRT to D-LoRT, a distributed variant that maintains statistical efficiency with minimal communication overhead. Theoretical analysis and experiments on tensor regression tasks, including compressed sensing and completion, validate the robustness and versatility of the proposed methods. These findings indicate the potential of LoRT as a robust method for tensor regression in settings with limited data and complex distributional structures. Andong Wang, Yuning Qiu, Zhong Jin, Guoxu Zhou, Qibin Zhao |
ICML | 1 |
| 2025 | Towards a Geometric Understanding of Tensor Learning via the t-ProductabstractDespite the growing success of transform-based tensor models such as the t-product, their underlying geometric principles remain poorly understood. Classical differential geometry, built on real-valued function spaces, is not well suited to capture the algebraic and spectral structure induced by transform-based tensor operations. In this work, we take an initial step toward a geometric framework for tensors equipped with tube-wise multiplication via orthogonal transforms. We introduce the notion of smooth t-manifolds, defined as topological spaces locally modeled on structured tensor modules over a commutative t-scalar ring. This formulation enables transform-consistent definitions of geometric objects, including metrics, gradients, Laplacians, and geodesics, thereby bridging discrete and continuous tensor settings within a unified algebraic-geometric perspective.
On this basis, we develop a statistical procedure for testing whether tensor data lie near a low-dimensional t-manifold, and provide nonasymptotic guarantees for manifold fitting under noise. We further establish approximation bounds for tensor neural networks that learn smooth functions over t-manifolds, with generalization rates determined by intrinsic geometric complexity. This framework offers a theoretical foundation for geometry-aware learning in structured tensor spaces and supports the development of models that align with transform-based tensor representations. Andong Wang, Yuning Qiu, Haonan Huang, Zhong Jin, Guoxu Zhou, Qibin Zhao |
NeurIPS | 1 |
| 2025 | Balanced Unfolding Induced Tensor Nuclear Norms for High-Order Tensor CompletionabstractThe recently proposed tensor tubal rank has been witnessed to obtain extraordinary success in real-world tensor data completion. However, existing works usually fix the transform orientation along the third mode and may fail to turn multidimensional low-tubal-rank structure into account. To alleviate these bottlenecks, we introduce two unfolding induced tensor nuclear norms (TNNs) for the tensor completion (TC) problem, which naturally extends tensor tubal rank to high-order data. Specifically, we show how multidimensional low-tubal-rank structure can be captured by utilizing a novel balanced unfolding strategy, upon which two TNNs, namely, overlapped TNN (OTNN) and latent TNN (LTNN), are developed. We also show the immediate relationship between the tubal rank of unfolding tensor and the existing tensor network (TN) rank, e.g., CANDECOMP/PARAFAC (CP) rank, Tucker rank, and tensor ring (TR) rank, to demonstrate its efficiency and practicality. Two efficient TC models are then proposed with theoretical guarantees by analyzing a unified nonasymptotic upper bound. To solve optimization problems, we develop two alternating direction methods of multipliers (ADMM) based algorithms. The proposed models have been demonstrated to exhibit superior performance based on experimental findings involving synthetic and real-world tensors, including facial images, light field images, and video sequences. Yuning Qiu, Guoxu Zhou, Andong Wang, Qibin Zhao, Shengli Xie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Towards Multi-Mode Outlier Robust Tensor Ring DecompositionabstractConventional Outlier Robust Tensor Decomposition (ORTD) approaches generally represent sparse outlier corruption within a specific mode. However, such an assumption, which may hold for matrices, proves inadequate when applied to high-order tensors. In the tensor domain, the outliers are prone to be corrupted in multiple modes simultaneously. Addressing this limitation, this study proposes a novel ORTD approach by recovering low-rank tensors contaminated by outliers spanning multiple modes. In particular, we conceptualize outliers within high-order tensors as latent tensor group sparsity by decomposing the corrupted tensor into a sum of multiple latent components, where each latent component is exclusive to outliers within a particular direction. Thus, it can effectively mitigate the outlier corruptions prevalent in high-order tensors across multiple modes. To theoretically guarantee recovery performance, we rigorously analyze a non-asymptotic upper bound of the estimation error for the proposed ORTD approach. In the optimization process, we develop an efficient alternate direction method of multipliers (ADMM) algorithm. Empirical validation of the approach's efficacy is undertaken through comprehensive experimentation. Yuning Qiu, Guoxu Zhou, Andong Wang, Qibin Zhao |
AAAI | 3 |
| 2024 | SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World KnowledgeabstractLearning commonsense reasoning from visual contexts and scenes in real-world is a crucial step toward advanced artificial intelligence. However, existing video reasoning benchmarks are still inadequate since they were mainly designed for factual or situated reasoning and rarely involve broader knowledge in the real world. Our work aims to delve deeper into reasoning evaluations, specifically within dynamic, open-world, and structured context knowledge. We propose a new benchmark (SOK-Bench), consisting of 44K questions and 10K situations with instance-level annotations depicted in the videos. The reasoning process is required to understand and apply situated knowledge and general knowledge for problem-solving. To create such a dataset, we propose an automatic and scalable gener-ation method to generate question-answer pairs, knowledge graphs, and rationales by instructing the combinations of LLMs and MLLMs. Concretely, we first extract observable situated entities, relations, and processes from videos for situated knowledge and then extend to open-world knowledge beyond the visible content. The task generation is facilitated through multiple dialogues as iterations and subsequently corrected and refined by our designed self-promptings and demonstrations. With a corpus of both explicit situated facts and implicit commonsense, we generate associated question-answer pairs and reasoning processes, finally followed by manual reviews for quality assurance. We evaluated recent mainstream large vision-language models on the benchmark and found several in-sightful conclusions. For more information, please refer to our benchmark at www.bobbywu.com/SOKBench. Andong Wang, Bo Wu 0018, Sunli Chen, Zhenfang Chen, Haotian Guan, Wei-Ning Lee, Li Erran Li, Chuang Gan 0001 |
CVPR | 1 |
| 2024 | Diffusion Models Demand Contrastive Guidance for Adversarial Purification to AdvanceabstractIn adversarial defense, adversarial purification can be viewed as a special generation task with the purpose to remove adversarial attacks and diffusion models excel in adversarial purification for their strong generative power. With different predetermined generation requirements, various types of guidance have been proposed, but few of them focuses on adversarial purification. In this work, we propose to guide diffusion models for adversarial purification using contrastive guidance. We theoretically derive the proper noise level added in the forward process diffusion models for adversarial purification from a feature learning perspective. For the reverse process, it is implied that the role of contrastive loss guidance is to facilitate the evolution towards the signal direction. From the theoretical findings and implications, we design the forward process with the proper amount of Gaussian noise added and the reverse process with the gradient of contrastive loss as the guidance of diffusion models for adversarial purification. Empirically, extensive experiments on CIFAR-10, CIFAR-100, the German Traffic Sign Recognition Benchmark and ImageNet datasets with ResNet and WideResNet classifiers show that our method outperforms most of current adversarial training and adversarial purification methods by a large improvement. Mingyuan Bai, Tenghui Li 0001, Andong Wang, Junbin Gao, Cesar F. Caiafa, Qibin Zhao |
ICML | 4 |
| 2024 | Adversarially Robust Deep Multi-View Clustering: A Novel Attack and Defense FrameworkabstractDeep Multi-view Clustering (DMVC) stands out as a widely adopted technique aiming at enhanced clustering performance by leveraging diverse data sources. However, the critical issue of vulnerability to adversarial attacks is unexplored due to the lack of well-defined attack objectives. To fill this crucial gap, this paper is the first work to investigate the possibility of adversarial attacks on DMVC models. Specifically, we introduce an adversarial attack with Generative Adversarial Networks (GANs) with the aim to maximally change the complementarity and consistency of multiple views, thus leading to wrong clustering. Building upon this adversarial context, in the realm of defense, we propose a novel Adversarially Robust Deep Multi-View Clustering by leveraging adversarial training. Based on the analysis from an information-theoretic perspective, we design an Attack Mitigator that provides a foundation to guarantee the adversarial robustness of our DMVC models. Experiments conducted on multi-view datasets confirmed that our attack framework effectively reduces the clustering performance of the target model. Furthermore, our proposed adversarially robust method is also demonstrated to be an effective defense against such attacks. This work is a pioneer in exploring adversarial threats and advancing both theoretical understanding and practical strategies for robust multi-view clustering. Code is available at https://github.com/libertyhhn/AR-DMVC. Haonan Huang, Guoxu Zhou, Yanghang Zheng, Yuning Qiu, Andong Wang, Qibin Zhao |
ICML | 5 |
| 2024 | Generalized Tensor Decomposition for Understanding Multi-Output Regression under Combinatorial ShiftsabstractIn multi-output regression, we identify a previously neglected challenge that arises from the inability of training distribution to cover all combinations of input features, leading to combinatorial distribution shift (CDS). To the best of our knowledge, this is the first work to formally define and address this problem. We tackle it through a novel tensor decomposition perspective, proposing the Functional t-Singular Value Decomposition (Ft-SVD) theorem which extends the classical tensor SVD to infinite and continuous feature domains, providing a natural tool for representing and analyzing multi-output functions. Within the Ft-SVD framework, we formulate the multi-output regression problem under CDS as a low-rank tensor estimation problem under the missing not at random (MNAR) setting, and introduce a series of assumptions about the true functions, training and testing distributions, and spectral properties of the ground-truth embeddings, making the problem more tractable.
To address the challenges posed by CDS in multi-output regression, we develop a tailored Double-Stage Empirical Risk Minimization (ERM-DS) algorithm that leverages the spectral properties of the embeddings and uses specific hypothesis classes in each frequency component to better capture the varying spectral decay patterns. We provide rigorous theoretical analyses that establish performance guarantees for the ERM-DS algorithm. This work lays a preliminary theoretical foundation for multi-output regression under CDS. Andong Wang, Yuning Qiu, Mingyuan Bai, Zhong Jin, Guoxu Zhou, Qibin Zhao |
NeurIPS | 1 |
| 2023 | Transformed Low-Rank Parameterization Can Help Robust Generalization for Tensor Neural NetworksabstractMulti-channel learning has gained significant attention in recent applications, where neural networks with t-product layers (t-NNs) have shown promising performance through novel feature mapping in the transformed domain.
However, despite the practical success of t-NNs, the theoretical analysis of their generalization remains unexplored. We address this gap by deriving upper bounds on the generalization error of t-NNs in both standard and adversarial settings. Notably, it reveals that t-NNs compressed with exact transformed low-rank parameterization can achieve tighter adversarial generalization bounds compared to non-compressed models. While exact transformed low-rank weights are rare in practice, the analysis demonstrates that through adversarial training with gradient flow, highly over-parameterized t-NNs with the ReLU activation can be implicitly regularized towards a transformed low-rank parameterization under certain conditions. Moreover, this paper establishes sharp adversarial generalization bounds for t-NNs with approximately transformed low-rank weights. Our analysis highlights the potential of transformed low-rank parameterization in enhancing the robust generalization of t-NNs, offering valuable insights for further research and development. Andong Wang, Chao Li 0013, Mingyuan Bai, Zhong Jin, Guoxu Zhou, Qibin Zhao |
NeurIPS | 1 |
| 2022 | HINT: Hierarchical Neuron Concept ExplainerabstractTo interpret deep networks, one main approach is to associate neurons with human-understandable concepts. However, existing methods often ignore the inherent connections of different concepts (e.g., dog and cat both belong to animals), and thus lose the chance to explain neurons responsible for higher-level concepts (e.g., animal). In this paper, we study hierarchical concepts inspired by the hierarchical cognition process of human beings. To this end, we propose HIerarchical Neuron concepT explainer (HINT) to effectively build bidirectional associations between neurons and hierarchical concepts in a low-cost and scalable manner. HINT enables us to systematically and quantitatively study whether and how the implicit hierarchical relationships of concepts are embedded into neurons. Specifically, HINT identifies collaborative neurons responsible for one concept and multimodal neurons pertinent to different concepts, at different semantic levels from concrete concepts (e.g., dog) to more abstract ones (e.g., animal). Finally, we verify the faithfulness of the associations using Weakly Supervised Object Localization, and demonstrate its applicability in various tasks, such as discovering saliency regions and explaining adversarial attacks. Code is available on https://github.com/AntonotnaWang/HINT. Andong Wang, Wei-Ning Lee, Xiaojuan Qi 0001 |
CVPR | 1 |
| 2022 | Semantic-Sparse Colorization Network for Deep Exemplar-Based Colorization
Yunpeng Bai, Chao Dong 0005, Zenghao Chai, Andong Wang, Zhengzhuo Xu, Chun Yuan 0003 |
ECCV (6) | 4 |
| 2020 | Robust Tensor Decomposition via Orientation Invariant Tubal Nuclear NormsabstractLow-rank tensor recovery has been widely applied to computer vision and machine learning. Recently, tubal nuclear norm (TNN) based optimization is proposed with superior performance as compared to other tensor nuclear norms. However, one major limitation is its orientation sensitivity due to low-rankness strictly defined along tubal orientation and it cannot simultaneously model spectral low-rankness in multiple orientations. To this end, we introduce two new tensor norms called OITNN-O and OITNN-L to exploit multi-orientational spectral low-rankness for an arbitrary K-way (K ≥ 3) tensors. We further formulate two robust tensor decomposition models via the proposed norms and develop two algorithms as the solutions. Theoretically, we establish non-asymptotic error bounds which can predict the scaling behavior of the estimation error. Experiments on real-world datasets demonstrate the superiority and effectiveness of the proposed norms. Andong Wang, Chao Li 0013, Zhong Jin, Qibin Zhao |
AAAI | 1 |
| 2020 | Classification of Epileptic IEEG Signals by CNN and Data AugmentationabstractEpileptic focus localization in patients with epileptic seizures is essential when surgery is needed. Recent studies show that this can be done automatically using machine learning approaches. However, well-designed feature extraction methods are often computationally demanding, requiring a large amount of data labeled by physicians, which is time consuming and impractical. In this paper, we firstly introduce a one-dimensional convolutional neural network (1D-CNN) model for epileptic seizure focus detection which avoids the manual, time-consuming feature extraction Moreover, to reduce the necessary number of training samples, we introduce an approach for data augmentation. The experimental results demonstrate the efficiency of the proposed method, with a nearly 3% improvement in performance using the data enhancement method compared to the best result obtained using the traditional feature extraction method. Jordi Solé i Casals, Binghua Li 0001, Zihao Huang 0003, Andong Wang, Jianting Cao, Toshihisa Tanaka 0001, Qibin Zhao |
ICASSP | 5 |
| 2020 | Robust tensor decomposition via t-SVD: Near-optimal statistical guarantee and scalable algorithms
Andong Wang, Zhong Jin, Guoqing Tang |
Signal Process. | 1 |
| 2019 | Latent Schatten TT Norm for Tensor CompletionabstractTensor completion arouses much attention in signal processing and machine learning. The tensor train (TT) decomposition has shown better performances than the Tucker decomposition in image and video inpainting. In this paper, we propose a novel tensor completion model based on a newly defined latent Schatten TT norm. Then, the statistical performance is analyzed by establishing a non-asymptotic upper bound on the estimation error. Further, a scalable algorithm is developed to efficiently solve the model. Experimental results of color image inpainting demonstrate that the proposed norm has promising performances compared to other variants of Schatten norm. Andong Wang, Xulin Song, Xiyin Wu, Zhihui Lai 0001, Zhong Jin |
ICASSP | 1 |
| 2019 | Generalized Dantzig Selector for Low-tubal-rank Tensor RecoveryabstractDue to the superiority in exploiting the ubiquitous "spatial-shifting" property in modern multi-way data, the recently proposed low-tubal-rank model has been successfully applied for tensor recovery in signal processing and computer vision. In this paper, we define the generalized tensor Dantzig selector to recover a low-tubal-rank tensor from noisy linear measurements. Algorithmically, we develop an efficient algorithm based on the ADMM framework. Statistically, we establish non-asymptotic upper bounds on the estimation error for the problems of tensor completion and compressive sensing. Numerical experiments illustrate that our bounds can predict the scaling behavior of the estimation error. Experiments on realword datasets show the effectiveness of the proposed model. Andong Wang, Xulin Song, Xiyin Wu, Zhihui Lai 0001, Zhong Jin |
ICASSP | 1 |
| 2019 | Robust Low-tubal-rank Tensor CompletionabstractReal multi-way data may suffer from missing entries, noise and outliers simultaneously. The recently proposed tubal nuclear norm (TNN) has shown its superiority in tensor completion. However, statistical analysis of TNN based models is still deficient. This paper aims to robustly recover a polluted incomplete tensor with rigorous statistical guarantee. Specifically, an estimator based on a weighed variant of TNN is proposed to complete a low-tubal-rank tensor corrupted by element sparse errors or slice sparse sample outliers from partial noisy observations. Non-asymptotic upper bounds on the estimation error are established and further proved to be minimax optimal up to a log factor. Sharpness of the upper bounds is verified on synthetic datasets and superiority of the proposed estimator is demonstrated through robust video inpainting. Andong Wang, Xulin Song, Xiyin Wu, Zhihui Lai 0001, Zhong Jin |
ICASSP | 1 |
| 2019 | Noisy low-tubal-rank tensor completion
Andong Wang, Zhihui Lai 0001, Zhong Jin |
Neurocomputing | 1 |
| 2019 | Statistical performance of convex low-rank and sparse tensor recovery
Xiangrui Li, Andong Wang, Jianfeng Lu 0003, Zhenmin Tang |
Pattern Recognit. | 2 |
| 2018 | Optimization of Urban Heating Network Design Using Genetic AlgorithmabstractAs the main energy source is coal burning, district heating in Northern China is an important driver of air pollution. Optimization of the performance of District Heating Network (DHN) carries both social and economic benefits. This study proposes an approach for optimizing urban heating network design based on Genetic Algorithm. Our case study shows that DHN can meet the users' requirements and achieve minimum cost in parallel. Andong Wang, Victor O. K. Li, Jacqueline C. K. Lam |
IEEE BigData | 1 |
| 2018 | Salient Object Detection Via Deformed Smoothness ConstraintabstractIn recent years, various graph-based salient object detection methods have been successfully proposed. Since existing methods may miss some object regions with low contrast to background, a novel propagation model via deformed smoothness constraint is proposed to address this problem. By regularizing nodes and their neighbors locally, the deformed smoothness constraint is able to prevent erroneous label propagation. Thus, the object regions with low contrast to background can be emerged. Besides, the deformed smoothness constraint is further utilized in a map refinement model, which can suppress the background noises in label propagation result. Experiments on three public datasets show that the proposed method outperforms eleven state-of-the-art salient object detection methods. Xiyin Wu, Xiaodi Ma, Jinxia Zhang, Andong Wang, Zhong Jin |
ICIP | 4 |