EDBT 2026 Demo / reviewers in the wild / expert
Liu Liu 0014
dblp:74/7037-14
· DBLP profile ↗
60ranked-venue papers
5as first author
54since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 5 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 26 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Remodeling Semantic Relationships in Vision-Language Fine-TuningabstractVision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus leading to suboptimal performance. Toward solving this problem, we propose a method that can improve multimodal alignment and fusion based on both semantics and relationships.Specifically, we first extract multilevel semantic features from different vision encoder to capture more visual cues of the relationships. Then, we learn to project the vision features to group related semantics, among which are more likely to have relationships. Finally, we fuse the visual features with the textual by using inheritable cross-attention, where we globally remove the redundant visual relationships by discarding visual-language feature pairs with low correlation. We evaluate our proposed method on eight foundation models and two downstream tasks, visual question answering and image captioning, and show that it outperforms all existing methods. Liu Liu 0014, Baosheng Yu, Jiayan Qiu |
AAAI | 2 |
| 2026 | CrystalDiT: Simple Diffusion Transformers for Crystal GenerationabstractWe present CrystalDiT, a diffusion transformer for crystal structure generation that achieves state-of-the-art performance by challenging the trend of architectural complexity. Instead of intricate, multi-stream designs, CrystalDiT employs a unified transformer that imposes a powerful inductive bias: treating lattice and atomic properties as a single, interdependent system. Combined with a periodic table-based atomic representation and a balanced training strategy, our approach achieves 8.78% SUN (Stable, Unique, Novel) rate on MP-20, substantially outperforming recent methods including FlowMM (4.21%) and MatterGen (3.66%). Notably, CrystalDiT generates 63.28% unique and novel structures while maintaining comparable stability rates, demonstrating that architectural simplicity can be more effective than complexity for materials discovery. Our results suggest that in data-limited scientific domains, carefully designed simple architectures outperform sophisticated alternatives that are prone to overfitting. Xiaohan Yi, Guikun Xu, Zhong Zhang 0014, Liu Liu 0014, Yatao Bian, Xi Xiao 0001, Peilin Zhao |
AAAI | 4 |
| 2026 | Towards alleviating hallucination in text-to-image retrieval for CLIP in zero-shot learning
Hanyao Wang, Yibing Zhan, Liu Liu 0014, Liang Ding 0006, Jun Yu 0002 |
Neurocomputing | 3 |
| 2026 | Structure-Aware Alignment for Day-Night Cross-Domain Vehicle Re-IdentificationabstractDay-night cross-domain vehicle re-identification (DN-ReID) is fundamentally challenged by drastic illumination changes that create substantial domain gaps and hinder consistent feature representation. Most existing methods focus on aligning distribution statistics but often overlook essential structural relationships, such as cross-domain clustering and geometric topology. To address this, we propose a structure-aware alignment (SAA) method that, for the first time, formulates centered kernel alignment as a trainable loss for structural alignment between domains. This approach explicitly aligns cross-domain relational structures, thereby supporting the transfer of intra-class compactness and inter-class separability from the source to the target domain for robust cross-domain matching. Extensive experiments on DN-348 and DN-Wild demonstrate that our approach consistently outperforms state-of-the-art methods, achieving a 2.7% mAP improvement on DN-348. Jingyi Zhuang, Baihui Sa, Jinjie Zheng, Liu Liu 0014, Jianqing Zhu, Huanqiang Zeng |
IEEE Signal Process. Lett. | 4 |
| 2026 | Diffusion-Based Text-Guided Image Generation With Fine-Grained Spatial Object-Attribute RelationshipsabstractExpressing and controlling fine-grained spatial attributes of objects in large-scale models presents significant challenges, as these spatial attributes are often difficult to describe textually and exhaustive enumeration is impractical. This hinders effective alignment with user preferences regarding spatial attribute-object relationships in fine-grained synthesis tasks. To tackle this problem, we propose AttrObjDiff, a novel framework built on the pre-trained Stable Diffusion model to integrate spatial attribute maps. Firstly, AttrObjDiff constrains the denoising step using trainable cross-attention fusion modules, attribute-enhancing cross-attention and LoRAs. The fusion modules take layout features extracted by a frozen ControlNet and corresponding fine-grained attribute maps as inputs to generate joint constraint features of spatial attribute-object relationships. We leverage attribute-enhancing cross-attention within the U-Net to further refine these spatial attributes. Finally, LoRAs are employed to align with these joint constraint features of finegrained relationships. Secondly, AttrObjDiff enhances the reverse process with lightweight noise reranking models to improve spatial object-attribute alignment. The reranking models select semantic noises related to fine-grained relationships, improving synthesis quality without significantly increasing computational costs. Experimental results demonstrate that our method can generate high-quality images guided by fine-grained spatial object-attribute relationships, improving synthesis controllability and semantic consistency. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Ziliang Ren, Dacheng Tao, Xinyu Wu 0001, Jun Cheng 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | NoisePO: Efficient Semantic Noise Generation and Ranking for Diffusion-Based Text-to-Image SynthesisabstractDiffusion-based methods have achieved remarkable success in photorealistic image generation, leveraging iterative denoising steps to improve image quality. However, multi-step denoising often suffers from error accumulation-similar to exposure bias in autoregressive models-due to suboptimal noise estimation, which can lead to degraded semantic alignment and image fidelity. To tackle the challenge of suboptimal inner latent representations in generation and improve the inner latent, this paper introduces a novel method NoisePO, an efficient semantic noise preference optimization framework. NoisePO employs a semantic noise preference optimization generative adversarial network (NPO-GAN) and noise ranking methods to search for semantically relevant noises based on textual conditions, thus eliminating undesired semantic features while emphasizing the necessary semantic ones. Specifically, NoisePO utilizes a light NPO-GAN to generate semantic noises that encourage the latent at the previous step to incorporate more semantic information from the caption. Then, light ranking models are employed to filter out low-quality noises and select the best noise. Experimental results demonstrate that NoisePO consistently outperforms the baselines across widely used frameworks, achieving notable improvements in image quality, semantic consistency, and user-specific alignment as measured by IS, FID, CLIP, and other metrics. These results indicate that NoisePO effectively enhances synthesis quality and strengthens text-image alignment. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Chengqun Song, Dacheng Tao, Jun Cheng 0002 |
IEEE Trans. Image Process. | 2 |
| 2025 | HDT: Hierarchical Discrete Transformer for Multivariate Time Series ForecastingabstractGenerative models have gained significant attention in multivariate time series forecasting (MTS), particularly due to their ability to generate high-fidelity samples. Forecasting the probability distribution of multivariate time series is a challenging yet practical task. Although some recent attempts have been made to handle this task, two major challenges persist: 1) some existing generative methods underperform in high-dimensional multivariate time series forecasting, which is hard to scale to higher dimensions; 2) The inherent high-dimensional multivariate attributes constrain the forecasting lengths of existing generative models. In this paper, we point out that discrete token representations can model high-dimensional MTS with faster inference time, and forecast the target with the long-term trends of itself can extend the forecasting length with high accuracy. Motivated by this, we propose a vector quantized framework called Hierarchical Discrete Transformer (HDT) that models time series into discrete token representations with l2 normalization enhanced vector quantized strategy, in which we transform the MTS forecasting into discrete tokens generation. To address the limitations of generative models in long-term forecasting, we propose a hierarchical discrete Transformer. This model captures the discrete long-term trend of the target at the low level and leverages this trend as a condition to generate the discrete representation of the target at the high level that introduces the features of target itself for extending the forecasting length in high-dimensional MTS. Extensive experiments on five popular MTS datasets verify the effectiveness of our proposed method. The source code will be released. Shibo Feng, Peilin Zhao, Liu Liu 0014, Zhiqi Shen 0001 |
AAAI | 3 |
| 2025 | SkipNode: On Alleviating Performance Degradation for Deep Graph Convolutional Networks (Extended Abstract)abstractGraph Convolutional Networks (GCNs) are powerful tools for learning representations in graph-structured data. However, their performance tends to degrade with increased model depth due to over-smoothing. Although previous studies attribute degradation to over-smoothing, this work identifies the mutually reinforcing effects of over-smoothing and gradient vanishing as the root cause. In this paper, we propose SkipNode, a plug-and-play module that mitigates degradation in deep GCNs. SkipNode introduces node-sampling in each convolutional layer to selectively skip convolutions, preventing over-smoothing by reducing the depth experienced by specific nodes and facilitating gradient backpropagation. We demonstrate both theoretically and experimentally that SkipNode effectively curtails over-smoothing and gradient vanishing, improving deep GCN performance across diverse tasks. Extensive evaluations show SkipNode's robustness and superior performance over state-of-the-art (SOTA) baselines, establishing it as a practical solution for training deep GCNs. Weigang Lu 0001, Yibing Zhan, Binbin Lin 0001, Ziyu Guan, Liu Liu 0014, Baosheng Yu, Wei Zhao 0019, Yaming Yang 0002, Dacheng Tao |
ICDE | 5 |
| 2025 | Principled Data Selection for Alignment: The Hidden Risks of Difficult ExamplesabstractThe alignment of large language models (LLMs) often assumes that using more clean data yields better outcomes, overlooking the match between model capacity and example difficulty. Challenging this, we propose a new principle: *Preference data vary in difficulty, and overly difficult examples hinder alignment, by exceeding the model's capacity*. Through systematic experimentation, we validate this principle with three key findings: (1) preference examples vary in difficulty, as evidenced by consistent learning orders across alignment runs; (2) overly difficult examples significantly degrade performance across four LLMs and two datasets; and (3) the capacity of a model dictates its threshold for handling difficult examples, underscoring a critical relationship between data selection and model capacity. Building on this principle, we introduce *Selective DPO*, which filters out overly difficult examples. This simple adjustment improves alignment performance by 9-16\% in win rates on the AlpacaEval 2 benchmark compared to the DPO baseline, surpassing a series of DPO variants with different algorithmic adjustments. These results together illuminate the importance of aligning data difficulty with model capacity, offering a transformative perspective for improving alignment strategies in LLMs. Code is available at https://github.com/glorgao/SelectiveDPO Chengqian Gao, Liu Liu 0014, Zeke Xie, Peilin Zhao, Zhiqiang Xu 0003 |
ICML | 3 |
| 2025 | Test-time Adapted Reinforcement Learning with Action Entropy RegularizationabstractOffline reinforcement learning is widely applied in multiple fields due to its advantages in efficiency and risk control. However, a major problem it faces is the distribution shift between offline datasets and online environments. This mismatch leads to out-of-distribution (OOD) state-action pairs that fall outside the scope of the training data. Therefore, existing conservative training policies may not provide reliable decisions when the test environment deviates greatly from the offline dataset. In this paper, we propose Test-time Adapted Reinforcement Learning (TARL) to address this problem. TARL constructs unsupervised test-time optimization objectives for discrete and continuous control tasks, using test data without depending on environmental rewards. In discrete control tasks, it minimizes the entropy of predicted action probabilities to decrease uncertainty and avoid OOD state-action pairs. For continuous control tasks, it represents and minimizes action uncertainty based on the normal distribution of policy network outputs. Moreover, to prevent model bias caused by overfitting and error accumulation during the test-time update process, TARL enforces a KL divergence constraint between the fine-tuned policy and the original policy. For efficiency, TARL only updates the layer normalization layer parameters during testing. Extensive experiments on popular Atari game benchmarks and the D4RL dataset demonstrate the superiority of our method. Our method achieved a significant improvement over CQL, with a 13.6% episode return relative increase on the hopper-expert-v2 task. Shoukai Xu, Zihao Lian, Mingkui Tan, Liu Liu 0014, Zhong Zhang 0014, Peilin Zhao |
ICML | 4 |
| 2025 | Injecting Imbalance Sensitivity for Multi-Task LearningabstractMulti-task learning (MTL) has emerged as a promising approach for deploying deep learning models in real-life applications. Recent studies have proposed optimization-based learning paradigms to establish task-shared representations in MTL. However, our paper empirically argues that these studies, specifically gradient-based ones, primarily emphasize the conflict issue while neglecting the potentially more significant impact of imbalance/dominance in MTL. In line with this perspective, we enhance the existing baseline method by injecting imbalance-sensitivity through the imposition of constraints on the projected norms. To demonstrate the effectiveness of our proposed IMbalance-sensitive Gradient (IMGrad) descent method, we evaluate it on multiple mainstream MTL benchmarks, encompassing supervised learning tasks as well as reinforcement learning. The experimental results consistently demonstrate competitive performance. Liu Liu 0014, Peilin Zhao, Wei Gong 0001 |
IJCAI | 2 |
| 2025 | Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDPabstractExisting neural methods for multi-task vehicle routing problems (VRPs) typically learn unified solvers to handle multiple constraints simultaneously. However, they often underutilize the compositional structure of VRP variants, each derivable from a common set of basis VRP variants. This critical oversight causes unified solvers to miss out the potential benefits of basis solvers, each specialized for a basis VRP variant. To overcome this limitation, we propose a framework that enables unified solvers to perceive the shared-component nature across VRP variants by proactively reusing basis solvers, while mitigating the exponential growth of trained neural solvers. Specifically, we introduce a State-Decomposable MDP (SDMDP) that reformulates VRPs by expressing the state space as the Cartesian product of basis state spaces associated with basis VRP variants. More crucially, this formulation inherently yields the optimal basis policy for each basis VRP variant. Furthermore, a Latent Space-based SDMDP extension is developed by incorporating both the optimal basis policies and a learnable mixture function to enable the policy reuse in the latent space. Under mild assumptions, this extension provably recovers the optimal unified policy of SDMDP through the mixture function that computes the state embedding as a mapping from the basis state embeddings generated by optimal basis policies. For practical implementation, we introduce the Mixture-of-Specialized-Experts Solver (MoSES), which realizes basis policies through specialized Low-Rank Adaptation (LoRA) experts, and implements the mixture function via an adaptive gating mechanism. Extensive experiments conducted across VRP variants showcase the superiority of MoSES over prior methods. Yuxin Pan, Zhiguang Cao, Chengyang Gu, Liu Liu 0014, Peilin Zhao, Yize Chen, Fangzhen Lin |
NeurIPS | 4 |
| 2025 | SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo AugmentationabstractEnhancing large language models by simply scaling up datasets has begun to yield diminishing returns, shifting the spotlight to data quality. Monte Carlo Tree Search (MCTS) has emerged as a powerful technique for generating high-quality chain-of-thought data, yet conventional approaches typically retain only the top-scoring trajectory from the search tree, discarding sibling nodes that often contain valuable partial insights, recurrent error patterns, and alternative reasoning strategies. This unconditional rejection of non-optimal reasoning branches may waste vast amounts of informative data in the whole search tree. We propose SIGMA (Sibling Guided Monte Carlo Augmentation), a novel framework that reintegrates these discarded sibling nodes to refine LLM reasoning. SIGMA forges semantic links among sibling nodes along each search path and applies a two-stage refinement: a critique model identifies overlooked strengths and weaknesses across the sibling set, and a revision model conducts text-based backpropagation to refine the top-scoring trajectory in light of this comparative feedback. By recovering and amplifying the underutilized but valuable signals from non-optimal reasoning branches, SIGMA substantially improves reasoning trajectories. On the challenging MATH benchmark, our SIGMA-tuned 7B model achieves 54.92\% accuracy using only 30K samples, outperforming state-of-the-art models trained on 590K samples. This result highlights that our sibling-guided optimization not only significantly reduces data usage but also significantly boosts LLM reasoning. Yanwei Ren, Fuxiang Wu, Jiayan Qiu, Jiaxing Huang 0001, Baosheng Yu, Liu Liu 0014 |
NeurIPS | 7 |
| 2025 | Memory-augmented shuffled meta learning for visible-infrared person re-identification
Hanxiao Wu, Jianqing Zhu, Liu Liu 0014, Huanqiang Zeng |
Neural Networks | 5 |
| 2025 | Textual Embeddings are Good Class-Aware Visual Prompts for Adapting Vision-Language ModelsabstractDue to the parallel nature of the textual and visual encoders, very little attention has been paid to developing prompt learning by using well-pretrained encoders in a serial manner, in which the low-biased high-level semantic information accessible to each other for these encoders is ignored. In this letter, we find that textual embeddings are good class-aware visual prompts for adapting vision-language models, which leads to a new framework called TVPrompt (Textual embeddings as class-aware Visual Prompts). To eliminate the modal gap between text and vision, we design a bridging module, which integrates textual embeddings and class token to produce class-aware visual prompts. To ensure that such prompts could effectively collect class-relevant information, we further propose using masked attention to block the unnecessary interactions. Experimental evidence on benchmark datasets demonstrates that our TVPrompt achieves competitive efficiency and performance. Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Harmonizing Metric Discrepancy for Cross-Modal Object Re-IdentificationabstractVisible and infrared cross-modal re-identification tasks often encounter significant modal discrepancies, which undermine the effectiveness of feature extraction and compromise the reliability of similarity metrics. These discrepancies pose a substantial challenge for accurately matching data across different modalities. To address these issues, we propose a novel approach centered on the maximum mean metric discrepancy (MMMD). We leverage kernel-based statistical techniques to effectively capture and quantify the disparities in cross-modal metrics, providing a robust framework for aligning metrics from different modalities. Building upon the foundation of MMMD, we develop the metric discrepancy harmonization (MDH) method. This method integrates a temperature-controlled optimization technique designed to enhance metric alignment across various modal configurations, ensuring more consistent and reliable performance. By focusing on metric alignment, our approach enhances the accuracy of cross-modal re-identification tasks. Comprehensive evaluations on the LLCM, RGBN300, and SYSU-MM01 datasets demonstrate that our approach achieves state-of-the-art performance. Linhan Huang, Liu Liu 0014, Jianqing Zhu, Huanqiang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Class-Irrelevant Feature Removal for Few-Shot Image ClassificationabstractMost existing few-shot image classification methods employ global pooling to aggregate class-relevant local features in a data-drive manner. Due to the difficulty and inaccuracy in locating class-relevant regions in complex scenarios, as well as the large semantic diversity of local features, the class-irrelevant information could reduce the robustness of the representations obtained by performing global pooling. Meanwhile, the scarcity of labeled images exacerbates the difficulties of data-hungry deep models in identifying class-relevant regions. These issues severely limit deep models' few-shot learning ability. In this work, we propose to remove the class-irrelevant information by making local features class relevant, thus bypassing the big challenge of identifying which local features are class irrelevant. The resulting class-irrelevant feature removal (CIFR) method consists of three phases. First, we employ the masked image modeling strategy to build an understanding of images' internal structures that generalizes well. Second, we design a semantic-complementary feature propagation module to make local features class relevant. Third, we introduce a weighted dense-connected similarity measure, based on which a loss function is raised to fine-tune the entire pipeline, with the aim of further enhancing the semantic consistency of the class-relevant local features. Visualization results show that CIFR achieves the removal of class-irrelevant information by making local features related to classes. Comparison results on four benchmark datasets indicate that CIFR yields very promising performance. Fusheng Hao, Liu Liu 0014, Fuxiang Wu, Qieshi Zhang, Jun Cheng 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Stochastic Optimization for Nonconvex Problem With Inexact Hessian Matrix, Gradient, and FunctionabstractTrust region (TR) and adaptive regularization using cubics (ARC) have proven to have some very appealing theoretical properties for nonconvex optimization by concurrently computing function value, gradient, and Hessian matrix to obtain the next search direction and the adjusted parameters. Although stochastic approximations help largely reduce the computational cost, it is challenging to theoretically guarantee the convergence rate. In this article, we explore a family of stochastic TR (STR) and stochastic ARC (SARC) methods that can simultaneously provide inexact computations of the Hessian matrix, gradient, and function values. Our algorithms require much fewer propagations overhead per iteration than TR and ARC. We prove that the iteration complexity to achieve -approximate second-order optimality is of the same order as the exact computations demonstrated in previous studies. In addition, the mild conditions on inexactness can be met by leveraging a random sampling technology in the finite-sum minimization problem. Numerical experiments with a nonconvex problem support these findings and demonstrate that, with the same or a similar number of iterations, our algorithms require less computational overhead per iteration than current second-order methods. Liu Liu 0014, Xuanqing Liu, Cho-Jui Hsieh, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Pareto Deep Long-Tailed Recognition: A Conflict-Averse SolutionabstractDeep long-tailed recognition (DTLR) has attracted much attention due to its close touch with realistic scenarios. Recent advances have focused on re-balancing across various aspects, e.g., sampling strategy, loss re-weighting, logit adjustment, and input/parameter perturbation, to name a few. However, few studies have considered dynamic re-balancing to address intrinsic optimization conflicts. In this paper, we first empirically argue that the optimizations of mainstream DLTR methods are still dominated by some categories (e.g., major) due to a fixed re-balancing strategy. Thus, they fail to deal with gradient conflicts among categories, which naturally deduces the motivation for reaching Pareto optimal solutions. Unfortunately, a naive integration of multi-objective optimization (MOO) with DLTR methods is not applicable due to the gap between multi-task learning (MTL) and DLTR, and can in turn lead to class-specific feature degradation. Thus, we provide effective alternatives by decoupling MOO-based MTL from the temporal rather than structure perspective, and enhancing it via optimizing variability collapse loss motivated by the derived MOO-based DLTR generalization bound. Moreover, we resort to anticipating worst-case optimization with theoretical insights to further ensure convergence. We build a Pareto deep long-tailed recognition method termed PLOT upon the proposed MOO framework. Extensive evaluations demonstrate that our method not only generally improves mainstream pipelines, but also achieves an augmented version to realize state-of-the-art performance across multiple benchmarks. Liu Liu 0014, Peilin Zhao, Wei Gong 0001 |
ICLR | 2 |
| 2024 | Cross-modal group-relation optimization for visible-infrared person re-identification
Jianqing Zhu, Hanxiao Wu, Yuqing Fu, Huanqiang Zeng, Liu Liu 0014, Zhen Lei 0001 |
Neural Networks | 7 |
| 2024 | Modality-Consistent Attention for Visible-Infrared Vehicle Re-IdentificationabstractVisible-infrared vehicle re-identification (VIVR) seeks to match vehicle images of the same identity taken by cameras of different modalities. The noticeable disparity between visible and infrared modalities leads to attention deviations, causing deep models to incorrectly focus on different local regions of vehicles in visible and infrared images. We observed that the spatial distributions of distinguishing local regions, such as logos, front windows, and wheels, exhibit similarity in average images obtained from both visible and infrared images. Based on this, we propose a modality-consistent attention (MCA) approach for VIVR. Unlike image-level attention, our MCA is identity-level attention that holistically emphasizes the distinguishing regions of a vehicle identity across multiple images captured from various viewpoints. Furthermore, we constrain the differences between the identity-level spatial attention masks resulting from visible and infrared modalities. This approach helps deep networks focus consistently on learning the distinguishing local characteristics of vehicles across different modalities and viewpoints. Our experiments on RGBN300 and MSVR310 datasets demonstrate that our approach achieves state-of-the-art performance. Jiajun Su, Jianqing Zhu, Liu Liu 0014, Huanqiang Zeng |
IEEE Signal Process. Lett. | 4 |
| 2024 | SkipNode: On Alleviating Performance Degradation for Deep Graph Convolutional NetworksabstractGraph Convolutional Networks (GCNs) suffer from performance degradation when models go deeper. However, earlier works only attributed the performance degeneration to over-smoothing. In this paper, we conduct theoretical and experimental analysis to explore the fundamental causes of performance degradation in deep GCNs: over-smoothing and gradient vanishing have a mutually reinforcing effect that causes the performance to deteriorate more quickly in deep GCNs. On the other hand, existing anti-over-smoothing methods all perform full convolutions up to the model depth. They could not well resist the exponential convergence of over-smoothing due to model depth increasing. In this work, we propose a simple yet effective plug-and-play module,SkipNode, to overcome the performance degradation of deep GCNs. It samples graph nodes in each convolutional layer to skip the convolution operation. In this way, both over-smoothing and gradient vanishing can be effectively suppressed since (1) not all nodes'features propagate through full layers and, (2) the gradient can be directly passed back through “skipped” nodes. We provide both theoretical analysis and empirical evaluation to demonstrate the efficacy ofSkipNodeand its superiority over SOTA baselines. Weigang Lu 0001, Yibing Zhan, Binbin Lin 0001, Ziyu Guan, Liu Liu 0014, Baosheng Yu, Wei Zhao 0019, Yaming Yang 0002, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Overcoming Catastrophic Forgetting in Continual Learning by Exploring Eigenvalues of Hessian MatrixabstractNeural networks tend to suffer performance deterioration on previous tasks when they are applied to multiple tasks sequentially without access to previous data. The problem is commonly known as catastrophic forgetting, a significant challenge in continual learning (CL). To overcome the catastrophic forgetting, regularization-based CL methods construct a regularization-based term, which can be considered as the approximation loss function of previous tasks, to penalize the update of parameters. However, the rigorous theoretical analysis of regularization-based methods is limited. Therefore, we theoretically analyze the forgetting and the convergence properties of regularization-based methods. The theoretical results demonstrate that the upper bound of the forgetting has a relationship with the maximum eigenvalue of the Hessian matrix. Hence, to decrease the upper bound of the forgetting, we propose eiGenvalues ExplorAtion Regularization-based (GEAR) method, which explores the geometric properties of the approximation loss of prior tasks regarding the maximum eigenvalue. Extensive experimental results demonstrate that our method mitigates catastrophic forgetting and outperforms existing regularization-based methods. Yajing Kong, Liu Liu 0014, Huanhuan Chen 0001, Janusz Kacprzyk, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Reject Decoding via Language-Vision Models for Text-to-Image SynthesisabstractTransformer-based text-to-image synthesis generates images from abstractive textual conditions and achieves prompt results. Since transformer-based models predict visual tokens step by step in testing, where the early error is hard to be corrected and would be propagated. To alleviate this issue, the common practice is drawing multi-paths from the transformer-based models and re-ranking the multi-images decoded from multi-paths to find the best one and filter out others. Therefore, the computing procedure of excluding images may be inefficient. To improve the effectiveness and efficiency of decoding, we exploit a reject decoding algorithm with tiny multi-modal models to enlarge the searching space and exclude the useless paths as early as possible. Specifically, we build tiny multi-modal models to evaluate the similarities between the partial paths and the caption at multi scales. Then, we propose a reject decoding algorithm to exclude some lowest quality partial paths at the inner steps. Thus, under the same computing load as the original decoding, we could search across more multi-paths to improve the decoding efficiency and synthesizing quality. The experiments conducted on the MS-COCO dataset and large-scale datasets show that the proposed reject decoding algorithm can exclude the useless paths and enlarge the searching paths to improve the synthesizing quality by consuming less time. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Lei Wang 0018, Jun Cheng 0002 |
AAAI | 2 |
| 2023 | Class-Aware Patch Embedding Adaptation for Few-Shot Image Classificationabstract"A picture is worth a thousand words", significantly beyond mere a categorization. Accompanied by that, many patches of the image could have completely irrelevant meanings with the categorization if they were independently observed. This could significantly reduce the efficiency of a large family of few-shot learning algorithms, which have limited data and highly rely on the comparison of image patches. To address this issue, we propose a Class-aware Patch Embedding Adaptation (CPEA) method to learn "class-aware embeddings" of the image patches. The key idea of CPEA is to integrate patch embeddings with class-aware embeddings to make them class-relevant. Furthermore, we define a dense score matrix between class-relevant patch embeddings across images, based on which the degree of similarity between paired images is quantified. Visualization results show that CPEA concentrates patch embeddings by class, thus making them class-relevant. Extensive experiments on four benchmark datasets, miniImageNet, tieredImageNet, CIFAR-FS, and FC-100, indicate that our CPEA significantly outperforms the existing state-of-the-art methods. The source code is available at https://github.com/FushengHao/CPEA. Fusheng Hao, Fengxiang He, Liu Liu 0014, Fuxiang Wu, Dacheng Tao, Jun Cheng 0002 |
ICCV | 3 |
| 2023 | BEEF: Bi-Compatible Class-Incremental Learning via Energy-Based Expansion and Fusion
Fu-Yun Wang, Da-Wei Zhou 0001, Liu Liu 0014, Han-Jia Ye, Yatao Bian, De-Chuan Zhan, Peilin Zhao |
ICLR | 3 |
| 2023 | Trust-Region Adaptive Frequency for Online Continual LearningabstractAbstract In the paradigm of online continual learning, one neural network is exposed to a sequence of tasks, where the data arrive in an online fashion and previously seen data are not accessible. Such online fashion causes insufficient learning and severe forgetting on past tasks issues, preventing a good stability-plasticity trade-off, where ideally the network is expected to have high plasticity to adapt to new tasks well and have the stability to prevent forgetting on old tasks simultaneously. To solve these issues, we propose a trust-region adaptive frequency approach, which alternates between standard-process and intra-process updates. Specifically, the standard-process replays data stored in a coreset and interleaves the data with current data, and the intra-process updates the network parameters based on the coreset. Furthermore, to improve the unsatisfactory performance stemming from online fashion, the frequency of the intra-process is adjusted based on a trust region, which is measured by the confidence score of current data. During the intra-process, we distill the dark knowledge to retain useful learned knowledge. Moreover, to store more representative data in the coreset, a confidence-based coreset selection is presented in an online manner. The experimental results on standard benchmarks show that the proposed method significantly outperforms state-of-art continual learning algorithms. Yajing Kong, Liu Liu 0014, Maoying Qiao, Zhen Wang 0030, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2023 | Attribute-Image Person Re-identification via Modal-Consistent Metric Learning
Jianqing Zhu, Liu Liu 0014, Yibing Zhan, Xiaobin Zhu 0001, Huanqiang Zeng, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2023 | On exploring node-feature and graph-structure diversities for node drop graph pooling
Chuang Liu 0008, Yibing Zhan, Baosheng Yu, Liu Liu 0014, Bo Du 0001, Wenbin Hu 0001, Tongliang Liu |
Neural Networks | 4 |
| 2023 | Prescribed Safety Performance Imitation Learning From a Single Expert DatasetabstractExisting safe imitation learning (safe IL) methods mainly focus on learning safe policies that are similar to expert ones, but may fail in applications requiring different safety constraints. In this paper, we propose the Lagrangian Generative Adversarial Imitation Learning (LGAIL) algorithm, which can adaptively learn safe policies from a single expert dataset under diverse prescribed safety constraints. To achieve this, we augment GAIL with safety constraints and then relax it as an unconstrained optimization problem by utilizing a Lagrange multiplier. The Lagrange multiplier enables explicit consideration of the safety and is dynamically adjusted to balance the imitation and safety performance during training. Then, we apply a two-stage optimization framework to solve LGAIL: (1) a discriminator is optimized to measure the similarity between the agent-generated data and the expert ones; (2) forward reinforcement learning is employed to improve the similarity while considering safety concerns enabled by a Lagrange multiplier. Furthermore, theoretical analyses on the convergence and safety of LGAIL demonstrate its capability of adaptively learning a safe policy given prescribed safety constraints. At last, extensive experiments in OpenAI Safety Gym conclude the effectiveness of our approach. Zhihao Cheng, Li Shen 0008, Miaoxi Zhu, Jiaxian Guo, Liu Liu 0014, Bo Du 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Mixer-Based Semantic Spread for Few-Shot LearningabstractKey semantics can come from everywhere on an image. Semantic alignment is a key part of few-shot learning but still remains challenging. In this paper, we design a Mixer-Based Semantic Spread (MBSS) algorithm that employs amixermodule to spread the key semantic on the whole image, so that one can directly compare the processed image pairs. We first adopt a convolutional neural network to extract features from both support and query images and separate each of them into multiple Local Descriptor-based Representations (LDRs). The LDRs are then fed into themixerfor semantic spread, where every LDR attracts complementary information from its peers. In this way, the objective semantic is made spread on the whole image in a data-driven manner. The overall pipeline is supervised by a voting-based loss, guaranteeing a goodmixer. Visualization results validate the feasibility of ourmixer. Comprehensive experiments on three benchmark datasets, miniImageNet, tieredImageNet, and CUB, show that our algorithm achieves the state-of-the-art performance in both 5-way 1-shot and 5-way 5-shot settings. Jun Cheng 0002, Fusheng Hao, Fengxiang He, Liu Liu 0014, Qieshi Zhang |
IEEE Trans. Multim. | 4 |
| 2023 | InDecGAN: Learning to Generate Complex Images From Captions via Independent Object-Level Decomposition and EnhancementabstractText-to-image synthesis is a challenging problem, in which a complex scene contains diverse objects of various sizes and sub-images of objects belonging to the same class have diverse forms from different perspectives. Thus, synthesis models have difficulty in capturing varied objects in the complex scene. To alleviate these problems, we devise an independent object-level decomposing and enhancing generative adversarial networks, denoted as InDecGAN, to synthesize complex images and capture varied objects in a complex scene. Specifically, InDecGAN fully utilizes the independent object-level information, bounding boxes and high-resolution images of objects in training, by employing independent object-level pathways to synthesize varied objects. The independent object-level pathway integrates an independent object-level adversarial loss and the bounding box information to learn the visual features of objects independently, then, the main pathway exploits the features provided by the object-level pathway to compose the full scene and synthesize images. In addition, we analyze the generalization properties of the proposed InDecGAN and demonstrate the improvement from the perspective of the model architecture. Moreover, extensive experiments conducted on a widely used dataset are presented to demonstrate that the proposed model with an independent object-level pathway produces synthesized images of significantly improved quality. Jun Cheng 0002, Fuxiang Wu, Liu Liu 0014, Qieshi Zhang, Leszek Rutkowski, Dacheng Tao |
IEEE Trans. Multim. | 3 |
| 2023 | Language-Based Image Manipulation Built on Language-Guided RankingabstractText-based image manipulation is a popular subject and has many applications. However, it is a challenging task because there is no ground-truth edited dataset and textual descriptions have abstractive and ambiguous properties. To alleviate the difficult issues, we propose a manipulation framework consisting of the proposal attentional GANs, language-related semantic mask, and language-guided ranker. Specially, we construct an editing proposal generator to generate the suitable edited proposals with and without semantic conditions, which supports the reorganization of sub-generators to output proposals in various aspects as many as possible. To distinguish the text-relevant and the text-irrelevant regions, we introduce a language-related semantic mask based on the source image and target caption. Then, we exploit a language-guided ranker to retrieve the best edited result from the edited proposals through using the multi-modal similarity and the language-related semantic mask. Extensive experiments on widely-used datasets demonstrate that our model could manipulate images interactively and improve the editing quality effectively. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Jun Cheng 0002 |
IEEE Trans. Multim. | 2 |
| 2023 | On the Guaranteed Almost Equivalence Between Imitation Learning From Observation and DemonstrationabstractImitation learning from observation (LfO) is more preferable than imitation learning from demonstration (LfD) because of the nonnecessity of expert actions when reconstructing the expert policy from the expert data. However, previous studies imply that the performance of LfO is inferior to LfD by a tremendous gap, which makes it challenging to employ LfO in practice. By contrast, this article proves that LfO is almost equivalent to LfD in the deterministic robot environment, and more generally even in the robot environment with bounded randomness. In the deterministic robot environment, from the perspective of the control theory, we show that the inverse dynamics disagreement between LfO and LfD approaches zero, meaning that LfO is almost equivalent to LfD. To further relax the deterministic constraint and better adapt to the practical environment, we consider bounded randomness in the robot environment and prove that the optimizing targets for both LfD and LfO remain almost the same in the more generalized setting. Extensive experiments for multiple robot tasks are conducted to demonstrate that LfO achieves comparable performance to LfD empirically. In fact, the most common robot systems in reality are the robot environment with bounded randomness (i.e., the environment this article considered). Hence, our findings greatly extend the potential of LfO and suggest that we can safely apply LfO in practice without sacrificing the performance compared to LfD. Zhihao Cheng, Liu Liu 0014, Aishan Liu, Hao Sun 0019, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Stochastically Controlled Compositional Gradient for Composition ProblemsabstractWe consider composition problems of the form$(1/n)\sum _{i= 1}^{n} F_{i} (1/m)\sum _{j = 1}^{m} G_{j}(x)$, which are important for machine learning. Although gradient descent and stochastic gradient descent are straightforward solutions, the essential computation of$G (x)= (1/m)\sum _{j = 1}^{m}{G_{j}(x)}$in each single iteration is expensive, let alone for large$m$. In this article, we devise a stochastically controlled compositional gradient algorithm. Specifically, we introduce two variants of stochastically controlled technique to estimate the inner function$G(x)$and the gradient of the objective function, respectively. The computational cost is largely reduced. However, the natural needs of two stochastic subsets${\mathcal D}_{1}$and${\mathcal D}_{2}$form direct barriers to guarantee the convergence of the algorithm, especially the theoretical proof of the convergence. To this end, we present a general convergence analysis by proving$|{\mathcal{ D}}_{1}|=\min \{1/\epsilon,m\}$and$|{\mathcal{ D}}_{2}|=\min \{1/\epsilon,n \}$, through which the proposed method significantly improve composition algorithms under low target accuracy (i.e.,$1/\epsilon \ll m$or$n$) in both strongly convex and nonconvex settings. Comprehensive experiments demonstrate the superiority of the proposed method over existing methods. Liu Liu 0014, Ji Liu 0002, Cho-Jui Hsieh, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | SIN: Semantic Inference Network for Few-Shot Streaming Label LearningabstractStreaming label learning aims to model newly emerged labels for multilabel classification systems, which requires plenty of new label data for training. However, in changing environments, only a small amount of new label data can practically be collected. In this work, we formulate and study few-shot streaming label learning (FSLL), which models emerging new labels with only a few annotated examples by utilizing the knowledge learned from past labels. We propose a meta-learning framework, semantic inference network (SIN), which can learn and infer the semantic correlation between new labels and past labels to adapt FSLL tasks from a few examples effectively. SIN leverages label semantic representation to regularize the output space and acquires labelwise meta-knowledge based on gradient-based meta-learning. Moreover, SIN incorporates a novel label decision module with a meta-threshold loss to find the optimal confidence thresholds for each new label. Theoretically, we illustrate that the proposed semantic inference mechanism could constrain the complexity of hypotheses space to reduce the risk of overfitting and achieve better generalizability. Experimentally, extensive empirical results and ablation studies demonstrate the performance of SIN is superior to the prior state-of-the-art methods on FSLL. Zhen Wang 0030, Liu Liu 0014, Yiqun Duan, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | The Visual Footsteps Planning System for Exoskeleton Robots Under Complex TerrainabstractThe lower limb power-assist exoskeletons are expected to help paraplegic people to walk again in daily life. However, most of these exoskeletons deal with walking in the scene that has been seen or has an external vision sensor, rather than in the unknown environment. It is a great challenge to understand the wear’s intention and plan the footstep sequence in an unknown scene. Moreover, the traditional visual footstep planning is dominated by the robot, which can lead to an awkward trajectory plan. Therefore, we construct a visual footstep planning system and propose an onboard vision planning algorithm based on the Bezier curve to address the previous two problems. Specially, our human–computer interaction system understands the environment and the wearer’s behavior intention by integrating Hololens and Realsense. Then, we apply the Bezier curve to plan footsteps for the first time in the field of the exoskeleton and define two parameters of the Bezier curve, which are more suitable for our exoskeleton system and could increase the planning speed. Finally, we add the tracking feature cost in the cost function, which could better fit the planned footprints to the planned path and make the gait smoother. Extended experimental results show that the average planning time is 67.46% less than that of the traditional search algorithm. Moreover, the effectiveness of our system is also verified on the visual interaction platform. Xinyu Wu 0001, Liu Liu 0014, Dacheng Tao |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Continual Learning through Retrieval and ImaginationabstractContinual learning is an intellectual ability of artificial agents to learn new streaming labels from sequential data. The main impediment to continual learning is catastrophic forgetting, a severe performance degradation on previously learned tasks. Although simply replaying all previous data or continuously adding the model parameters could alleviate the issue, it is impractical in real-world applications due to the limited available resources. Inspired by the mechanism of the human brain to deepen its past impression, we propose a novel framework, Deep Retrieval and Imagination (DRI), which consists of two components: 1) an embedding network that constructs a unified embedding space without adding model parameters on the arrival of new tasks; and 2) a generative model to produce additional (imaginary) data based on the limited memory. By retrieving the past experiences and corresponding imaginary data, DRI distills knowledge and rebalances the embedding space to further mitigate forgetting. Theoretical analysis demonstrates that DRI can reduce the loss approximation error and improve the robustness through retrieval and imagination, bringing better generalizability to the network. Extensive experiments show that DRI performs significantly better than the existing state-of-the-art continual learning methods and effectively alleviates catastrophic forgetting. Zhen Wang 0030, Liu Liu 0014, Yiqun Duan, Dacheng Tao |
AAAI | 2 |
| 2022 | Resistance Training Using Prior Bias: Toward Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) aims to build a structured representation of a scene using objects and pairwise relationships, which benefits downstream tasks. However, current SGG methods usually suffer from sub-optimal scene graph generation because of the long-tailed distribution of training data. To address this problem, we propose Resistance Training using Prior Bias (RTPB) for the scene graph generation. Specifically, RTPB uses a distributed-based prior bias to improve models' detecting ability on less frequent relationships during training, thus improving the model generalizability on tail categories. In addition, to further explore the contextual information of objects and relationships, we design a contextual encoding backbone network, termed as Dual Transformer (DTrans). We perform extensive experiments on a very popular benchmark, VG150, to demonstrate the effectiveness of our method for the unbiased scene graph generation. In specific, our RTPB achieves an improvement of over 10% under the mean recall when applied to current SGG methods. Furthermore, DTrans with RTPB outperforms nearly all state-of-the-art methods with a large margin. Code is available at https://github.com/ChCh1999/RTPB Yibing Zhan, Baosheng Yu, Liu Liu 0014, Yong Luo 0002, Bo Du 0001 |
AAAI | 4 |
| 2022 | Channelized Axial Attention - considering Channel Relation within Spatial Attention for Semantic SegmentationabstractSpatial and channel attentions, modelling the semantic interdependencies in spatial and channel dimensions respectively, have recently been widely used for semantic segmentation. However, computing spatial and channel attentions separately sometimes causes errors, especially for those difficult cases. In this paper, we propose Channelized Axial Attention (CAA) to seamlessly integrate channel attention and spatial attention into a single operation with negligible computation overhead. Specifically, we break down the dot-product operation of the spatial attention into two parts and insert channel relation in between, allowing for independently optimized channel attention on each spatial location. We further develop grouped vectorization, which allows our model to run with very little memory consumption without slowing down the running speed. Comparative experiments conducted on multiple benchmark datasets, including Cityscapes, PASCAL Context, and COCO-Stuff, demonstrate that our CAA outperforms many state-of-the-art segmentation models (including dual attention) on all tested datasets. Wenjing Jia, Liu Liu 0014, Xiangjian He |
AAAI | 4 |
| 2022 | Continual Learning with Lifelong Vision TransformerabstractContinual learning methods aim at training a neural network from sequential data with streaming labels, relieving catastrophic forgetting. However, existing methods are based on and designed for convolutional neural networks (CNNs), which have not utilized the full potential of newly emerged powerful vision transformers. In this paper, we propose a novel attention-based framework Lifelong Vision Transformer (LVT), to achieve a better stability-plasticity trade-off for continual learning. Specifically, an inter-task attention mechanism is presented in LVT, which implicitly absorbs the previous tasks' information and slows down the drift of important attention between previous tasks and the current task. LVT designs a dual-classifier structure that independently injects new representation to avoid catas-trophic interference and accumulates the new and previous knowledge in a balanced manner to improve the overall performance. Moreover, we develop a confidence-aware memory update strategy to deepen the impression of the previous tasks. The extensive experimental results show that our approach achieves state-of-the-art performance with even fewer parameters on continual learning benchmarks. Zhen Wang 0030, Liu Liu 0014, Yiqun Duan, Yajing Kong, Dacheng Tao |
CVPR | 2 |
| 2022 | Text-to-Image Synthesis based on Object-Guided Joint-Decoding TransformerabstractObject-guided text-to-image synthesis aims to generate images from natural language descriptions built by two-step frameworks, i.e., the model generates the layout and then synthesizes images from the layout and captions. However, such frameworks have two issues: 1) complex structure, since generating language-related layout is not a trivial task; 2) error propagation, because the inappropriate layout will mislead the image synthesis and is hard to be revised. In this paper, we propose an object-guided joint-decoding module to simultaneously generate the image and the corresponding layout. Specially, we present the joint-decoding transformer to model the joint probability on images tokens and the corresponding layouts tokens, where layout tokens provide additional observed data to model the complex scene better. Then, we describe a novel Layout-Vqgan for layout encoding and decoding to provide more information about the complex scene. After that, we present the detail-enhanced module to enrich the language-related details based on two facts: 1) visual details could be omitted in the compression of VQGANs; 2) the joint-decoding transformer would not have sufficient generating capacity. The experiments show that our approach is competitive with previous object-centered models and can generate diverse and high-quality objects under the given layouts. Fuxiang Wu, Liu Liu 0014, Fusheng Hao, Fengxiang He, Jun Cheng 0002 |
CVPR | 2 |
| 2022 | Balancing Stability and Plasticity Through Advanced Null Space in Continual Learning
Yajing Kong, Liu Liu 0014, Zhen Wang 0030, Dacheng Tao |
ECCV (26) | 2 |
| 2022 | Online Continual Learning with Contrastive Vision Transformer
Zhen Wang 0030, Liu Liu 0014, Yajing Kong, Jiaxian Guo, Dacheng Tao |
ECCV (20) | 2 |
| 2022 | UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware MixupabstractSubpopulation shift widely exists in many real-world machine learning applications, referring to the training and test distributions containing the same subpopulation groups but varying in subpopulation frequencies. Importance reweighting is a normal way to handle the subpopulation shift issue by imposing constant or adaptive sampling weights on each sample in the training dataset. However, some recent studies have recognized that most of these approaches fail to improve the performance over empirical risk minimization especially when applied to over-parameterized neural networks. In this work, we propose a simple yet practical framework, called uncertainty-aware mixup (UMIX), to mitigate the overfitting issue in over-parameterized models by reweighting the ''mixed'' samples according to the sample uncertainty. The training-trajectories-based uncertainty estimation is equipped in the proposed UMIX for each sample to flexibly characterize the subpopulation distribution. We also provide insightful theoretical analysis to verify that UMIX achieves better generalization bounds over prior works. Further, we conduct extensive empirical studies across a wide range of tasks to validate the effectiveness of our method both qualitatively and quantitatively. Code is available at https://github.com/TencentAILabHealthcare/UMIX. Zongbo Han, Fan Yang 0081, Liu Liu 0014, Lanqing Li, Yatao Bian, Peilin Zhao, Bingzhe Wu, Changqing Zhang 0002, Jianhua Yao 0001 |
NeurIPS | 4 |
| 2022 | Escaping from the Barren Plateau via Gaussian Initializations in Deep Variational Quantum CircuitsabstractVariational quantum circuits have been widely employed in quantum simulation and quantum machine learning in recent years. However, quantum circuits with random structures have poor trainability due to the exponentially vanishing gradient with respect to the circuit depth and the qubit number. This result leads to a general standpoint that deep quantum circuits would not be feasible for practical tasks. In this work, we propose an initialization strategy with theoretical guarantees for the vanishing gradient problem in general deep quantum circuits. Specifically, we prove that under proper Gaussian initialized parameters, the norm of the gradient decays at most polynomially when the qubit number and the circuit depth increase. Our theoretical results hold for both the local and the global observable cases, where the latter was believed to have vanishing gradients even for very shallow circuits. Experimental results verify our theoretical findings in quantum simulation and quantum chemistry. Kaining Zhang, Liu Liu 0014, Min-Hsiu Hsieh, Dacheng Tao |
NeurIPS | 2 |
| 2022 | Dual-branch Density Ratio Estimation for Signed Network EmbeddingabstractSigned network embedding (SNE) has received considerable attention in recent years. A mainstream idea of SNE is to learn node representations by estimating the ratio of sampling densities. Though achieving promising performance, these methods based on density ratio estimation are limited to the issues of confusing sample, expected error, and fixed priori. To alleviate the above-mentioned issues, in this paper, we propose a novel dual-branch density ratio estimation (DDRE) architecture for SNE. Specifically, DDRE 1) consists of a dual-branch network, dealing with the confusing sample; 2) proposes the expected matrix factorization without sampling to avoid the expected error; and 3) devises an adaptive cross noise sampling to alleviate the fixed priori. We perform sign prediction and node classification experiments on four real-world and three artificial datasets, respectively. Extensive empirical results demonstrate that DDRE not only significantly outperforms the methods based on density ratio estimation but also achieves competitive performance compared with other types of methods such as graph likelihood, generative adversarial networks, and graph convolutional networks. Code is publicly available at https://github.com/WHU-SNA/DDRE. Pinghua Xu, Yibing Zhan, Liu Liu 0014, Baosheng Yu, Bo Du 0001, Jia Wu 0001, Wenbin Hu 0001 |
WWW | 3 |
| 2022 | Variance Reduced Methods for Non-Convex Composition OptimizationabstractThis paper explores the non-convex composition optimization consisting of inner and outer finite-sum functions with a large number of component functions. This problem arises in important applications such as nonlinear embedding and reinforcement learning. Although existing approaches such as stochastic gradient descent (SGD) and stochastic variance reduced gradient (SVRG) descent can be applied to solve this problem, their query complexities tend to be high, especially when the number of inner component functions is large. Therefore, to significantly improve the query complexity of current approaches, we have devised the stochastic composition via variance reduction (SCVR). What's more, we analyze the query complexity under different numbers of inner function and outer function. Based on different kinds of estimation of inner component function, we also present the SCVRII algorithm, though the order of query complexities are the same with SCVR. Additionally, we propose an extension to handle the mini-batch cases, which improve the query complexity under the optimal mini-batch size. The experimental results validate our proposed algorithms and theoretical analyses. Liu Liu 0014, Ji Liu 0002, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Imposing Semantic Consistency of Local Descriptors for Few-Shot LearningabstractFew-shot learning suffers from the scarcity of labeled training data. Regarding local descriptors of an image as representations for the image could greatly augment existing labeled training data. Existing local descriptor based few-shot learning methods have taken advantage of this fact but ignore that the semantics exhibited by local descriptors may not be relevant to the image semantic. In this paper, we deal with this issue from a new perspective of imposing semantic consistency of local descriptors of an image. Our proposed method consists of three modules. The first one is a local descriptor extractor module, which can extract a large number of local descriptors in a single forward pass. The second one is a local descriptor compensator module, which compensates the local descriptors with the image-level representation, in order to align the semantics between local descriptors and the image semantic. The third one is a local descriptor based contrastive loss function, which supervises the learning of the whole pipeline, with the aim of making the semantics carried by the local descriptors of an image relevant and consistent with the image semantic. Theoretical analysis demonstrates the generalization ability of our proposed method. Comprehensive experiments conducted on benchmark datasets indicate that our proposed method achieves the semantic consistency of local descriptors and the state-of-the-art performance. Jun Cheng 0002, Fusheng Hao, Liu Liu 0014, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2021 | Multi-label Few-shot Learning with Semantic Inference (Student Abstract)abstractFew-shot learning can adapt the classification model to new labels with only a few labeled examples. Previous studies mainly focus on the scenario of a single category label per example but have not solved the more challenging multi-label scenario with exponential-sized output space and low-data effectively. In this paper, we propose a semantic-aware meta-learning model for multi-label few-shot learning. Our approach can learn and infer the semantic correlation between unseen labels and historical labels to quickly adapt multi-label tasks from only a few examples. Specifically, features can be mapped into the semantic embedding space via label word vectors to explore and exploit the label correlation, and thus cope with the challenge on the overwhelming size of the output space. Then a novel semantic inference mechanism is designed for leveraging prior knowledge learned from historical labels, which will produce good generalization performance on new labels to alleviate the low-data problem. Finally, extensive empirical results show that the proposed method significantly outperforms the existing state-of-the-art methods on the multi-label few-shot learning tasks. Zhen Wang 0030, Yiqun Duan, Liu Liu 0014, Dacheng Tao |
AAAI | 3 |
| 2021 | Adaptive Curriculum LearningabstractInspired by the human learning principle that learning easier concepts first and then gradually paying more attention to harder ones, curriculum learning uses the nonuniform sampling of mini-batches according to the order of examples’ difficulty. Just as a teacher adjusts the curriculum according to the learning progress of each student, a proper curriculum should be adapted to the current state of the model. Therefore, in contrast to recent works using a fixed curriculum, we devise a new curriculum learning method, Adaptive Curriculum Learning (Adaptive CL), adapting the difficulty of examples to the current state of the model. Specifically, we make use of the loss of the current model to adjust the difficulty score while retaining previous useful learned knowledge by KL divergence. Moreover, under a non-linear model and binary classification, we theoretically prove that the expected convergence rate of curriculum learning monotonically decreases with respect to the loss of a point regarding the optimal hypothesis, and monotonically increases with respect to the loss of a point regarding the current hypothesis. The analyses indicate that Adaptive CL could improve the convergence properties during the early stages of learning. Extensive experimental results demonstrate the superiority of the proposed approach over existing competitive curriculum learning methods. Yajing Kong, Liu Liu 0014, Jun Wang 0002, Dacheng Tao |
ICCV | 2 |
| 2021 | Contrastive Graph Poisson Networks: Semi-Supervised Learning with Extremely Limited LabelsabstractGraph Neural Networks (GNNs) have achieved remarkable performance in the task of semi-supervised node classification. However, most existing GNN models require sufficient labeled data for effective network training. Their performance can be seriously degraded when labels are extremely limited. To address this issue, we propose a new framework termed Contrastive Graph Poisson Networks (CGPN) for node classification under extremely limited labeled data. Specifically, our CGPN derives from variational inference; integrates a newly designed Graph Poisson Network (GPN) to effectively propagate the limited labels to the entire graph and a normal GNN, such as Graph Attention Network, that flexibly guides the propagation of GPN; applies a contrastive objective to further exploit the supervision information from the learning process of GPN and GNN models. Essentially, our CGPN can enhance the learning performance of GNNs under extremely limited labels by contrastively propagating the limited labels to the entire graph. We conducted extensive experiments on different types of datasets to demonstrate the superiority of CGPN. Sheng Wan, Yibing Zhan, Liu Liu 0014, Baosheng Yu, Shirui Pan, Chen Gong 0002 |
NeurIPS | 3 |
| 2021 | A spatial structural similarity triplet loss for auxiliary vehicle re-identification
Jianqing Zhu, Liu Liu 0014, Xiaobin Zhu 0001, Huanqiang Zeng |
Sci. China Inf. Sci. | 2 |
| 2021 | Solving Jigsaw Puzzles via Nonconvex Quadratic Programming With the Projected Power MethodabstractJigsaw puzzles consist of reconstructing a picture that has been divided into many interlocking pieces. This paper describes an automatic global method for solving the square-piece jigsaw puzzle problem in which neither the orientations nor the locations of the jigsaw pieces are known. This hard combinatorial sorting task is formulated as a nonconvex quadratic programming problem that is solved via the projected power method. Specifically, this work aims to specify the locations and orientations of puzzle pieces by maximizing a constrained quadratic function that resolves an optimized permutation matrix composed of the noisy pairwise affinities between jigsaw pieces. The experimental results obtained in the MIT, McGill and Pomeranz datasets indicate that our method outperforms state-of-the-art techniques. Fang Yan 0003, Yuanjie Zheng, Jinyu Cong, Liu Liu 0014, Dacheng Tao, Sujuan Hou |
IEEE Trans. Multim. | 4 |
| 2020 | Deep Streaming Label LearningabstractIn multi-label learning, each instance can be associated with multiple and non-exclusive labels. Previous studies assume that all the labels in the learning process are fixed and static; however, they ignore the fact that the labels will emerge continuously in changing environments. In order to fill in these research gaps, we propose a novel deep neural network (DNN) based framework, Deep Streaming Label Learning (DSLL), to classify instances with newly emerged labels effectively. DSLL can explore and incorporate the knowledge from past labels and historical models to understand and develop emerging new labels. DSLL consists of three components: 1) a streaming label mapping to extract deep relationships between new labels and past labels with a novel label-correlation aware loss; 2) a streaming feature distillation propagating feature-level knowledge from the historical model to a new model; 3) a senior student network to model new labels with the help of knowledge learned from the past. Theoretically, we prove that DSLL admits tight generalization error bounds for new labels in the DNN framework. Experimentally, extensive empirical results show that the proposed method performs significantly better than the existing state-of-the-art multi-label learning methods to handle the continually emerging new labels. Zhen Wang 0030, Liu Liu 0014, Dacheng Tao |
ICML | 2 |
| 2019 | Dualityfree Methods for Stochastic Composition OptimizationabstractIn this paper, we consider the composition optimization with two expected-value functions in the form of$({1}/{n})\sum _{i = 1}^{n} F_{i}\left({({1}/{m})\sum _{j = 1}^{m} G_{j}(x)}\right)+R(x)$, which formulates many important problems in statistical learning and machine learning such as solving Bellman equations in reinforcement learning and nonlinear embedding. Full gradient- or classical stochastic gradient descent-based optimization algorithms are unsuitable or computationally expensive to solve this problem due to the inner expectation$({1}/{m})\sum _{j = 1}^{m} G_{j}(x)$. We propose a dualityfree-based stochastic composition method that combines the variance reduction methods to address the stochastic composition problem. We apply the stochastic variance reduction gradient- and stochastic average gradient algorithm-based methods to estimate the inner function and the dualityfree method to estimate the outer function. We prove the linear convergence rate not only for the convex composition problem but also for the case that the individual outer functions are nonconvex, while the objective function is strongly convex. We also provide the results of experiments that show the effectiveness of our proposed methods. Liu Liu 0014, Ji Liu 0002, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | GoDec+: Fast and Robust Low-Rank Matrix Decomposition Based on Maximum CorrentropyabstractGoDec is an efficient low-rank matrix decomposition algorithm. However, optimal performance depends on sparse errors and Gaussian noise. This paper aims to address the problem that a matrix is composed of a low-rank component and unknown corruptions. We introduce a robust local similarity measure called correntropy to describe the corruptions and, in doing so, obtain a more robust and faster low-rank decomposition algorithm: GoDec+. Based on half-quadratic optimization and greedy bilateral paradigm, we deliver a solution to the maximum correntropy criterion (MCC)-based low-rank decomposition problem. Experimental results show that GoDec+ is efficient and robust to different corruptions including Gaussian noise, Laplacian noise, salt & pepper noise, and occlusion on both synthetic and real vision data. We further apply GoDec+ to more general applications including classification and subspace clustering. For classification, we construct an ensemble subspace from the low-rank GoDec+ matrix and introduce an MCC-based classifier. For subspace clustering, we utilize GoDec+ values low-rank matrix for MCC-based self-expression and combine it with spectral clustering. Face recognition, motion segmentation, and face clustering experiments show that the proposed methods are effective and robust. In particular, we achieve the state-of-the-art performance on the Hopkins 155 data set and the first 10 subjects of extended Yale B for subspace clustering. Kailing Guo, Liu Liu 0014, Xiangmin Xu 0001, Dong Xu 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Diversified dictionaries for multi-instance learning
Maoying Qiao, Liu Liu 0014, Jun Yu 0002, Chang Xu 0002, Dacheng Tao |
Pattern Recognit. | 2 |
| 2017 | Label Propagation via Teaching-to-Learn and Learning-to-TeachabstractHow to propagate label information from labeled examples to unlabeled examples over a graph has been intensively studied for a long time. Existing graph-based propagation algorithms usually treat unlabeled examples equally, and transmit seed labels to the unlabeled examples that are connected to the labeled examples in a neighborhood graph. However, such a popular propagation scheme is very likely to yield inaccurate propagation, because it falls short of tackling ambiguous but critical data points (e.g., outliers). To this end, this paper treats the unlabeled examples in different levels of difficulties by assessing their reliability and discriminability, and explicitly optimizes the propagation quality by manipulating the propagation sequence to move from simple to difficult examples. In particular, we propose a novel iterative label propagation algorithm in which each propagation alternates between two paradigms, teaching-to-learn and learning-to-teach (TLLT). In the teaching-to-learn step, the learner conducts the propagation on the simplest unlabeled examples designated by the teacher. In the learning-to-teach step, the teacher incorporates the learner's feedback to adjust the choice of the subsequent simplest examples. The proposed TLLT strategy critically improves the accuracy of label propagation, making our algorithm substantially robust to the values of tuning parameters, such as the Gaussian kernel width used in graph construction. The merits of our algorithm are theoretically justified and empirically demonstrated through experiments performed on both synthetic and real-world data sets. Chen Gong 0002, Dacheng Tao, Wei Liu 0005, Liu Liu 0014, Jie Yang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | Sublinear Dual Coordinate Ascent for Regularized Loss MinimizationabstractWe present a sublinear version of the dual coordinate ascent method for solving a group of regularized loss minimization problems in machine learning. The proposed method seamlessly integrates sampling techniques, the dual coordinate ascent method, and a multiplicative update algorithm. The sampling techniques choose the "expected" examples, and estimate the corresponding inner products. The dual coordinate ascent method generates an updated iterative step, which outperforms the time-learning step used in the previous sublinear perceptron algorithm. The multiplicative update algorithm updates the example weighting. The proposed method is implemented with an iterative step of order O(log(n)), where n is the size of examples, and achieves a better result than other methods, with high probability. We present a theoretical analysis of the sublinear iterative in order to justify its benefits. We then apply the proposed optimization method to support vector machine and conduct experiments on three large-scale datasets. Our experimental results validate our theoretical findings. Liu Liu 0014, Dacheng Tao |
ICDM | 1 |