Zhongzhan Huang

dblp:241/9753 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
abstract
Reconstructing animatable 3D animals from videos traditionally depends on sparse semantic keypoints to fit parametric models. Acquiring these keypoints is labor-intensive, and detectors trained on limited animal datasets are often unreliable. We propose 4D-Animal, a keypoint-free framework that reconstructs animatable 3D animals directly from videos. Our method employs a dense feature network to map 2D image representations to SMAL parameters, improving both efficiency and stability. Additionally, we introduce a hierarchical alignment strategy that leverages silhouette, part-level, pixel-level, and temporal cues from pretrained 2D models, ensuring accurate and temporally coherent reconstructions. Extensive experiments demonstrate that 4D-Animal outperforms both model-based and model-free baselines on dog dataset. Moreover, the high-quality 3D assets generated by our method can benefit other 3D tasks, underscoring its potential for large-scale applications. The code is released at https://github.com/zhongshsh/4D-Animal.
Shanshan Zhong, Zehan Zheng, Zhongzhan Huang, Wufei Ma, Guofeng Zhang 0025, Qihao Liu, Alan L. Yuille, Jieneng Chen
WACV4
2025 MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
abstract
Long Context Understanding (LCU) is a critical area for exploration in current large language models (LLMs).However, due to the inherently lengthy nature of long-text data, existing LCU benchmarks for LLMs often result in prohibitively high evaluation costs, like testing time and inference expenses.Through extensive experimentation, we discover that existing LCU benchmarks exhibit significant redundancy, which means the inefficiency in evaluation.In this paper, we propose a concise data compression method tailored for longtext data with sparse information characteristics.By pruning the well-known LCU benchmark LongBench, we create MiniLongBench.This benchmark includes only 237 test samples across six major task categories and 21 distinct tasks.Through empirical analysis of over 60 LLMs, MiniLongBench achieves an average evaluation cost reduced to only 4.5% of the original while maintaining an average rank correlation coefficient of 0.97 with Long-Bench results.Therefore, our MiniLongBench, as a low-cost benchmark, holds great potential to substantially drive future research into the LCU capabilities of LLMs.See Github for our code, data and tutorial.
Zhongzhan Huang, Guoming Ling, Shanshan Zhong, Hefeng Wu, Liang Lin 0004
ACL (1)1
2025 AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
abstract
Recent advancements in multimodal large language models (MLLMs) have garnered significant attention, offering a promising pathway toward artificial general intelligence (AGI).Among the essential capabilities required for AGI, creativity has emerged as a critical trait for MLLMs, with association serving as its foundation.Association reflects a model's ability to think creatively, making it vital to evaluate and understand.While several frameworks have been proposed to assess associative ability, they often overlook the inherent ambiguity in association tasks, which arises from the divergent nature of associations and undermines the reliability of evaluations.To address this issue, we decompose ambiguity into two types-internal ambiguity and external ambiguity-and introduce AssoCiAm, a benchmark designed to evaluate associative ability while circumventing the ambiguity through a hybrid computational method.We then conduct extensive experiments on MLLMs, revealing a strong positive correlation between cognition and association.Additionally, we observe that the presence of ambiguity in the evaluation process causes MLLMs' behavior to become more random-like.Finally, we validate the effectiveness of our method in ensuring more accurate and reliable evaluations.See Project Page for the data and codes.
Wenkuan Zhao, Shanshan Zhong, Jinghui Qin, Mingfu Liang, Zhongzhan Huang, Wushao Wen
EMNLP6
2025 Anima2: Cross-Species Animal Animation through Image-to-Video Synthesis with Subject Alignment
abstract
Recent video editing advancements rely on accurate pose sequences to animate human actors. However, these efforts are not suitable for cross-species animation due to pose misalignment between species (for example, the poses of a cat differ greatly from that of a pig due to their distinct body structures). In this paper, we present Anima2, a zero-shot diffusion-based video generator to address this issue, aiming to accurately ANIMAte ANIMAls while preserving the background. The key technique involves two-fold subject alignment. First, we improve appearance feature extraction by integrating a Laplacian detail booster and a prompt-tuning identity extractor. They capture essential appearance information, including identity and fine details. Second, we align shape features and address conflicts from differing animals by introducing a scale-information remover and an adaptive rescaling module. They both enhance subject alignment for accurate cross-species animation. Additionally, we introduce two high-quality animal video datasets with diverse species to benchmark cross-species animation. Trained on these extensive datasets, our model directly generates videos with accurate movements, consistent appearances, and high-fidelity frames, eliminating the need for test-time training. Extensive experiments demonstrate our method’s superiority in cross-species animation, showcasing robust adaptability and generality.
Yuanfeng Xu, Zhongzhan Huang, Guangrun Wang, Liang Lin 0004
ICASSP3
2025 Flat Local Minima for Continual Learning on Semantic Segmentation
Zhongzhan Huang, Mingfu Liang, Senwei Liang, Shanshan Zhong
MMM (1)1
2025 DVIB: Towards Robust Multimodal Recommender Systems via Variational Information Bottleneck Distillation
abstract
In multimodal recommender systems (MRS), integrating various modalities helps to model user preferences and item characteristics more accurately, thereby assisting users in discovering items that match their interests. Although the introduction of multimodal information offers opportunities for performance improvement, it will increase the risks of inherent noise and information redundancy, posing challenges to the robustness of MRS. Many existing methods typically address these two issues separately either by introducing perturbations at the model input for robust training to handle noise or by designing complex network structures to filter out redundant information. In contrast, we propose the DVIB framework to simultaneously address both issues in a simple manner. We found that moving the perturbations from the input layer to the hidden layer, combined with feature self-distillation, can mitigate noise and handle information redundancy without altering the original network architecture. Additionally, we also provide theoretical evidence for the effectiveness of DVIB, demonstrating that the framework not only explicitly enhances the robustness of model training but also implicitly exhibits an information bottleneck effect, which effectively reduces redundant information during multimodal fusion and improves feature extraction quality. Extensive experiments show that DVIB consistently improves the performance of MRS across different datasets and model settings, and it can complement existing robust training methods, representing a promising new paradigm in MRS. See code at https://github.com/MarshmallowLight/DVIB.git.
Wenkuan Zhao, Shanshan Zhong, Wushao Wen, Jinghui Qin, Mingfu Liang, Zhongzhan Huang
WWW7
2025 A generic shared attention mechanism for various backbone neural networks
Zhongzhan Huang, Senwei Liang, Mingfu Liang
Neurocomputing1
2025 A Causality-Aware Paradigm for Evaluating Creativity of Multimodal Large Language Models
abstract
Recently, numerous benchmarks have been developed to evaluate the logical reasoning abilities of large language models (LLMs). However, assessing the equally important creative capabilities of LLMs is challenging due to the subjective, diverse, and data-scarce nature of creativity, especially in multimodal scenarios. In this paper, we consider the comprehensive pipeline for evaluating the creativity of multimodal LLMs, with a focus on suitable evaluation platforms and methodologies. First, we find the Oogiri game-a creativity-driven task requiring humor, associative thinking, and the ability to produce unexpected responses to text, images, or both. This game aligns well with the input-output structure of modern multimodal LLMs and benefits from a rich repository of high-quality, human-annotated creative responses, making it an ideal platform for studying LLM creativity. Next, beyond using the Oogiri game for standard evaluations like ranking and selection, we propose LoTbench, an interactive, causality-aware evaluation framework, to further address some intrinsic risks in standard evaluations, such as information leakage and limited interpretability. The proposed LoTbench not only quantifies LLM creativity more effectively but also visualizes the underlying creative thought processes. Our results show that while most LLMs exhibit constrained creativity, the performance gap between LLMs and humans is not insurmountable. Furthermore, we observe a strong correlation between results from the multimodal cognition benchmark MMMU and LoTbench, but only a weak connection with traditional creativity metrics. This suggests that LoTbench better aligns with human cognitive theories, highlighting cognition as a critical foundation in the early stages of creativity and enabling the bridging of diverse concepts. Project Page.
Zhongzhan Huang, Shanshan Zhong, Pan Zhou 0002, Shanghua Gao, Marinka Zitnik, Liang Lin 0004
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Continuous Value Assignment: A Doubly Robust Data Augmentation for Off-Policy Learning
abstract
Deep reinforcement learning (RL) has witnessed remarkable success in a wide range of control tasks. To overcome RL's notorious sample inefficiency, prior studies have explored data augmentation techniques leveraging collected transition data. However, these methods face challenges in synthesizing transitions adhering to the authentic environment dynamics, especially when the transition is high-dimensional and includes many redundant/irrelevant features to the task. In this article, we introduce continuous value assignment (CVA), an innovative optimization-level data augmentation approach that directly synthesizes novel training data in the state-action value space, effectively bypassing the need for explicit transition modeling. The key intuition of our method is that the transition plays an intermediate role in calculating the state-action value during optimization, and therefore directly augmenting the state-action value is more causally related to the optimization process. Specifically, our CVA combines parameterized value prediction and nonparametric value interpolation from neighboring states, resulting in doubly robust target values w.r.t. novel states and actions. Extensive experiments demonstrate CVA's substantial improvements in sample efficiency across complex continuous control tasks, surpassing several advanced baselines.
Junfan Lin, Zhongzhan Huang, Keze Wang, Lingbo Liu, Liang Lin 0004
IEEE Trans. Neural Networks Learn. Syst.2
2024 Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
abstract
Chain-of-Thought (CoT) [2], [3] guides large language models (LLMs) to reason step-by-step, and can motivate their logical reasoning ability. While effective for logi-cal tasks, CoT is not conducive to creative problem-solving which often requires out-of-box thoughts and is crucial for innovation advancements. In this paper, we explore the Leap-of-Thought (LoT) abilities within LLMs - a non-sequential, creative paradigm involving strong associations and knowledge leaps. To this end, we study LLMs on the popular Oogiri game which needs participants to have good creativity and strong associative thinking for responding unexpectedly and humorously to the given image, text, or both, and thus is suitable for LoT study. Then to investi-gate LLMs' LoT ability in the Oogiri game, we first build a multimodal and multilingual Oogiri-GO dataset which contains over 130,000 samples from the Oogiri game, and observe the insufficient LoT ability or failures of most existing LLMs on the Oogiri game. Accordingly, we introduce a creative Leap-of-Thought (CLoT) paradigm to improve LLM's LoT ability. CLoT first formulates the Oogiri-GO dataset into LoT-oriented instruction tuning data to train pre-trained LLM for achieving certain LoT humor generation and discrimination abilities. Then CLoT designs an ex-plorative self-refinement that encourages the LLM to gener-ate more creative LoT data via exploring parallels between seemingly unrelated concepts and selects high-quality data to train itself for self-refinement. CLoT not only excels in humor generation in the Oogiri game as shown in Fig. 1 but also boosts creative abilities in various tasks like “cloud guessing game” and “divergent association task”. These findings advance our understanding and offer a pathway to improve LLMs' creative capacities for innovative applications across domains. The dataset, code, and models have been released online: https://zhongshsh.github.io/CLoT.
Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin 0004, Marinka Zitnik, Pan Zhou 0002
CVPR2
2024 Stripe Observation Guided Inference Cost-Free Attention Mechanism
Zhongzhan Huang, Shanshan Zhong, Wushao Wen, Jinghui Qin, Liang Lin 0004
ECCV (24)1
2024 DEEPAM: Toward Deeper Attention Module in Residual Convolutional Neural Networks
Shanshan Zhong, Wushao Wen, Jinghui Qin, Zhongzhan Huang
ICANN (1)4
2024 Lottery Ticket Hypothesis for Attention Mechanism in Residual Convolutional Neural Network*
abstract
Recently many plug-and-play self-attention modules (SAMs) are proposed to enhance the model performance by exploiting the internal information of deep convolutional neural networks. However, most SAMs connect individually with each block of the backbone for granted, leading to incremental computational cost and the number of parameters with the growth of network depth. To address this issue, we first empirically and theoretically explore the Lottery Ticket Hypothesis for Self-attention Networks (LTH4SA): a full self-attention network contains a subnetwork with sparse self-attention connections that can (1) accelerate inference, (2) reduce extra parameter increment, and (3) maintain accuracy. Furthermore, we propose a simple yet effective policy-gradient-based baseline method to search the ticket, i.e., the connection scheme that satisfies the three above-mentioned conditions. Extensive experiments on widely-used benchmarks and popular self-attention networks show the effectiveness of our method. Besides, our experiments illustrate that our searched ticket has the capacity of transferring to other vision tasks, e.g., crowd counting and segmentation. https://github.com/gbup-group/EAN-efficient-attention-network.
Zhongzhan Huang, Senwei Liang, Mingfu Liang, Wei He 0025, Haizhao Yang
ICME1
2024 AttNS: Attention-Inspired Numerical Solving For Limited Data Scenarios
abstract
We propose the attention-inspired numerical solver (AttNS), a concise method that helps the generalization and robustness issues faced by the AI-Hybrid numerical solver in solving differential equations due to limited data. AttNS is inspired by the effectiveness of attention modules in Residual Neural Networks (ResNet) in enhancing model generalization and robustness for conventional deep learning tasks. Drawing from the dynamical system perspective of ResNet, We seamlessly incorporate attention mechanisms into the design of numerical methods tailored for the characteristics of solving differential equations. Our results on benchmarks, ranging from high-dimensional problems to chaotic systems, showcase AttNS consistently enhancing various numerical solvers without any intricate model crafting. Finally, we analyze AttNS experimentally and theoretically, demonstrating its ability to achieve strong generalization and robustness while ensuring the convergence of the solver. This includes requiring less data compared to other advanced methods to achieve comparable generalization errors and better prevention of numerical explosion issues when solving differential equations.
Zhongzhan Huang, Mingfu Liang, Shanshan Zhong, Liang Lin 0004
ICML1
2024 Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local Minima
abstract
Multimodal recommender systems utilize various types of information to model user preferences and item features, helping users discover items aligned with their interests. The integration of multimodal information mitigates the inherent challenges in recommender systems, e.g., the data sparsity problem and cold-start issues. However, it simultaneously magnifies certain risks from multimodal information inputs, such as information adjustment risk and inherent noise risk. These risks pose crucial challenges to the robustness of recommendation models. In this paper, we analyze multimodal recommender systems from the novel perspective of flat local minima and propose a concise yet effective gradient strategy called Mirror Gradient (MG). This strategy can implicitly enhance the model's robustness during the optimization process, mitigating instability risks arising from multimodal information inputs. We also provide strong theoretical evidence and conduct extensive empirical experiments to show the superiority of MG across various multimodal recommendation models and benchmarks. Furthermore, we find that the proposed MG can complement existing robust training methods and be easily extended to diverse advanced recommendation models, making it a promising new and fundamental paradigm for training multimodal recommender systems. The code is released at https://github.com/Qrange-group/Mirror-Gradient.
Shanshan Zhong, Zhongzhan Huang, Daifeng Li, Wushao Wen, Jinghui Qin, Liang Lin 0004
WWW2
2024 An Introspective Data Augmentation Method for Training Math Word Problem Solvers
abstract
Though quite challenging, training a deep neural network for automatically solving Math Word Problems (MWPs) has increasingly attracted attention due to its significance in investigating how a machine can understand and reason complex problems like a human. However, the data volume of existing high-quality MWP datasets is far from sufficient to train a robust solver since collecting these datasets would cost a very high price, i.e., they require professional knowledge that accords with the educational standard and massive accessible data. This data bottleneck inspires us to consider using cost-effective data augmentation methods to improve the utilization of the existing data and enhance the performance of an MWP solver. Nevertheless, the traditional input-based data augmentation methods for training natural image or language models are incompetent for training MWP solvers due to the following two reasons. First, MWPs are concise yet comprehensive, so these data augmentation methods are prone to make them more ambiguous. Second, the mathematical dependencies grounded in the problems must be maintained when a batch of augmented examples is generated during the data augmentation process. To address these issues, we propose a simple yet effective data augmentation method called the Introspective Data Augmentation Method (IDAM) that allows the MWP examples to be augmented in latent space during the training of the neural network, instead of making perturbations over the input data. In particular, our IDAM is capable of applying different data augmentation operations on the latent feature representations of MWPs to produce new examples. Moreover, a new training objective is developed to constrain the mathematical dependency consistency between the original MWP and the produced ones. Extensive experiments conducted on standard benchmarks demonstrate the effectiveness of IDAM in generally improving the performance of existing MWP solvers without any elaborated model crafting.
Jinghui Qin, Zhongzhan Huang, Quanshi Zhang, Liang Lin 0004
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Understanding Self-attention Mechanism via Dynamical System Perspective
abstract
The self-attention mechanism (SAM) is widely used in various fields of artificial intelligence and has successfully boosted the performance of different models. However, current explanations of this mechanism are mainly based on intuitions and experiences, while there still lacks direct modeling for how the SAM helps performance. To mitigate this issue, in this paper, based on the dynamical system perspective of the residual neural network, we first show that the intrinsic stiffness phenomenon (SP) in the high-precision solution of ordinary differential equations (ODEs) also widely exists in high-performance neural networks (NN). Thus the ability of NN to measure SP at the feature level is necessary to obtain high performance and is an important factor in the difficulty of training NN. Similar to the adaptive step-size method which is effective in solving stiff ODEs, we show that the SAM is also a stiffness-aware step size adaptor that can enhance the model's representational ability to measure intrinsic SP by refining the estimation of stiffness information and generating adaptive attention values, which provides a new understanding about why and how the SAM can benefit the model performance. This novel perspective can also explain the lottery ticket hypothesis in SAM, design new quantitative metrics of representational ability, and inspire a new theoretic-inspired approach, StepNet. Extensive experiments on several popular benchmarks demonstrate that StepNet can extract fine-grained stiffness information and measure SP accurately, leading to significant improvements in various visual tasks.
Zhongzhan Huang, Mingfu Liang, Jinghui Qin, Shanshan Zhong, Liang Lin 0004
ICCV1
2023 LSAS: Lightweight Sub-attention Strategy for Alleviating Attention Bias Problem
abstract
In computer vision, the performance of deep neural networks (DNNs) is highly related to the feature extraction ability, i.e., the ability to recognize and focus on key pixel regions in an image. However, in this paper, we quantitatively and statistically illustrate that DNNs have a serious attention bias problem on many samples from some popular datasets: (1) Position bias: DNNs fully focus on label-independent regions; (2) Range bias: The focused regions from DNN are not completely contained in the ideal region. Moreover, we find that the existing self-attention modules can alleviate these biases to a certain extent, but the biases are still non-negligible. To further mitigate them, we propose a lightweight sub-attention strategy (LSAS), which utilizes high-order sub-attention modules to improve the original self-attention modules. The effectiveness of LSAS is demonstrated by extensive experiments on widely-used benchmark datasets and popular attention networks. We release our code to help other researchers to reproduce the results of LSAS1
Shanshan Zhong, Wushao Wen, Jinghui Qin, Qiangpu Chen, Zhongzhan Huang
ICME5
2023 SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language Models
abstract
Diffusion models, which have emerged to become popular text-to-image generation models, can produce high-quality and content-rich images guided by textual prompts. However, there are limitations to semantic understanding and commonsense reasoning in existing models when the input prompts are concise narrative, resulting in low-quality image generation. To improve the capacities for narrative prompts, we propose a simple-yet-effective parameter-efficient fine-tuning approach called the Semantic Understanding and Reasoning adapter (SUR-adapter) for pre-trained diffusion models. To reach this goal, we first collect and annotate a new dataset SURD which consists of more than 57,000 semantically corrected multi-modal samples. Each sample contains a simple narrative prompt, a complex keyword-based prompt, and a high-quality image. Then, we align the semantic representation of narrative prompts to the complex prompts and transfer knowledge of large language models (LLMs) to our SUR-adapter via knowledge distillation so that it can acquire the powerful semantic understanding and reasoning capabilities to build a high-quality textual semantic representation for text-to-image generation. We conduct experiments by integrating multiple LLMs and popular pre-trained diffusion models to show the effectiveness of our approach in enabling diffusion models to understand and reason concise natural language without image quality degradation. Our approach can make text-to-image diffusion models easier to use with better user experience, which demonstrates our approach has the potential for further advancing the development of user-friendly text-to-image generation models by bridging the semantic gap between simple narrative prompts and complex keyword-based prompts. The code is released at https://github.com/Qrange-group/SUR-adapter.
Shanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin, Liang Lin 0004
ACM Multimedia2
2023 ScaleLong: Towards More Stable Training of Diffusion Model via Scaling Network Long Skip Connection
abstract
In diffusion models, UNet is the most popular network backbone, since its long skip connects (LSCs) to connect distant network blocks can aggregate long-distant information and alleviate vanishing gradient. Unfortunately, UNet often suffers from unstable training in diffusion models which can be alleviated by scaling its LSC coefficients smaller. However, theoretical understandings of the instability of UNet in diffusion models and also the performance improvement of LSC scaling remain absent yet. To solve this issue, we theoretically show that the coefficients of LSCs in UNet have big effects on the stableness of the forward and backward propagation and robustness of UNet. Specifically, the hidden feature and gradient of UNet at any layer can oscillate and their oscillation ranges are actually large which explains the instability of UNet training. Moreover, UNet is also provably sensitive to perturbed input, and predicts an output distant from the desired output, yielding oscillatory loss and thus oscillatory gradient. Besides, we also observe the theoretical benefits of the LSC coefficient scaling of UNet in the stableness of hidden features and gradient and also robustness. Finally, inspired by our theory, we propose an effective coefficient scaling framework ScaleLong that scales the coefficients of LSC in UNet and better improve the training stability of UNet. Experimental results on CIFAR10, CelebA, ImageNet and COCO show that our methods are superior to stabilize training, and yield about 1.5x training acceleration on different diffusion models with UNet or UViT backbones.
Zhongzhan Huang, Pan Zhou 0002, Shuicheng Yan, Liang Lin 0004
NeurIPS1
2023 ESA: Excitation-Switchable Attention for convolutional neural networks
Shanshan Zhong, Zhongzhan Huang, Wushao Wen, Zhijing Yang, Jinghui Qin
Neurocomputing2
2022 CEM: Machine-Human Chatting Handoff via Causal-Enhance Module
abstract
Aiming to ensure chatbot quality by predicting chatbot failure and enabling human-agent collaboration, Machine-Human Chatting Handoff (MHCH) has attracted lots of attention from both industry and academia in recent years.However, most existing methods mainly focus on the dialogue context or assist with global satisfaction prediction based on multi-task learning, which ignore the grounded relationships among the causal variables, like the user state and labor cost.These variables are significantly associated with handoff decisions, resulting in prediction bias and cost increase.Therefore, we propose Causal-Enhance Module (CEM) by establishing the causal graph of MHCH based on these two variables, which is a simple yet effective module and can be easy to plug into the existing MHCH methods.For the impact of users, we use the user state to correct the prediction bias according to the causal relationship of multi-task.For the labor cost, we train an auxiliary cost simulator to calculate unbiased labor cost through counterfactual learning so that a model becomes cost-aware.Extensive experiments conducted on four real-world benchmarks demonstrate the effectiveness of CEM in generally improving the performance of existing MHCH methods without any elaborated model crafting.
Shanshan Zhong, Jinghui Qin, Zhongzhan Huang, Daifeng Li
EMNLP3
2022 Stiffness-aware neural network for learning Hamiltonian systems
Senwei Liang, Zhongzhan Huang
ICLR2
2021 Blending Pruning Criteria for Convolutional Neural Networks
Wei He 0025, Zhongzhan Huang, Mingfu Liang, Senwei Liang, Haizhao Yang
ICANN (4)2
2021 Continuous Transition: Improving Sample Efficiency for Continuous Control Problems via MixUp
abstract
Although deep reinforcement learning (RL) has been successfully applied to a variety of robotic control tasks, it’s still challenging to apply it to real-world tasks, due to the poor sample efficiency. Attempting to overcome this shortcoming, several works focus on reusing the collected trajectory data during the training by decomposing them into a set of policy-irrelevant discrete transitions. However, their improvements are somewhat marginal since i) the amount of the transitions is usually small, and ii) the value assignment only happens in the joint states. To address these issues, this paper introduces a concise yet powerful method to construct Continuous Transition, which exploits the trajectory information by exploiting the potential transitions along the trajectory. Specifically, we propose to synthesize new transitions for training by linearly interpolating the consecutive transitions. To keep the constructed transitions authentic, we also develop a discriminator to guide the construction process automatically. Extensive experiments demonstrate that our proposed method achieves a significant improvement in sample efficiency on various complex continuous robotic control problems in MuJoCo and outperforms the advanced model-based / model-free RL methods. The source code is available1.
Junfan Lin, Zhongzhan Huang, Keze Wang, Xiaodan Liang, Liang Lin 0004
ICRA2
2021 Rethinking the Pruning Criteria for Convolutional Neural Network
abstract
Channel pruning is a popular technique for compressing convolutional neural networks (CNNs), where various pruning criteria have been proposed to remove the redundant filters. From our comprehensive experiments, we found two blind spots of pruning criteria: (1) Similarity: There are some strong similarities among several primary pruning criteria that are widely cited and compared. According to these criteria, the ranks of filters’ Importance Score are almost identical, resulting in similar pruned structures. (2) Applicability: The filters' Importance Score measured by some pruning criteria are too close to distinguish the network redundancy well. In this paper, we analyze the above blind spots on different types of pruning criteria with layer-wise pruning or global pruning. We also break some stereotypes, such as that the results of $\ell_1$ and $\ell_2$ pruning are not always similar. These analyses are based on the empirical experiments and our assumption (Convolutional Weight Distribution Assumption) that the well-trained convolutional filters in each layer approximately follow a Gaussian-alike distribution. This assumption has been verified through systematic and extensive statistical tests.
Zhongzhan Huang, Wenqi Shao, Xinjiang Wang, Liang Lin 0004, Ping Luo 0002
NeurIPS1
2020 DIANet: Dense-and-Implicit Attention Network
abstract
Attention networks have successfully boosted the performance in various vision problems. Previous works lay emphasis on designing a new attention module and individually plug them into the networks. Our paper proposes a novel-and-simple framework that shares an attention module throughout different network layers to encourage the integration of layer-wise information and this parameter-sharing module is referred to as Dense-and-Implicit-Attention (DIA) unit. Many choices of modules can be used in the DIA unit. Since Long Short Term Memory (LSTM) has a capacity of capturing long-distance dependency, we focus on the case when the DIA unit is the modified LSTM (called DIA-LSTM). Experiments on benchmark datasets show that the DIA-LSTM unit is capable of emphasizing layer-wise feature interrelation and leads to significant improvement of image classification accuracy. We further empirically show that the DIA-LSTM has a strong regularization ability on stabilizing the training of deep networks by the experiments with the removal of skip connections (He et al. 2016a) or Batch Normalization (Ioffe and Szegedy 2015) in the whole residual network.
Zhongzhan Huang, Senwei Liang, Mingfu Liang, Haizhao Yang
AAAI1
2020 Instance Enhancement Batch Normalization: An Adaptive Regulator of Batch Noise
abstract
Batch Normalization (BN) (Ioffe and Szegedy 2015) normalizes the features of an input image via statistics of a batch of images and hence BN will bring the noise to the gradient of training loss. Previous works indicate that the noise is important for the optimization and generalization of deep neural networks, but too much noise will harm the performance of networks. In our paper, we offer a new point of view that the self-attention mechanism can help to regulate the noise by enhancing instance-specific information to obtain a better regularization effect. Therefore, we propose an attention-based BN called Instance Enhancement Batch Normalization (IEBN) that recalibrates the information of each channel by a simple linear transformation. IEBN has a good capacity of regulating the batch noise and stabilizing network training to improve generalization even in the presence of two kinds of noise attacks during training. Finally, IEBN outperforms BN with only a light parameter increment in image classification tasks under different network structures and benchmark datasets.
Senwei Liang, Zhongzhan Huang, Mingfu Liang, Haizhao Yang
AAAI2