Shanshan Zhong

dblp:297/7816 · also ShanShan Zhong · DBLP profile ↗
← Back
19ranked-venue papers
10as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
abstract
Reconstructing animatable 3D animals from videos traditionally depends on sparse semantic keypoints to fit parametric models. Acquiring these keypoints is labor-intensive, and detectors trained on limited animal datasets are often unreliable. We propose 4D-Animal, a keypoint-free framework that reconstructs animatable 3D animals directly from videos. Our method employs a dense feature network to map 2D image representations to SMAL parameters, improving both efficiency and stability. Additionally, we introduce a hierarchical alignment strategy that leverages silhouette, part-level, pixel-level, and temporal cues from pretrained 2D models, ensuring accurate and temporally coherent reconstructions. Extensive experiments demonstrate that 4D-Animal outperforms both model-based and model-free baselines on dog dataset. Moreover, the high-quality 3D assets generated by our method can benefit other 3D tasks, underscoring its potential for large-scale applications. The code is released at https://github.com/zhongshsh/4D-Animal.
Shanshan Zhong, Zehan Zheng, Zhongzhan Huang, Wufei Ma, Guofeng Zhang 0025, Qihao Liu, Alan L. Yuille, Jieneng Chen
WACV1
2025 MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
abstract
Long Context Understanding (LCU) is a critical area for exploration in current large language models (LLMs).However, due to the inherently lengthy nature of long-text data, existing LCU benchmarks for LLMs often result in prohibitively high evaluation costs, like testing time and inference expenses.Through extensive experimentation, we discover that existing LCU benchmarks exhibit significant redundancy, which means the inefficiency in evaluation.In this paper, we propose a concise data compression method tailored for longtext data with sparse information characteristics.By pruning the well-known LCU benchmark LongBench, we create MiniLongBench.This benchmark includes only 237 test samples across six major task categories and 21 distinct tasks.Through empirical analysis of over 60 LLMs, MiniLongBench achieves an average evaluation cost reduced to only 4.5% of the original while maintaining an average rank correlation coefficient of 0.97 with Long-Bench results.Therefore, our MiniLongBench, as a low-cost benchmark, holds great potential to substantially drive future research into the LCU capabilities of LLMs.See Github for our code, data and tutorial.
Zhongzhan Huang, Guoming Ling, Shanshan Zhong, Hefeng Wu, Liang Lin 0004
ACL (1)3
2025 AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
abstract
Recent advancements in multimodal large language models (MLLMs) have garnered significant attention, offering a promising pathway toward artificial general intelligence (AGI).Among the essential capabilities required for AGI, creativity has emerged as a critical trait for MLLMs, with association serving as its foundation.Association reflects a model's ability to think creatively, making it vital to evaluate and understand.While several frameworks have been proposed to assess associative ability, they often overlook the inherent ambiguity in association tasks, which arises from the divergent nature of associations and undermines the reliability of evaluations.To address this issue, we decompose ambiguity into two types-internal ambiguity and external ambiguity-and introduce AssoCiAm, a benchmark designed to evaluate associative ability while circumventing the ambiguity through a hybrid computational method.We then conduct extensive experiments on MLLMs, revealing a strong positive correlation between cognition and association.Additionally, we observe that the presence of ambiguity in the evaluation process causes MLLMs' behavior to become more random-like.Finally, we validate the effectiveness of our method in ensuring more accurate and reliable evaluations.See Project Page for the data and codes.
Wenkuan Zhao, Shanshan Zhong, Jinghui Qin, Mingfu Liang, Zhongzhan Huang, Wushao Wen
EMNLP3
2025 KSINet: Keypoint-Enhanced Single-Head Inference for Sitting Posture Recognition
abstract
Poor posture, a common phenomenon in modern lifestyles, has led to an increased incidence of health issues such as cervical spondylosis and lumbar disc herniation. These conditions not only severely affect the quality of people’s daily lives but also result in a significant rise in healthcare costs. Consequently, an accurate posture detection model is crucial. To address the limitations of existing posture classification methods, such as slow inference speeds and low accuracy in complex environments, we propose the Keypoint-Enhanced Single-Head Inference Network (KSINet). During training, both the classification head and heatmap head are utilized, with their losses integrated to optimize the model’s performance. The heatmap head guides the model to focus on key regions within the image, enhancing classification accuracy. For inference, only the classification head is activated, reducing computational costs and significantly accelerating inference speed. To further improve KSINet’s ability to capture long-range dependencies, we introduce the SmoothSharp Attention (SSA) mechanism and Gated Dilated Residual (GDR) module. Experiments demonstrate that our model achieves 98.64% accuracy with a 7ms inference delay in our constructed high-quality dataset of 8 sitting categories, outperforming state-of-the-art models.
Yongheng Tai, Shoudong Shi, Shanshan Zhong, Ting Lan 0001, Tianxiang Zhao 0005, Kedi Qiu
IJCNN3
2025 Flat Local Minima for Continual Learning on Semantic Segmentation
Zhongzhan Huang, Mingfu Liang, Senwei Liang, Shanshan Zhong
MMM (1)4
2025 DPHNet: A Dynamic Parallel Hybrid Network for Sitting Posture Recognition
abstract
Poor sitting posture can lead to a variety of diseases [1]. Therefore, using visual technology to recognize and correct poor sitting posture is of great significance in promoting a healthier lifestyle. In recent years, hybrid architectures that combine the advantages of Convolutional Neural Networks and Vision Transformers have achieved success in a variety of vision tasks. However, challenges remain in effectively coordinating information fusion between the two models and controlling computational cost without compromising performance. To address these issues, we propose a dynamic parallel hybrid model, DPHNet. DPHNet introduces a Dual-Path Control (DPC) module that dynamically enables optional attention branche based on feature information and adjusts the weights between attention and fixed convolutional branches. In order to strike a balance between computational cost and feature representation accuracy, DPHNet introduce the Patch Dynamic Selection (PDS) module, which dynamically selects the patch size based on feature information. With this design, DPHNet demonstrates excellent performance in sitting posture recognition: on our constructed dataset containing 31,020 images covering 8 sitting postures, DPHNet attains a recognition accuracy of 96.1% while maintaining a minimal computational cost of just 3.1 GFLOPs, outperforming existing state-of-the-art models.
Shanshan Zhong, Shoudong Shi, Yongheng Tai, Ting Lan 0001, Tianxiang Zhao 0005, Kedi Qiu
SMC1
2025 DVIB: Towards Robust Multimodal Recommender Systems via Variational Information Bottleneck Distillation
abstract
In multimodal recommender systems (MRS), integrating various modalities helps to model user preferences and item characteristics more accurately, thereby assisting users in discovering items that match their interests. Although the introduction of multimodal information offers opportunities for performance improvement, it will increase the risks of inherent noise and information redundancy, posing challenges to the robustness of MRS. Many existing methods typically address these two issues separately either by introducing perturbations at the model input for robust training to handle noise or by designing complex network structures to filter out redundant information. In contrast, we propose the DVIB framework to simultaneously address both issues in a simple manner. We found that moving the perturbations from the input layer to the hidden layer, combined with feature self-distillation, can mitigate noise and handle information redundancy without altering the original network architecture. Additionally, we also provide theoretical evidence for the effectiveness of DVIB, demonstrating that the framework not only explicitly enhances the robustness of model training but also implicitly exhibits an information bottleneck effect, which effectively reduces redundant information during multimodal fusion and improves feature extraction quality. Extensive experiments show that DVIB consistently improves the performance of MRS across different datasets and model settings, and it can complement existing robust training methods, representing a promising new paradigm in MRS. See code at https://github.com/MarshmallowLight/DVIB.git.
Wenkuan Zhao, Shanshan Zhong, Wushao Wen, Jinghui Qin, Mingfu Liang, Zhongzhan Huang
WWW2
2025 A Causality-Aware Paradigm for Evaluating Creativity of Multimodal Large Language Models
abstract
Recently, numerous benchmarks have been developed to evaluate the logical reasoning abilities of large language models (LLMs). However, assessing the equally important creative capabilities of LLMs is challenging due to the subjective, diverse, and data-scarce nature of creativity, especially in multimodal scenarios. In this paper, we consider the comprehensive pipeline for evaluating the creativity of multimodal LLMs, with a focus on suitable evaluation platforms and methodologies. First, we find the Oogiri game-a creativity-driven task requiring humor, associative thinking, and the ability to produce unexpected responses to text, images, or both. This game aligns well with the input-output structure of modern multimodal LLMs and benefits from a rich repository of high-quality, human-annotated creative responses, making it an ideal platform for studying LLM creativity. Next, beyond using the Oogiri game for standard evaluations like ranking and selection, we propose LoTbench, an interactive, causality-aware evaluation framework, to further address some intrinsic risks in standard evaluations, such as information leakage and limited interpretability. The proposed LoTbench not only quantifies LLM creativity more effectively but also visualizes the underlying creative thought processes. Our results show that while most LLMs exhibit constrained creativity, the performance gap between LLMs and humans is not insurmountable. Furthermore, we observe a strong correlation between results from the multimodal cognition benchmark MMMU and LoTbench, but only a weak connection with traditional creativity metrics. This suggests that LoTbench better aligns with human cognitive theories, highlighting cognition as a critical foundation in the early stages of creativity and enabling the bridging of diverse concepts. Project Page.
Zhongzhan Huang, Shanshan Zhong, Pan Zhou 0002, Shanghua Gao, Marinka Zitnik, Liang Lin 0004
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation
abstract
Chain-of-Thought (CoT) [2], [3] guides large language models (LLMs) to reason step-by-step, and can motivate their logical reasoning ability. While effective for logi-cal tasks, CoT is not conducive to creative problem-solving which often requires out-of-box thoughts and is crucial for innovation advancements. In this paper, we explore the Leap-of-Thought (LoT) abilities within LLMs - a non-sequential, creative paradigm involving strong associations and knowledge leaps. To this end, we study LLMs on the popular Oogiri game which needs participants to have good creativity and strong associative thinking for responding unexpectedly and humorously to the given image, text, or both, and thus is suitable for LoT study. Then to investi-gate LLMs' LoT ability in the Oogiri game, we first build a multimodal and multilingual Oogiri-GO dataset which contains over 130,000 samples from the Oogiri game, and observe the insufficient LoT ability or failures of most existing LLMs on the Oogiri game. Accordingly, we introduce a creative Leap-of-Thought (CLoT) paradigm to improve LLM's LoT ability. CLoT first formulates the Oogiri-GO dataset into LoT-oriented instruction tuning data to train pre-trained LLM for achieving certain LoT humor generation and discrimination abilities. Then CLoT designs an ex-plorative self-refinement that encourages the LLM to gener-ate more creative LoT data via exploring parallels between seemingly unrelated concepts and selects high-quality data to train itself for self-refinement. CLoT not only excels in humor generation in the Oogiri game as shown in Fig. 1 but also boosts creative abilities in various tasks like “cloud guessing game” and “divergent association task”. These findings advance our understanding and offer a pathway to improve LLMs' creative capacities for innovative applications across domains. The dataset, code, and models have been released online: https://zhongshsh.github.io/CLoT.
Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin 0004, Marinka Zitnik, Pan Zhou 0002
CVPR1
2024 Stripe Observation Guided Inference Cost-Free Attention Mechanism
Zhongzhan Huang, Shanshan Zhong, Wushao Wen, Jinghui Qin, Liang Lin 0004
ECCV (24)2
2024 DEEPAM: Toward Deeper Attention Module in Residual Convolutional Neural Networks
Shanshan Zhong, Wushao Wen, Jinghui Qin, Zhongzhan Huang
ICANN (1)1
2024 AttNS: Attention-Inspired Numerical Solving For Limited Data Scenarios
abstract
We propose the attention-inspired numerical solver (AttNS), a concise method that helps the generalization and robustness issues faced by the AI-Hybrid numerical solver in solving differential equations due to limited data. AttNS is inspired by the effectiveness of attention modules in Residual Neural Networks (ResNet) in enhancing model generalization and robustness for conventional deep learning tasks. Drawing from the dynamical system perspective of ResNet, We seamlessly incorporate attention mechanisms into the design of numerical methods tailored for the characteristics of solving differential equations. Our results on benchmarks, ranging from high-dimensional problems to chaotic systems, showcase AttNS consistently enhancing various numerical solvers without any intricate model crafting. Finally, we analyze AttNS experimentally and theoretically, demonstrating its ability to achieve strong generalization and robustness while ensuring the convergence of the solver. This includes requiring less data compared to other advanced methods to achieve comparable generalization errors and better prevention of numerical explosion issues when solving differential equations.
Zhongzhan Huang, Mingfu Liang, Shanshan Zhong, Liang Lin 0004
ICML3
2024 Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local Minima
abstract
Multimodal recommender systems utilize various types of information to model user preferences and item features, helping users discover items aligned with their interests. The integration of multimodal information mitigates the inherent challenges in recommender systems, e.g., the data sparsity problem and cold-start issues. However, it simultaneously magnifies certain risks from multimodal information inputs, such as information adjustment risk and inherent noise risk. These risks pose crucial challenges to the robustness of recommendation models. In this paper, we analyze multimodal recommender systems from the novel perspective of flat local minima and propose a concise yet effective gradient strategy called Mirror Gradient (MG). This strategy can implicitly enhance the model's robustness during the optimization process, mitigating instability risks arising from multimodal information inputs. We also provide strong theoretical evidence and conduct extensive empirical experiments to show the superiority of MG across various multimodal recommendation models and benchmarks. Furthermore, we find that the proposed MG can complement existing robust training methods and be easily extended to diverse advanced recommendation models, making it a promising new and fundamental paradigm for training multimodal recommender systems. The code is released at https://github.com/Qrange-group/Mirror-Gradient.
Shanshan Zhong, Zhongzhan Huang, Daifeng Li, Wushao Wen, Jinghui Qin, Liang Lin 0004
WWW1
2023 Understanding Self-attention Mechanism via Dynamical System Perspective
abstract
The self-attention mechanism (SAM) is widely used in various fields of artificial intelligence and has successfully boosted the performance of different models. However, current explanations of this mechanism are mainly based on intuitions and experiences, while there still lacks direct modeling for how the SAM helps performance. To mitigate this issue, in this paper, based on the dynamical system perspective of the residual neural network, we first show that the intrinsic stiffness phenomenon (SP) in the high-precision solution of ordinary differential equations (ODEs) also widely exists in high-performance neural networks (NN). Thus the ability of NN to measure SP at the feature level is necessary to obtain high performance and is an important factor in the difficulty of training NN. Similar to the adaptive step-size method which is effective in solving stiff ODEs, we show that the SAM is also a stiffness-aware step size adaptor that can enhance the model's representational ability to measure intrinsic SP by refining the estimation of stiffness information and generating adaptive attention values, which provides a new understanding about why and how the SAM can benefit the model performance. This novel perspective can also explain the lottery ticket hypothesis in SAM, design new quantitative metrics of representational ability, and inspire a new theoretic-inspired approach, StepNet. Extensive experiments on several popular benchmarks demonstrate that StepNet can extract fine-grained stiffness information and measure SP accurately, leading to significant improvements in various visual tasks.
Zhongzhan Huang, Mingfu Liang, Jinghui Qin, Shanshan Zhong, Liang Lin 0004
ICCV4
2023 LSAS: Lightweight Sub-attention Strategy for Alleviating Attention Bias Problem
abstract
In computer vision, the performance of deep neural networks (DNNs) is highly related to the feature extraction ability, i.e., the ability to recognize and focus on key pixel regions in an image. However, in this paper, we quantitatively and statistically illustrate that DNNs have a serious attention bias problem on many samples from some popular datasets: (1) Position bias: DNNs fully focus on label-independent regions; (2) Range bias: The focused regions from DNN are not completely contained in the ideal region. Moreover, we find that the existing self-attention modules can alleviate these biases to a certain extent, but the biases are still non-negligible. To further mitigate them, we propose a lightweight sub-attention strategy (LSAS), which utilizes high-order sub-attention modules to improve the original self-attention modules. The effectiveness of LSAS is demonstrated by extensive experiments on widely-used benchmark datasets and popular attention networks. We release our code to help other researchers to reproduce the results of LSAS1
Shanshan Zhong, Wushao Wen, Jinghui Qin, Qiangpu Chen, Zhongzhan Huang
ICME1
2023 SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language Models
abstract
Diffusion models, which have emerged to become popular text-to-image generation models, can produce high-quality and content-rich images guided by textual prompts. However, there are limitations to semantic understanding and commonsense reasoning in existing models when the input prompts are concise narrative, resulting in low-quality image generation. To improve the capacities for narrative prompts, we propose a simple-yet-effective parameter-efficient fine-tuning approach called the Semantic Understanding and Reasoning adapter (SUR-adapter) for pre-trained diffusion models. To reach this goal, we first collect and annotate a new dataset SURD which consists of more than 57,000 semantically corrected multi-modal samples. Each sample contains a simple narrative prompt, a complex keyword-based prompt, and a high-quality image. Then, we align the semantic representation of narrative prompts to the complex prompts and transfer knowledge of large language models (LLMs) to our SUR-adapter via knowledge distillation so that it can acquire the powerful semantic understanding and reasoning capabilities to build a high-quality textual semantic representation for text-to-image generation. We conduct experiments by integrating multiple LLMs and popular pre-trained diffusion models to show the effectiveness of our approach in enabling diffusion models to understand and reason concise natural language without image quality degradation. Our approach can make text-to-image diffusion models easier to use with better user experience, which demonstrates our approach has the potential for further advancing the development of user-friendly text-to-image generation models by bridging the semantic gap between simple narrative prompts and complex keyword-based prompts. The code is released at https://github.com/Qrange-group/SUR-adapter.
Shanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin, Liang Lin 0004
ACM Multimedia1
2023 SPEM: Self-adaptive Pooling Enhanced Attention Module for Image Recognition
Shanshan Zhong, Wushao Wen, Jinghui Qin
MMM (2)1
2023 ESA: Excitation-Switchable Attention for convolutional neural networks
Shanshan Zhong, Zhongzhan Huang, Wushao Wen, Zhijing Yang, Jinghui Qin
Neurocomputing1
2022 CEM: Machine-Human Chatting Handoff via Causal-Enhance Module
abstract
Aiming to ensure chatbot quality by predicting chatbot failure and enabling human-agent collaboration, Machine-Human Chatting Handoff (MHCH) has attracted lots of attention from both industry and academia in recent years.However, most existing methods mainly focus on the dialogue context or assist with global satisfaction prediction based on multi-task learning, which ignore the grounded relationships among the causal variables, like the user state and labor cost.These variables are significantly associated with handoff decisions, resulting in prediction bias and cost increase.Therefore, we propose Causal-Enhance Module (CEM) by establishing the causal graph of MHCH based on these two variables, which is a simple yet effective module and can be easy to plug into the existing MHCH methods.For the impact of users, we use the user state to correct the prediction bias according to the causal relationship of multi-task.For the labor cost, we train an auxiliary cost simulator to calculate unbiased labor cost through counterfactual learning so that a model becomes cost-aware.Extensive experiments conducted on four real-world benchmarks demonstrate the effectiveness of CEM in generally improving the performance of existing MHCH methods without any elaborated model crafting.
Shanshan Zhong, Jinghui Qin, Zhongzhan Huang, Daifeng Li
EMNLP1