Wang Shen

dblp:250/8374 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Beyond Unidimensional Burnout: Exploring Social Media Fatigue Through a Dynamic and Multidimensional Lens
abstract
Despite the widely recognition of social media fatigue (SMF), its multidimensional nature remains underexplored. This study proposes a dynamic model dividing SMF into cognitive, emotional, and motivational dimensions. Utilizing the Stressor-Strain-Outcome (SSO) framework, we investigate how different stressors activate these dimensions and affect discontinuous use. The model was tested using PLS-SEM on data from 356 users. Three main findings emerged. Cognitive fatigue increases emotional fatigue, and together they increase motivational fatigue, indicating a potential reinforcing pattern. Secondly, stressors have varying effects across SMF dimensions. Information and system feature overload mainly induce cognitive and motivational fatigue but not emotional fatigue. By contrast, social overload mainly induces emotional fatigue. Social comparison raises cognitive and emotional fatigue but reduces motivational fatigue, acting as both a stressor and engagement driver. Thirdly, only motivational fatigue directly predicts discontinuous usage intention; cognitive and emotional fatigue act indirectly via it. The study contributes to the ongoing development of social media research.
Qianru Shi, Wang Shen
Int. J. Hum. Comput. Interact.2
2025 AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
abstract
Vision-language models (VLMs) show remarkable performance in multimodal tasks. However, excessively long multimodal inputs lead to oversized Key-Value (KV) caches, resulting in significant memory consumption and I/O bottlenecks. Previous KV quantization methods for Large Language Models (LLMs) may alleviate these issues but overlook the attention saliency differences of multimodal tokens, resulting in suboptimal performance. In this paper, we investigate the attention-aware token saliency patterns in VLM and propose AKVQ-VL. AKVQ-VL leverages the proposed Text-Salient Attention (TSA) and Pivot-Token-Salient Attention (PSA) patterns to adaptively allocate bit budgets. Moreover, achieving extremely low-bit quantization requires effectively addressing outliers in KV tensors. AKVQ-VL utilizes the Walsh-Hadamard transform (WHT) to construct outlier-free KV caches, thereby reducing quantization difficulty. Evaluations of 2-bit quantization on 12 long-context and multimodal tasks demonstrate that AKVQ-VL maintains or even improves accuracy, outperforming LLM-oriented methods. AKVQ-VL can reduce peak memory usage by 2.13×, support up to 3.25× larger batch sizes and 2.46× throughput.
Zunhai Su, Wang Shen, Linge Li, Hanyu Wei, Huangqi Yu, Kehong Yuan
ICME2
2025 RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
abstract
Key-Value (KV) cache facilitates efficient large language models (LLMs) inference by avoiding recomputation of past KVs. As the batch size and context length increase, the oversized KV caches become a significant memory bottleneck, highlighting the need for efficient compression. Existing KV quantization rely on fine-grained quantization or the retention of a significant portion of high bit-widths caches, both of which compromise compression ratio and often fail to maintain robustness at extremely low average bit-widths. In this work, we explore the potential of rotation technique for 2-bit KV quantization and propose RotateKV, which achieves accurate and robust performance through the following innovations: (i) Outlier-Aware Rotation, which utilizes channel-reordering to adapt the rotations to varying channel-wise outlier distributions without sacrificing the computational efficiency of the fast Walsh-Hadamard transform (FWHT); (ii) Pre-RoPE Grouped-Head Rotation, which mitigates the impact of rotary position embedding (RoPE) on proposed outlier-aware rotation and further smooths outliers across heads; (iii) Attention-Sink-Aware Quantization, which leverages the massive activations to precisely identify and protect attention sinks. RotateKV achieves less than 0.3 perplexity (PPL) degradation with 2-bit quantization on WikiText-2 using LLaMA-2-13B, maintains strong CoT reasoning and long-context capabilities, with less than 1.7% degradation on GSM8K, outperforming existing methods even at lower average bit-widths. RotateKV also showcases a 3.97× reduction in peak memory usage, supports 5.75× larger batch sizes, and achieves a 2.32× speedup in decoding stage.
Zunhai Su, Hanyu Wei, Wang Shen, Linge Li, Huangqi Yu, Kehong Yuan
IJCAI4
2025 How to cope with the negative health information avoidance behavior in a pandemic: the role of resilience
abstract
While social media has become an increasingly important channel for updating risk news and getting early warnings in the pandemic, it has also led to an information overload and misinformation which has been shown to trigger negative emotions and impact mental health. Emotional state and cognitive characteristics are crucial factors that contribute to information avoidance. In this paper, we integrate the S-S-O framework with Resilience Theory to explore the factors that influence health information avoidance and coping mechanisms during the pandemic. Our investigation reveals that perceived health information overload causes psychological strain, including health information anxiety and time pressure, which are associated with health information avoidance. Our findings also show that resilience is a significant inhibitor of the S-S-O chain. These insights have significant implications for improving pandemic management and promoting individuals’ resilience during public health emergencies.
Wang Shen, Yuming He
Behav. Inf. Technol.3
2022 Enhanced Deep Animation Video Interpolation
abstract
Existing learning-based frame interpolation algorithms extract consecutive frames from high-speed natural videos to train the model. Compared to natural videos, cartoon videos are usually in a low frame rate. Besides, the motion between consecutive cartoon frames is typically nonlinear, which breaks the linear motion assumption of interpolation algorithms. Thus, it is unsuitable for generating a training set directly from cartoon videos. For better adapting frame interpolation algorithms from nature video to animation video, we present AutoFI, a simple and effective method to automatically render training data for deep animation video interpolation. AutoFI takes a layered architecture to render synthetic data, which ensures the assumption of linear motion. Experimental results show that AutoFI performs favorably in training both DAIN and ANIN. However, most frame interpolation algorithms will still fail in error-prone areas, such as fast motion or large occlusion. Besides AutoFI, we also propose a plug-and-play sketch-based post-processing module, named SktFI, to refine the final results using user-provided sketches manually. With AutoFI and SktFI, the interpolated animation frames show high perceptual quality.
Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021
ICIP1
2022 Spatial Temporal Video Enhancement Using Alternating Exposures
abstract
High-speed video acquisition under poor illumination conditions is a challenging task. Imaging using long exposure can ensure brightness and suppress noise. However, the captured images may be blurry due to fast object movements or camera shakes. Imaging with short exposure can record sharp textures, but the high camera gain may cause noticeable noise. To alleviate this dilemma, we design a camera system using alternating exposures, where frames expose cyclically in a short-long way. The system consists of restoration and interpolation modules to reconstruct sharp, noise-reduced, high-frame-rate frames from low-frame-rate alternate-exposed input images. We design an optical-flow-based alternate-complementary alignment architecture for spatial enhancement, which effectively aligns the short-exposed and long-exposed images in a two-stage progressive way. Moreover, it explores complementary information from short-exposed and long-exposed inputs to ensure consistency between outputs. We propose a flow-enhanced frame interpolation module for temporal enhancement, which refines the intermediate flows and reconstructs the intermediate images based on the restored images of the alignment network and warped input neighboring frames. The whole network with two modules is end-to-end jointly learnable. We first evaluate the algorithm on simulation data. To demonstrate practicality, we then test it on real data by setting up a prototype camera. We propose an effective spatial degradation regularization strategy to reduce the domain gap between simulation and real data. Besides, we extend our method by integrating multi-frame exposure fusion technology to reduce overexposure areas in real scenarios. Experimental results show that our method performs favorably against state-of-the-art methods on both synthetic data and real-world data.
Wang Shen, Guo Lu, Guangtao Zhai, Li Chen 0021, Muhammad Salman Asif
IEEE Trans. Circuits Syst. Video Technol.1
2021 Dual Attention Guided Gaze Target Detection in the Wild
abstract
Gaze target detection aims to infer where each person in a scene is looking. Existing works focus on 2D gaze and 2D saliency, but fail to exploit 3D contexts. In this work, we propose a three-stage method to simulate the human gaze inference behavior in 3D space. In the first stage, we introduce a coarse-to-fine strategy to robustly estimate a 3D gaze orientation from the head. The predicted gaze is decomposed into a planar gaze on the image plane and a depth-channel gaze. In the second stage, we develop a Dual Attention Module (DAM), which takes the planar gaze to produce the filed of view and masks interfering objects regulated by depth information according to the depth-channel gaze. In the third stage, we use the generated dual attention as guidance to perform two sub-tasks: (1) identifying whether the gaze target is inside or out of the image; (2) locating the target if inside. Extensive experiments demonstrate that our approach performs favorably against state-of-the-art methods on GazeFollow and VideoAttentionTarget datasets.
Yi Fang 0009, Jiapeng Tang, Wang Shen, Wei Shen 0002, Xiao Gu 0001, Li Song 0001, Guangtao Zhai
CVPR3
2021 Video Frame Interpolation and Enhancement via Pyramid Recurrent Framework
abstract
Video frame interpolation aims to improve users' watching experiences by generating high-frame-rate videos from low-frame-rate ones. Existing approaches typically focus on synthesizing intermediate frames using high-quality reference images. However, the captured reference frames may suffer from inevitable spatial degradations such as motion blur, sensor noise, etc. Few studies have approached the joint video enhancement problem, namely synthesizing high-frame-rate and high-quality results from low-frame-rate degraded inputs. In this paper, we propose a unified optimization framework for video frame interpolation with spatial degradations. Specifically, we develop a frame interpolation module with a pyramid structure to cyclically synthesize high-quality intermediate frames. The pyramid module features adjustable spatial receptive field and temporal scope, thus contributing to controllable computational complexity and restoration ability. Besides, we propose an inter-pyramid recurrent module to connect sequential models to exploit the temporal relationship. The pyramid module integrates the recurrent module, thus can iteratively synthesize temporally smooth results. And the pyramid modules share weights across iterations, thus it does not expand the model's parameter size. Our model can be generalized to several applications such as up-converting the frame rate of videos with motion blur, reducing compression artifacts, and jointly super-resolving low-resolution videos. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art methods on various video frame interpolation and enhancement tasks.
Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021, Xiongkuo Min
IEEE Trans. Image Process.1
2020 Blurry Video Frame Interpolation
abstract
Existing works reduce motion blur and up-convert frame rate through two separate ways, including frame deblurring and frame interpolation. However, few studies have approached the joint video enhancement problem, namely synthesizing high-frame-rate clear results from low-frame-rate blurry inputs. In this paper, we propose a blurry video frame interpolation method to reduce motion blur and up-convert frame rate simultaneously. Specifically, we develop a pyramid module to cyclically synthesize clear intermediate frames. The pyramid module features adjustable spatial receptive field and temporal scope, thus contributing to controllable computational complexity and restoration ability. Besides, we propose an inter-pyramid recurrent module to connect sequential models to exploit the temporal relationship. The pyramid module integrates a recurrent module, thus can iteratively synthesize temporally smooth results without significantly increasing the model size. Extensive experimental results demonstrate that our method performs favorably against state-of-the-art methods. The source code and pre-trained model are available at https://github.com/laomao0/BIN.
Wang Shen, Wenbo Bao, Guangtao Zhai, Li Chen 0021, Xiongkuo Min
CVPR1
2020 Improving biomedical named entity recognition with syntactic information
abstract
BACKGROUND: Biomedical named entity recognition (BioNER) is an important task for understanding biomedical texts, which can be challenging due to the lack of large-scale labeled training data and domain knowledge. To address the challenge, in addition to using powerful encoders (e.g., biLSTM and BioBERT), one possible method is to leverage extra knowledge that is easy to obtain. Previous studies have shown that auto-processed syntactic information can be a useful resource to improve model performance, but their approaches are limited to directly concatenating the embeddings of syntactic information to the input word embeddings. Therefore, such syntactic information is leveraged in an inflexible way, where inaccurate one may hurt model performance. RESULTS: In this paper, we propose BIOKMNER, a BioNER model for biomedical texts with key-value memory networks (KVMN) to incorporate auto-processed syntactic information. We evaluate BIOKMNER on six English biomedical datasets, where our method with KVMN outperforms the strong baseline method, namely, BioBERT, from the previous study on all datasets. Specifically, the F1 scores of our best performing model are 85.29% on BC2GM, 77.83% on JNLPBA, 94.22% on BC5CDR-chemical, 90.08% on NCBI-disease, 89.24% on LINNAEUS, and 76.33% on Species-800, where state-of-the-art performance is obtained on four of them (i.e., BC2GM, BC5CDR-chemical, NCBI-disease, and Species-800). CONCLUSION: The experimental results on six English benchmark datasets demonstrate that auto-processed syntactic information can be a useful resource for BioNER and our method with KVMN can appropriately leverage such information to improve model performance.
Yuanhe Tian, Wang Shen, Yan Song 0003, Fei Xia 0004, Kenli Li 0001
BMC Bioinform.2
2019 A QoE-oriented Saliency-aware Approach for 360-degree Video Transmission
abstract
The tradeoff between bandwidth efficiency and quality of experience (QoE) is a key issue in 360 video transmission. In this paper, we propose a QoE-oriented saliency-aware 360 video transmission framework to balance this tradeoff. The target is to reduce the bandwidth demand without declining the QoE. Specifically, the proposed model is based on the decision-making process. We use Lyapunov optimization to solve the decisionmaking problem. Furthermore, we integrate saliency information into the model to influence the decision policy, so that the model has the advantage of bandwidth efficiency. The simulation results show that the tradeoff parameter of Lyapunov optimization can balance the tradeoff between QoE and bandwidth efficiency, and 360 video saliency entropy limits the upper and lower bounds of QoE and bandwidth efficiency.
Wang Shen, Lianghui Ding, Guangtao Zhai, Ying Cui 0001
VCIP1