EDBT 2026 Demo / reviewers in the wild / expert
Shurun Wang
dblp:192/8544
· DBLP profile ↗
16ranked-venue papers
9as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improved data-driven model-free adaptive control method for an upper extremity power-assist exoskeleton
Shurun Wang, Hao Tang 0009, Zhaowu Ping, Bin Wang 0095 |
Appl. Intell. | 1 |
| 2025 | High Efficiency Image Compression for Large Visual-Language ModelsabstractIn recent years, large visual language models (LVLMs) have shown impressive performance and promising generalization capability in multi-modal tasks, thus replacing humans as receivers of visual information in various application scenarios. In this paper, we pioneer to propose a variable bitrate image compression scheme consisting of a pre-editing module and an end-to-end codec to achieve promising rate-accuracy performance for different LVLMs. In particular, instead of optimizing an adaptive pre-editing network towards a particular task or several representative tasks, we propose a new optimization strategy tailored for LVLMs, which is designed based on the representation and discrimination capability with token-level distortion and rank. The pre-editing module and the variable bitrate end-to-end image codec are jointly trained by the losses based on semantic tokens of the large model, which introduce enhanced generalization capability for various data and tasks. Experimental results demonstrate that the proposed framework could efficiently achieve much better rate-accuracy performance compared to the state-of-the-art coding standard, Versatile Video Coding. Meanwhile, experiments with multi-modal tasks have revealed the robustness and generalization capability of the proposed framework. Binzhe Li, Shurun Wang, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Interactive Face Video Coding: A Generative Compression FrameworkabstractIn this paper, we propose a novel framework for Interactive Face Video Coding (IFVC), which allows humans to interact with the intrinsic visual representations instead of the signals. The proposed solution enjoys several distinct advantages, including ultra-compact representation, low delay interaction, and vivid expression/headpose animation. In particular, we propose the Internal Dimension Increase (IDI) based representation, greatly enhancing the fidelity and flexibility in rendering the appearance while maintaining reasonable representation cost. By leveraging strong statistical regularities, the visual signals can be effectively projected into controllable semantics in the three dimensional space (e.g., mouth motion, eye blinking, head rotation, head translation and head location), which are compressed and transmitted. The editable bitstream, which naturally supports the interactivity at the semantic level, can synthesize the face frames via the strong inference ability of the deep generative model. Experimental results have demonstrated the performance superiority and application prospects of our proposed IFVC scheme. In particular, the proposed scheme not only outperforms the state-of-the-art video coding standard Versatile Video Coding (VVC) and the latest generative compression schemes in terms of rate-distortion performance for face videos, but also enables the interactive coding without introducing additional manipulation processes. Furthermore, the proposed framework is expected to shed lights on the future design of the digital human communication in the metaverse. The project page can be found at https://github.com/Berlin0610/Interactive_Face_Video_Coding. Zhao Wang 0004, Binzhe Li, Shurun Wang, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Image Process. | 4 |
| 2024 | Signaling of object masks with the assistance of the object mask information SEI messageabstractTo reduce the decoder burden of performing the video analysis tasks and increase the accuracy of the task results, an effective solution is to perform the video analysis task in the encoder side and send the results to the decoder. In this paper, the object masks which are obtained by invoking a video analysis model in the encoder are represented by object mask pictures and encoded as auxiliary pictures. To enable this idea, a new auxiliary picture type was specified and a new supplemental enhancement information (SEI) message, the object mask information (OMI) SEI message, is introduced. The experiments are conducted by using versatile video coding test model (VTM) layered coding, and according to the experimental results, it is concluded that coding the object mask pictures as auxiliary pictures is feasible and efficient. And thus, the proposed OMI SEI message was adopted to working draft of versatile supplemental enhancement information version 4 and included in the technical report of optimization of encoders and receiving systems for machine analysis of coded video content in January 2024. Jie Chen 0006, Zixiang Zhang, Yan Ye 0003, Shurun Wang |
VCIP | 4 |
| 2024 | Integrated block-wise neural network with auto-learning search framework for finger gesture recognition using sEMG signals
Shurun Wang, Hao Tang 0009, Qi Jiang 0004 |
Artif. Intell. Medicine | 1 |
| 2024 | Sparse-to-Dense: High Efficiency Rate Control for End-to-End Scale-Adaptive Video CodingabstractTraditional end-to-end video coding is typically featured with sparsely distributed operational rate-distortion (R-D) points. This creates daunting challenges to rate control which is typically regarded as the indispensable coding optimization module. To tackle this problem, this paper proposes high efficiency rate control for end-to-end scale-adaptive video coding which enables the conversion from sparsely to densely distributed R-D points. The proposed scheme does not increase the number of models in the sparse-to-dense conversion and provides more flexibility in end-to-end video coding thereby leading to better coding performance. More specifically, R-D analyses for scale-adaptive coding are first conducted, shedding light on the design of the rate control algorithm. Subsequently, generalized R-D models are presented, based on which high efficiency rate control is achieved. Extensive experimental results provide evidence of the efficiency of the proposed method in terms of R-D performance, control accuracy and computational complexity.. Jiancong Chen, Meng Wang 0017, Shurun Wang, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Ultra-Low Bitrate Face Video Compression Based on Conversions From 3D Keypoints to 2D Motion MapabstractHow to compress face video is a crucial problem for a series of online applications, such as video chat/conference, live broadcasting and remote education. Compared to other natural videos, these face-centric videos owning abundant structural information can be compactly represented and high-quality reconstructed via deep generative models, such that the promising compression performance can be achieved. However, the existing generative face video compression schemes are faced with the inconsistency between the 3D facial motion in the physical world and the face content evolution in the 2D view. To solve this drawback, we propose a 3D-Keypoint-and-2D-Motion based generative method for Face Video Compression, namely FVC-3K2M, which can well ensure perceptual compensation and visual consistency between motion description and face reconstruction. In particular, the temporal evolution of face video can be characterized into separate 3D keypoints from the global and local perspectives, entailing great coding flexibility and accurate motion representation. Moreover, a cascade motion conversion mechanism is further proposed to internally convert 3D keypoints to 2D dense motion, enforcing the face video reconstruction to be perceptually realistic. Finally, an adaptive reference frame selection scheme is developed to enhance the adaptation of various temporal movements. Experimental results show that the proposed scheme can realize reliable video communication in the extremely limited bandwidth, e.g., 2 kbps. Compared to the state-of-the-art video coding standards and the latest face video compression methods, extensive comparisons demonstrate that our proposed scheme achieves superior compression performance in terms of multiple quality evaluations. Zhao Wang 0004, Shurun Wang, Shiqi Wang 0001, Yan Ye 0003, Siwei Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Peering into The Sketch: Ultra-Low Bitrate Face Compression for Joint Human and Machine PerceptionabstractWe propose a novel face compression framework that leverages the external priors for joint human and machine perception under ultra-low bitrate scenarios. The proposed framework leverages the semantic richness of face images by representing the faces into sketches and thumbnails, resulting in improved bitrate utility for both human and machine vision. At the decoder side, the framework introduces a two-stage generative reconstruction, which faithfully enhances the reconstructed image via semi-parametric modeling and retrieved guidance from the external database. In particular, this coarse-to-fine strategy also results in improved identity consistency and analysis performance of the reconstructed image. Extensive evaluations of the proposed method have been conducted on the public face dataset by comparing it with end-to-end image compression techniques as well as traditional image compression standards. The experimental results demonstrate the effectiveness of the proposed method via superior perceptual and analytical performance under ultra-low bitrate conditions. Yudong Mao, Peilin Chen 0001, Shurun Wang, Shiqi Wang 0001, Dapeng Oliver Wu |
ACM Multimedia | 3 |
| 2023 | Deep Image Compression Toward Machine Vision: A Unified Optimization FrameworkabstractThere has been an increasing consensus that the machine vision is gradually replacing human vision in numerous tasks, with the demonstrated success of artificial intelligence. In this paper, we propose a deep image compression scheme towards machine vision, with the principle of “begin with the end in mind”. In particular, a unified optimization scheme for end-to-end image compression towards machine vision is proposed, accompanied with the dedicated variable bitrate coding and generalized rate-accuracy optimization. The presented framework, which jointly optimizes the compression and the machine vision networks, exploits the utmost potential of robust machine vision for compressed images. The variable bitrate modules towards machine vision, which effectively shrink the storage space for model parameters, are further developed to accommodate to the real-world applications. Moreover, an iterative algorithm is presented to achieve the optimality in terms of the generalized rate-accuracy towards machine vision. Experimental results show that the proposed framework achieves the state-of-the-art object detection performance among the end-to-end image compression methods: in the exploration of Video Coding for Machines (VCM) in Moving Picture Experts Group (MPEG), and the proposed framework achieves 31.69% and 23.96% BD-rate gains compared with the VCM official test datasets, the Open Images dataset and the TVD dataset respectively, which are generated using the state-of-the-art standard Versatile Video Coding (VVC) standard. The generalization capability of the proposed framework is also verified with instance segmentation under various scenarios. Shurun Wang, Zhao Wang 0004, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | A Novel Approach to Detecting Muscle Fatigue Based on sEMG by Using Neural Architecture Search FrameworkabstractMuscle fatigue detection is of great significance to human physiological activities, but many complex factors increase the difficulty of this task. In this article, we integrate several effective techniques to distinguish muscle states under fatigue and nonfatigue conditions via surface electromyography (sEMG) signals. First, we perform an isometric contraction experiment of biceps brachii to collect sEMG signals. Second, we propose a neural architecture search (NAS) framework based on reinforcement learning to autogenerate neural networks. Finally, we present an effective two-step training strategy to improve the performance by combining CNN with three types of commonly used statistical algorithms. Meanwhile, we propose a data enhancement algorithm based on empirical mode decomposition (EMD) to generate time-series data for expanding the dataset. The results show that this search algorithm can hunt for high-performing networks, and the accuracy of the best-selected model combined with support vector machine (SVM) for the group is 96.5%. With the same architecture, the average accuracy in individual models is 97.8%. The proposed data enhancement technique can effectively improve the fatigue detection performance, which allows further implementations in the human-exoskeleton interaction systems. Shurun Wang, Hao Tang 0009, Bin Wang 0095, Jia Mo |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Continuous Estimation of Human Joint Angles From sEMG Using a Multi-Feature Temporal Convolutional Attention-Based NetworkabstractIntention recognition based on surface electromyography (sEMG) signals is pivotal in human-machine interaction (HMI), where continuous motion estimation with high accuracy has been the challenge. The convolutional neural network (CNN) possesses excellent feature extraction capability. Still, it is difficult for ordinary CNN to explore the dependencies of time-series data, so most researchers adopt the recurrent neural network or its variants (e.g., LSTM) for motion estimation tasks. This paper proposes a multi-feature temporal convolutional attention-based network (MFTCAN) to recognize joint angles continuously. First, we recruited ten subjects to accomplish the signal acquisition experiments in different motion patterns. Then, we developed a joint training mechanism that integrates MFTCAN with commonly used statistical algorithms, and the integrated architectures were named MFTCAN-KNR, MFTCAN-SVR and MFTCAN-LR. Last, we utilized two performance indicators (RMSE and [Formula: see text]) to evaluate the effect of different methods. Moreover, we further validated the performance of the proposed method on the open dataset (Ninapro DB2). When evaluating on the original dataset, the average RMSE of the estimations obtained by MFTCAN-KNR is 0.14, which is significantly less than the results obtained by LSTM (0.20) and BP (0.21). The average [Formula: see text] of the estimations obtained by MFTCAN-KNR is 0.87, indicating the anti-disturbance ability of the architecture. Moreover, MFTCAN-KNR also achieves high performance when evaluating on the open dataset. The proposed methods can effectively accomplish the task of motion estimation, allowing further implementations in the human-exoskeleton interaction systems. Shurun Wang, Hao Tang 0009, Lifu Gao |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Towards Analysis-Friendly Face Representation With Scalable Feature and Texture CompressionabstractCompactly representing visual information plays a fundamental role in optimizing the ultimate utility of myriad visual data-centered applications. Numerous approaches have been proposed to efficiently compress the texture and visual features for human visual perception and machine intelligence, respectively; however, much less work has been dedicated to studying the interactions between them. Here, we investigate the integration of feature and texture compression and show that a universal and collaborative visual information representation can be achieved in a hierarchical way. In particular, we study feature and texture compression in a scalable coding framework, where the base layer serves as the deep learning feature and the enhancement layer targets to perfectly reconstruct the texture. Based on the strong generative capability of deep neural networks, the gap between the base feature layer and enhancement layer is further filled with feature-level texture reconstruction, with the goal of further constructing texture representations from features. As such, the residuals between the original and reconstructed texture could be further conveyed in the enhancement layer. To improve the efficiency of the proposed framework, the base layer neural network is trained in a multitask manner such that the learned features enjoy both high-quality reconstruction and high-accuracy analysis. The framework and optimization strategies are further applied in face image compression, and promising coding performance has been achieved in terms of both rate-fidelity and rate-accuracy evaluations. Shurun Wang, Shiqi Wang 0001, Wenhan Yang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Teacher-Student Learning With Multi-Granularity Constraint Towards Compact Facial Feature RepresentationabstractIn this paper, we propose a novel end-to-end feature compression scheme by leveraging the representation and learning capability of deep neural networks, towards intelligent front-end equipped analysis with promising accuracy and efficiency. In particular, the extracted features are compactly coded in an end-to-end manner by optimizing the rate- distortion cost to achieve feature-in-feature representation. The multi-granularity constraint is further imposed, serving as the optimization objective to make the feature compression more "healthier" from the perspective of ultimate utility. More specifically, the analysis accuracy is considered in the coarse granularity level constraint, ensuring the capability of facial analysis with the reconstructed feature. Furthermore, at the fine granularity level the feature fidelity is involved to preserve the original feature quality. Moreover, a latent code level teacher-student enhancement model is proposed to efficiently transfer the low bit-rate representation into a high bit- rate one. Such a strategy further allows us to adaptively shift the representation cost to decoding computations, leading to more flexible feature compression with enhanced decoding capability. We verify the effectiveness of the proposed model with the facial feature, and experimental results reveal better compression performance in terms of rate-accuracy compared with existing models. Shurun Wang, Shiqi Wang 0001, Wenhan Yang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001 |
ICASSP | 1 |
| 2021 | Defending Against Noise by Characterizing the Rate-Distortion Functions in End-to-End Noisy Image CompressionabstractThere has been an increasing consensus that precise understanding of the rate-distortion (RD) characteristics plays a critical role in image and video coding. In this paper, we explore the RD behaviors of end-to-end image compression in the real-world application scenario that the images could be corrupted by noise at different levels. With the RD behaviors that all images share, we develop a deep learning driven pre-analytical model which fully exploits the properties of RD functions and allows us to improve the quality with economized coding bits. The proposed approach does not require any prior knowledge of the noise level, and could effectively defend against the noise through the end-to-end compression. Extensive experimental results show that the proposed scheme offers the best promise in predicting RD behaviors, and naturally avoids the unnecessary bits consumption. Binzhe Li, Shurun Wang, Shiqi Wang 0001 |
ICIP | 2 |
| 2019 | Scalable Facial Image Compression with Deep Feature ReconstructionabstractIn this paper, we propose a scalable image compression scheme, including the base layer for feature representation and enhancement layer for texture representation. More specifically, the base layer is designed as the deep learning feature for analysis purpose, and it can also be converted to the fine structure with deep feature reconstruction. The enhancement layer, which serves to compress the residuals between the input image and the signals generated from the base layer, aims to faithfully reconstruct the input texture. The proposed scheme can feasibly inherit the advantages of both compress-then-analyze and analyze-then-compress schemes in surveillance applications. The performance of this framework is validated with facial images, and the conducted experiments provide useful evidences to show that the proposed framework can achieve better rate-accuracy and rate-distortion performance over conventional image compression schemes. Shurun Wang, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 1 |
| 2016 | Improved entropy of primitive for visual information estimationabstractSparse representation has been observed to be highly efficient in dealing with rich, varied and directional information in natural scenes. Based on the statistical analysis of primitives in sparse coding, the entropy of primitive (EoP) was proposed for measuring visual information of images, and its changing tendency has been shown to be highly relevant with the human visual system (HVS). But the sparse coefficient energy was ignored when calculating EoP, which may be critical in accounting for the primitive characteristics. To tackle this, an improved EoP is developed in this work via ℓ2norm calculation. We further give mathematical derivations for its convergence verification. Experimental evaluations have also demonstrated that the improved EoP can achieve more stable convergence tendencies, which is consistent with the perceptual experiences. Shurun Wang, Zhenghui Zhao, Xiang Zhang 0004, Jian Zhang 0018, Shiqi Wang 0001, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 1 |