Shiqing Zhang

dblp:94/2470 · DBLP profile ↗
← Back
53ranked-venue papers
16as first author
36since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 11 since 2021Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Two-stage multiple instance learning networks with attention-based hybrid aggregation for speech emotion recognition
Shiqing Zhang, Xiaoming Zhao 0002
Comput. Speech Lang.1
2026 Flame detection technology for buildings safety based on wireless sensing
Jiaqian Bao, Yong Xiong, Shiqing Zhang, Liangliang Lou
Eng. Appl. Artif. Intell.5
2026 A global-to-local state space model with context-mixing dynamic kernels for medical image classification
Yuehui Liao, Pengxia Yue, Panfei Li, Guang Yang 0006, Qun Jin, Shiqing Zhang, Xiaobo Lai, Qi Tian 0001
Expert Syst. Appl.8
2026 A hybrid CNN-Mamba state space model with pyramid-pooled skip connections for prostate tumor segmentation
Xueting Wei, Yuehui Liao, Shiqing Zhang, Guang Yang 0006, Qun Jin, Xiaobo Lai, Qi Tian 0001
Expert Syst. Appl.4
2026 Decoupled Hierarchical Distillation for Multimodal Emotion Recognition
abstract
Human multimodal emotion recognition (MER) seeks to infer human emotions by integrating information from language, visual, and acoustic modalities. Although existing MER approaches have achieved promising results, they still struggle with inherent multimodal heterogeneities and varying contributions from different modalities. To address these challenges, we propose a novel framework, Decoupled Hierarchical Multimodal Distillation (DHMD). DHMD decouples each modality's features into modality-irrelevant (homogeneous) and modality-exclusive (heterogeneous) components using a self-regression mechanism. The framework employs a two-stage knowledge distillation (KD) strategy: (1) coarse-grained KD via a Graph Distillation Unit (GD-Unit) in each decoupled feature space, where a dynamic graph facilitates adaptive distillation among modalities, and (2) fine-grained KD through a cross-modal dictionary matching mechanism, which aligns semantic granularities across modalities to produce more discriminative MER representations. This hierarchical distillation approach enables flexible knowledge transfer and effectively improves cross-modal feature alignment. Experimental results demonstrate that DHMD consistently outperforms state-of-the-art MER methods, achieving 1.3%/2.4% (ACC$_{7}$7), 1.3%/1.9% (ACC$_{2}$2) and 1.9%/1.8% (F1) relative improvement on CMU-MOSI/CMU-MOSEI dataset, respectively. Meanwhile, visualization results reveal that both the graph edges and dictionary activations in DHMD exhibit meaningful distribution patterns across modality-irrelevant/-exclusive feature spaces.
Yong Li 0032, Yuanzhi Wang, Yi Ding 0012, Shiqing Zhang, Ke Lu 0002, Cuntai Guan
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Convolution-aware multi-view graph contrastive learning
Yuhao Tao, Shuchang Zhao, Shiqing Zhang
Pattern Recognit.6
2026 Multi-Modal Cross-Attention-Guided Network for Audio-Visual Quality Evaluation via Visual Saliency and Mel-Spectrum Features
abstract
The quality evaluation of audio-visual (A/V) content has become increasingly critical in modern multimedia communication systems. Traditional single-modality quality evaluation methods and existing dedicated A/V quality models often fail to accurately assess the quality of A/V signals. To address this challenge, we propose a novel multi-modal cross-attention guided network specifically designed for A/V quality evaluation. By leveraging visual saliency and Mel-spectrum features, our network aims to achieve accurate and comprehensive quality evaluation. Specifically, distorted video frames are first converted into saliency maps, from which perceptually salient patches are selectively extracted and fed into a Convolutional Neural Network (CNN) for intra-frame visual feature extraction. Concurrently, the distorted audio signal is transformed into a Mel-spectrum, and time-frequency patches are extracted via sliding window techniques for CNN-based audio feature extraction. To effectively integrate these features and capture the long-term dependencies across consecutive A/V segments, we design a multi-modal cross-attention module that explicitly models complex inter-modal interactions. The resulting representations are then passed through a series of fully-connected (FC) layers for dimensionality reduction, ultimately deriving the quality score. Extensive experiments on three publicly available A/V quality datasets indicate that our metric outperforms the traditional quality metrics and newly-developed A/V quality metrics. The source code will be released at https://github.com/Jour3141/avqa.
Yueli Cui, Chenli Fang, Binghong Pan, Chencheng Pan, Gangyi Jiang, Shiqing Zhang, Siwei Ma 0001, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 Joint Error Detection and Correction for Safety Communication: Packet Fragmentation and Assembling
abstract
In today's industrial Internet of Things systems, functional safety communication protocols are widely adopted to transmit safety protocol data unit (SPDU). While the guessing random additive noise decoding (GRAND) algorithm can improve the reliability of cyclic redundancy check (CRC)-coded SPDU, the decoding complexity of long SPDU remains too high for practical deployment. To address this, we extend the GRAND-based joint error detection and correction (JEDeC) strategy to long SPDU and propose a JEDeC-based packet fragmentation/assembling mechanism that fragments a long SPDU into multiple short SPDUs for parallel correction of erroneous bits. We clarify the decoder input settings and channel model used in this work. We show that the proposed approach achieves tractable decoding complexity and latency: for an assembled SPDU length of 1024 bits, fragment length of 64 bits, CRC signature length of 32 bits, and maximum error-correction capability of 4 bits, under a representative bit error rate (BER)$P_{e}=10^{-3}$, the BER is reduced from$10^{-3}$to$1.49\times 10^{-8}$, the packet error rate from$6.46\times 10^{-1}$to$8.11\times 10^{-7}$, and the residual error probability from$6.36\times 10^{-11}$to$3.42\times 10^{-16}$, with an average of$2.86\times 10^{3}$guessing attempts per assembled SPDU. These results indicate that the fragmentation/assembling mechanism can substantially improve safety-communication dependability with implementable cost.
Ming Zhan, Zhibo Pang, Jiangwu Zhang, Shiqing Zhang, Kan Yu 0002
IEEE Trans. Ind. Informatics6
2025 Lifting Wavelet Transform-Based Network for Liver Segmentation in CT Scans
abstract
Liver segmentation plays a crucial role in the diagnosis and surgical planning of hepatocellular carcinoma. However, manual delineation of liver contours by radiologists is time-consuming, error-prone, and highly dependent on individual expertise. To address these challenges, we propose a Lifting Wavelet Transform-based Network (LWT-Net) for liver segmentation in CT scans. Built upon an encoder-decoder architecture, the proposed model incorporates a lifting wavelet transform module to perform multi-scale frequency-domain features. Specifically, the low-frequency components capture the global structural information, facilitating the modeling of contextual semantics, while the high-frequency components preserve finegrained details, thereby enhancing segmentation accuracy. In addition, to effectively integrate the multi-scale wavelet features with global contextual representations extracted by a CNN-based encoder, the double attention modules are applied to capture longrange spatial dependencies. Extensive experiments on two public datasets, LiTS2017 and FLARE22, demonstrate that LWT-Net achieves superior performance, with average Dice coefficients of 95.75% and 95.97%, respectively, significantly outperforming existing state-of-the-art methods. The source code will be publicly available on GitHub.
Huaxiang Liu, Baicheng Qu, Youyao Fu, Shiqing Zhang, Wenbin Ji, Jiangxiong Fang
BIBM5
2025 Flex8: A Flexible Precision Co-design for 8-bit Neural Network
abstract
The rapid growth of neural network parameters poses a significant challenge for resource-constrained training, particularly in terms of storage and computation.While 8-bit quantization techniques have shown promise, their practical application remains limited, and mixed-precision methods have yet to achieve optimal results.This paper systematically evaluates the strengths and limitations of FP8 quantization by analyzing the impact of network architecture, exponent width, and mantissa width on FP8 training performance.Additionally, we investigate the effects of low-precision quantization across different layers, parameter types, and training stages.Based on these insights, we propose Flex8, a hardware-software co-design framework for mixed-precision optimization that integrates fine-grained quantization strategies, tunable exponent-width instructions, and a flexible 8-bit FPU design.Flex8 significantly improves FP8 training accuracy to levels comparable to singleprecision, offering a novel and effective approach to low-precision training.
Lu Wang 0019, Guangda Zhang, Xia Zhao 0004, Shiqing Zhang
CF5
2025 NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUs
abstract
As Graphics Processing Units (GPUs) face increasing computing demands that surpass single-module capabilities due to transistor scaling and lithography constraints, the necessity for expanding the module count within GPUs grows. This escalation faces a significant challenge: the total inter-module bandwidth in many-chip-module GPUs is limited by manufacturing constraints in organic substrates or silicon interposers. Unlike Central Processing Units (CPUs), which are latency-sensitive, GPUs leverage their high thread-level parallelism to effectively hide memory access latency through simultaneous multithreading. This attribute makes GPUs inherently sensitive to bandwidth constraints, making the efficient exploitation of available inter-module bandwidth important. In this paper, we identify that fetching data from faraway memory in many-chip-module GPUs can easily cause bandwidth contention which degrades the real achieved data bandwidth per GPU module compared to fetching data from nearby memory. To further analyze this problem, we introduce the Inter-Module Bandwidth per Access (IBPA) metric for quantifying bandwidth usage and finding the network hop count directly impacts the IBPA and network contention. Next, we propose NearFetch, a routing-based solution to reduce IBPA. NearFetch works due to the fact that GPU modules along the routing path are typically much closer to the source GPU module while these GPU modules can supply $29.1 \%$ of data for high-sharing applications. NearFetch consists of two primary components: a data forwarding scheme, enabling data forwarding when the data resides in a remote GPU module, and a topology-aware Miss Status Handling Register (MSHR) coalescing scheme, responsible for recording the memory address information for future use in case of a data miss. By leveraging the data locality among various GPU modules, NearFetch substantially minimizes inter-module bandwidth usage, eliminating the need to fetch data from distant memory partitions. Our evaluation of NearFetch within the context of many-chip-module GPUs, across applications exhibiting diverse degrees of data locality, reveals that it reduces IBPA by $4 2. 6 \%$ and enhances performance by an average of $52.2 \%$ (with up to $9 8. 1 \%$ improvement) for high-sharing workloads.
Guangda Zhang, Shiqing Zhang, Huadong Dai
HPCA4
2025 Symmetric Bi-branch Modality-search Aggregation Network for Multi-modal Liver Segmentation
abstract
Medical image segmentation is crucial for diagnosis and surgical planning of liver diseases. The existing methods mainly focus on global or local features and neglect spatial dependencies among modalities and blurred boundaries. To tackle these challenges, we propose a symmetric bi-branch modality-search aggregation network (SBMANet). Specifically, we first design a symmetric network with dual encoder-decoder structure. Each encoder fuses two adjacent modal features to improve the intra-modal spatial information while the decoder can achieve accurate localization of segmented targets. To fully exploit multi-modal inter-modal dependencies, a hybrid Hadamard multimodal fusion module (HHMF) is proposed. Finally, we establish an adaptive modality-channel-search module by incorporating bi-branch features to automatically compute weights for each channel in different modalities. Extensive experiments on DLDS demonstrate that the proposed network outperforms existing state-of-the-art 3D segmentation networks. The code is available at the website: https://github.com/fangchj2002/SBMANet.
Huaxiang Liu, Youyao Fu, Shiqing Zhang, Wenbin Ji, Jiangxiong Fang
ICASSP5
2025 Distribution-Aware Multi-Attention Tri-branch Networks with Feedforward Differential Features for semi-supervised medical image segmentation
Peilian Shi, Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Hongsheng Lu, Jun Yu 0002
Expert Syst. Appl.5
2025 Human-Fire Classification Method Based on Wireless Sensing Technology
abstract
High-rise fires pose significant risks due to their complex structures and high population density, making timely and effective detection methods essential. Compared to traditional sensor-or camera-based detection methods, wireless sensingbased human-fire classification (HFC) methods offer significant cost-effectiveness advantages. While Channel State Information (CSI) provides more detailed and accurate channel data compared to Received Signal Strength (RSS), its acquisition and processing are more challenging. On the other hand, RSS-based methods are easier to implement and more cost-effective. Therefore, selecting appropriate methods for specific scenarios is crucial to achieving a balance between performance and cost-effectiveness. To this end, this paper proposes an innovative method called Wi-HFC, which leverages deep learning to evaluate the performance of RSS and CSI in various human-fire classification tasks. Specifically, a dataset of RSS and CSI data from five classification tasks in real fire scenarios was collected and evaluated using a custom-designed convolutional neural network-based deep learning model. Experimental results indicate that RSS is a cost-effective choice for short distances or simple environments, whereas CSI demonstrates significant advantages in scenarios requiring higher accuracy and involving greater environmental complexity. Furthermore, the developed dataset is publicly available at https://github.com/T-bjq/Wi-HFC-dataset, providing resources for further research.
Liangliang Lou, Jiaqian Bao, Yong Xiong, Shiqing Zhang
IEEE Internet Things J.6
2025 PACMR: Progressive Adaptive Crossmodal Reinforcement for Multimodal Apparent Personality Traits Analysis
abstract
Multimodal apparent personality traits analysis is a challenging issue due to the asynchrony among modalities. To address this issue, this paper proposes a Progressive Adaptive Crossmodal Reinforcement (PACMR) approach for multimodal apparent personality traits analysis. PACMR adopts a progressive reinforcement strategy to provide a multi-level information exchange among different modalities for crossmodal interactions, resulting in reinforcing the source and target modalities simultaneously. Specifically, PACMR introduces an Adaptive Modality Reinforcement Unit (AMRU) to adaptively adjust the weights of self-attention and crossmodal attention for capturing reliable contextual dependencies of multimodal sequence data. Experiment results on the public First Impressions dataset demonstrate the effectiveness of the proposed method.
Shiqing Zhang, Xiaoming Zhao 0002
IEEE Signal Process. Lett.4
2025 Examining the Fourier Spectrum of Speech Signal From a Time-Frequency Perspective for Automatic Depression Level Prediction
abstract
Currently, many studies use Fourier amplitude spectra of speech signals to predict depression levels. However, those works often treat Fourier amplitude spectra as images or sequences to capture depression cues using convolutional neural networks or multilayer perceptrons. Therefore, they ignore the complex element composition and time-frequency attributes of Fourier spectra, which is not conducive to capturing the differences among individuals with different depression levels. For this reason, we construct a Time-Frequency Self-Embedding (TFSE) module, which not only stores the correlation relationship among real (imaginary) parts of Fourier spectra of different subjects from the time-frequency perspective, but also maintain the physical properties of data through the weight embedding process. Besides, Global Average Pooling (GAP) or linear layers are difficult to balance both temporal and frequency dimensions in the vectorization process. Therefore, we construct a Time-Frequency Tensor Vectorization (TFTV) module, which summarizes each channel along time and frequency dimensions, and then generates the vectorization result by integrating various channels. In this way, we combine TFSE and TFTV modules to form our SpectrumFormer model for predicting depression levels. Evaluation indicators on AVEC 2013 and AVEC 2014 depression databases imply the progressiveness of our model.
Mingyue Niu, Jianhua Tao 0001, Yongjun He 0002, Shiqing Zhang, Ming Li 0065
IEEE Trans. Affect. Comput.4
2025 Theoretical Bound and Compensation for Residual Error Probability of GRAND-CRC-Based Functional Safety Communication
abstract
In modern Industrial Internet of Things (IIoT) ecosystems, functional safety protocols are increasingly utilized to transmit safety protocol data unit (SPDU). Integrating the universal guessing random additive noise decoding (GRAND) algorithm with cyclic redundancy check (CRC)-coded SPDU can minimize SPDU retransmissions. However, it introduces residual error probability (REP) degradation that requires careful consideration. Using the IEC 61784-3 Standard and CRC assumptions, we derive a closed-form REP evaluation formula specific to SPDU length and maximum error correction capability. Our analysis reveals that, under the worst industrial conditions and for an SPDU length of 128 bits, the REP performance degrades by approximately 2:2 × 102times when the CRC signature length is 24 bits and up to one bit is guessing decoded. This degradation becomes more pronounced with increased error correction capability. To address this, we propose to compensate the REP degradation by adopting longer CRC signature. This paper provides a theoretical framework for adopting the GRAND algorithm in functional safety communication, setting a foundation for enhancing reliability in IIoT applications..
Ming Zhan, Zhibo Pang, Shiqing Zhang, Jianwu Zhang, Kan Yu 0002
IEEE Trans. Commun.5
2025 MDKAT: Multimodal Decoupling With Knowledge Aggregation and Transfer for Video Emotion Recognition
abstract
Multimodal Emotion Recognition (MER) leverages multiple input signals to identify the expressed emotions in user-generated data. Currently, effectively addressing both modality heterogeneity and homogeneity on MER tasks is a challenging issue due to the diversity of multimodal inputs in videos. To address this issue, this work proposes an efficient Multimodal Decoupling Method with Knowledge Aggregation and Transfer (MDKAT) for robust multimodal feature learning in emotional videos. MDKAT is consisted of three key steps: modality-independent feature extraction, modality-specific feature extraction, and multi-loss integration for decoupling. In these three steps, four crucial modules are individually designed to improve different aspects of multimodal learning on MER tasks, including a Cross-modal Feature Fusion (CFF) module for enhancing modality-independent features, an Adaptive Masked Self-Attention (AMSA) module for feature refinement, a Knowledge Aggregation (KA) module for ensuring the semantic similarity of modality-independent features, and a Knowledge Transfer (KT) module for balancing the strengths of different modalities. Experimental results on the typical CMU-MOSI and CMU-MOSEI datasets show that MDKAT obtains superior performance over state-of-the-art methods, demonstrating the effectiveness of MDKAT on MER tasks.
Jian Wang 0066, Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jun Yu 0002, Yaowei Wang 0001, Yi Yang 0001, Siwei Ma 0001, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Multi-Scale Group Agent Attention-Based Graph Convolutional Decoding Networks for 2D Medical Image Segmentation
abstract
Automated medical image segmentation plays a crucial role in assisting doctors in diagnosing diseases. Feature decoding is a critical yet challenging issue for medical image segmentation. To address this issue, this work proposes a novel feature decoding network, called multi-scale group agent attention-based graph convolutional decoding networks (MSGAA-GCDN), to learn local-global features in graph structures for 2D medical image segmentation. The proposed MSGAA-GCDN combines graph convolutional network (GCN) and a lightweight multi-scale group agent attention (MSGAA) mechanism to represent features globally and locally within a graph structure. Moreover, in skip connections a simple yet efficient attention-based upsampling convolution fusion (AUCF) module is designed to enhance encoder-decoder feature fusion in both channel and spatial dimensions. Extensive experiments are conducted on three typical medical image segmentation tasks, namely Synapse abdominal multi-organs, Cardiac organs, and Polyp lesions. Experimental results demonstrate that the proposed MSGAA-GCDN outperforms the state-of-the-art methods, and the designed MSGAA is a lightweight yet effective attention architecture. The proposed MSGAA-GCDN can be easily taken as a plug-and-play decoder cascaded with other encoders for general medical image segmentation tasks.
Shuchang Zhao, Shiqing Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Hongsheng Lu, Jun Yu 0002, Qi Tian 0001
IEEE J. Biomed. Health Informatics4
2024 Integrating Representation Subspace Mapping with Unimodal Auxiliary Loss for Attention-based Multimodal Emotion Recognition
abstract
Multimodal emotion recognition (MER) aims to identify emotions by utilizing affective information from multiple modalities. Due to the inherent disparities among these heterogeneous modalities, there is a large modality gap in their representations, leading to the challenge of fusing multiple modalities for MER. To address this issue, this work proposes a novel attention-based MER framework by integrating representation subspace mapping with unimodal auxiliary loss for enhancing multimodal fusion capabilities. Initially, a representation subspace mapping module is proposed to map each modality into two distinct subspaces. One is modality-public, enabling the acquisition of common representations and reducing the discrepancies across modalities. The other is modality-unique, retaining the unique characteristics of each modality while eliminating redundant inter-modal attributes. Then, a cross-modality attention is leveraged to bridge the modality gap in unique representations and facilitate modality adaptation. Additionally, our method designs an unimodal auxiliary loss to remove the noise unrelated to emotion classification, resulting in robust and meaningful representations for MER. Comprehensive experiments are conducted on the IEMOCAP and MSP-Improv datasets, and experiment results show that our method achieves superior performance to state-of-the-art MER methods. Keywords: Multimodal emotion recognition, representation subspace mapping, cross-modality attention, unimodal auxiliary loss, fusion
Xulong Du, Xingnan Zhang, Shiqing Zhang, Xiaoming Zhao 0002, Jun Yu 0002, Liangliang Lou
LREC/COLING6
2024 AdCoalescer: An Adaptive Coalescer to Reduce the Inter-Module Traffic in MCM-GPUs
abstract
The demand for greater computing power has driven the development of Multi-chip-module GPUs (MCM-GPUs), which greatly improve parallel processing capabilities. Unfortunately, MCM-GPUs have encountered a notable challenge, the performance bottleneck caused by remote accesses through the inter-module network. In this work, we found significant data access redundancy among SMs within a GPU module which can be coalesced to reduce network pressure. However, how to design the coalescing scheme to identify memory addresses with high data locality is still not clear.
Xu Zhang 0086, Guangda Zhang, Lu Wang 0019, Shiqing Zhang, Xia Zhao 0004
ICPP4
2024 Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects
Shiqing Zhang, Yijiao Yang, Xingnan Zhang, Qingming Leng, Xiaoming Zhao 0002
Expert Syst. Appl.1
2024 Wireless-Sensing-Based Human-Vehicle Classification Method via Deep Learning: Analysis and Implementation
abstract
Wireless sensing methods for human-vehicle classification (WHVC) offer cost-effective advantages and enhance the detection efficiency of traffic parameters in intelligent transportation systems (ITSs). Existing WHVC methods primarily utilize channel state information (CSI) or received signal strength (RSS) features extracted from the surrounding wireless signals. Although CSI data provides more detailed and accurate channel information compared to RSS data, extracting and processing CSI is more challenging than RSS. Moreover, for applications that do not require fine-grained human-vehicle classification, such as intelligent street lighting systems, RSS-based WHVC has the advantages of easy implementation and low cost. Therefore, investigating the performance of CSI-and RSS-based WHVC methods in different application scenarios could provide valuable insights for the WHVC domain. To address this issue, this paper proposes a deep learning-based WHVC method, which employs deep learning as a tool to evaluate the performance of RSS and CSI methods in various classification tasks. Specifically, this paper collects CSI and RSS data for seven different classification tasks in real traffic road scenarios and evaluates these tasks using a convolutional neural network-based deep learning model designed in this paper. Experimental results demonstrate that for road user categories less than four, RSS-based WHVC achieves higher accuracy than CSI-based WHVC. However, as the number of categories increases, CSI-based WHVC exhibits superior accuracy compared to RSS-based WHVC. Additionally, the developed dataset is publicly available at https://github.com/TZ-mx/mixeddataset.
Liangliang Lou, Mingxin Song, Xiaoming Zhao 0002, Shiqing Zhang, Ming Zhan
IEEE Internet Things J.5
2024 Blind quality evaluation for tone-mapped images by exploiting statistical characteristics and deep perceptual features
Qiuzi Ruan, Siwen Cai, Yueli Cui, Yonglong Cui, Shuitu Li, Shiqing Zhang
Multim. Syst.10
2024 DSC-Ghost-Conv: A compact convolution module for building efficient neural network architectures
Shiqing Zhang
Multim. Tools Appl.2
2024 MTDAN: A Lightweight Multi-Scale Temporal Difference Attention Networks for Automated Video Depression Detection
abstract
Deep learning based video depression analysis has been recently an interesting and challenging topic. Most of existing works focus on learning single-scale facial dynamics of participants for depression detection. Besides, they usually adopt expensive deep learning models with high computational complexity, resulting in difficulty in real-time clinical applications. To address these two issues, this work proposes a lightweight Multi-scale Temporal Difference Attention Networks (MTDAN) integrating the temporal difference and attention mechanism to model both short-term and long-term temporal facial behaviors for automated video depression detection. Initially, two simple yet effective sub-branches, i.e., a Short-term Temporal Difference Attention Network (ST-TDAN), and a Long-term Temporal Difference Attention Network (LT-TDAN), are designed to perform individually short-term and long-term depressive behavior modeling. Then, a simple Interactive Multi-head Attention Fusion (IMHAF) strategy is employed for integrating short-term and long-term spatiotemporal features, followed by a linear fully-collected layer for depression score prediction. Experiments on two public AVEC2013 and AVEC2014 datasets show that our proposed method not only achieves highly competitive performance to state-of-the-art methods, but also has much smaller computational complexity than them on video depression detection tasks.
Shiqing Zhang, Xingnan Zhang, Xiaoming Zhao 0002, Jiangxiong Fang, Mingyue Niu, Ziping Zhao 0001, Jun Yu 0002, Qi Tian 0001
IEEE Trans. Affect. Comput.1
2024 Optimized Wireless Sensing and Deep Learning for Enhanced Human-Vehicle Recognition
abstract
In the realm of traffic parameter measurement, wireless sensing-based human-vehicle recognition methods have been pivotal due to their low cost and non-invasive nature. Traditionally, these methods have relied on the 2.4 GHz frequency band, often neglecting the rich potential of the sub-GHz bands. Furthermore, the energy attributes of wireless signals are influenced by both antenna height and carrier frequency, yet few studies have explored their impact on human-vehicle recognition performance. Addressing this critical research gap, this study introduces an innovative convolutional neural network-based method that leverages both sub-GHz bands and variable antenna heights. Specifically, this paper focuses on two key aspects: received signal strength signal-to-noise ratio analysis and wireless sensing-based human-vehicle recognition performance analysis. Experimental results demonstrate that the optimal human-vehicle recognition performance is achieved with 2.4 GHz wireless signals and an antenna height of 0.8 m, resulting in an average vehicle recognition accuracy of 95.96%. Besides, the dataset with various carrier frequencies and antenna heights has been publicly available athttps://github.com/TZ-mx/mixed-RSS-dataset.
Liangliang Lou, Mingxin Song, Xiaoming Zhao 0002, Shiqing Zhang
IEEE Trans. Intell. Transp. Syst.5
2023 SAC: Sharing-Aware Caching in Multi-Chip GPUs
abstract
Bandwidth non-uniformity in multi-chip GPUs poses a major design challenge for its last-level cache (LLC) architecture. Whereas a memory-side LLC caches data from the local memory partition while being accessible by all chips, an SM-side LLC is private to a chip while caching data from all memory partitions. We find that some workloads prefer a memory-side LLC while others prefer an SM-side LLC, and this preference solely depends on which organization maximizes the effective LLC bandwidth. In contrast to prior work which optimizes bandwidth beyond the LLC, we make the observation that the effective bandwidth ahead of the LLC is critical to end-to-end application performance. We propose Sharing-Aware Caching (SAC) to adopt either a memory-side or SM-side LLC organization by dynamically reconfiguring the routing policies in the intra-chip interconnection network and LLC controllers. SAC is driven by a simple and lightweight analytical model that predicts the impact of data sharing across chips on the effective LLC bandwidth. SAC improves average performance by 76% and 12% (and up to 157% and 49%) compared to a memory-side and SM-side LLC, respectively. We demonstrate significant performance improvements across the design space and across workloads.
Shiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven Eeckhout
ISCA1
2023 Deep learning-based panoptic segmentation: Recent advances and perspectives
abstract
Abstract In recent years, panoptic segmentation has drawn increasing amounts of attention, leading to the rapid emergence of numerous related algorithms. A variety of deep neural networks have been used more frequently for panoptic segmentation, which is motivated by the significant success of deep learning methods in other tasks. This article presents a comprehensive exploration of panoptic segmentation, focusing on the analysis and understanding of RGB image data. Initially, the authors introduce the background of panoptic segmentation, including deep learning models and image segmentation. Then, the authors thoroughly cover a variety of panoptic segmentation‐related topics, such as datasets connected to the field, evaluation metrics, panoptic segmentation models, and derived subfields based on panoptic segmentation. Finally, the authors examine the difficulties and possibilities in this area and identify its future paths.
Yuelong Chuang, Shiqing Zhang, Xiaoming Zhao 0002
IET Image Process.2
2023 Learning inter-class optical flow difference using generative adversarial networks for facial expression recognition
abstract
Abstract Facial expression recognition is a fine-grained task because different emotions have subtle facial movements. This paper proposes to learn inter-class optical flow difference using generative adversarial networks (GANs) for facial expression recognition. Initially, the proposed method employs a GAN to produce inter-class optical flow images from the difference between the static fully expressive samples and neutral expression samples. Such inter-class optical flow difference is used to highlight the displacement of facial parts between the neutral facial images and fully expressive facial images, which can avoid the disadvantage that the optical flow change between adjacent frames of the same video expression image is not obvious. Then, the proposed method designs four-channel convolutional neural networks (CNNs) to learn high-level optical flow features from the produced inter-class optical flow images, and high-level static appearance features from the fully expressive facial images, respectively. Finally, a decision-level fusion strategy is adopted to implement facial expression classification. The proposed method is validated on two public facial expression databases, BAUM_1a, SAMM and AFEW5.0, demonstrating its promising performance.
Wenping Guo, Xiaoming Zhao 0002, Shiqing Zhang, Xianzhang Pan
Multim. Tools Appl.3
2023 Characterizing Multi-Chip GPU Data Sharing
abstract
Multi-chip Graphics Processing Unit (GPU) systems are critical to scale performance beyond a single GPU chip for a wide variety of important emerging applications. A key challenge for multi-chip GPUs, though, is how to overcome the bandwidth gap between inter-chip and intra-chip communication. Accesses to shared data, i.e., data accessed by multiple chips, pose a major performance challenge as they incur remote memory accesses possibly congesting the inter-chip links and degrading overall system performance. This article characterizes the shared dataset in multi-chip GPUs in terms of (1) truly versus falsely shared data, (2) how the shared dataset scales with input size, (3) along which dimensions the shared dataset scales, and (4) how sensitive the shared dataset is with respect to the input’s characteristics, i.e., node degree and connectivity in graph workloads. We observe significant variety in scaling behavior across workloads: some workloads feature a shared dataset that scales linearly with input size, whereas others feature sublinear scaling (following a \(\sqrt {2}\) or \(\sqrt [3]{2}\) relationship). We further demonstrate how the shared dataset affects the optimum last-level cache organization (memory-side versus SM-side) in multi-chip GPUs, as well as optimum memory page allocation and thread scheduling policy. Sensitivity analyses demonstrate the insights across the broad design space.
Shiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven Eeckhout
ACM Trans. Archit. Code Optim.1
2023 An Environmental Energy Harvesting-Driven Wireless Parking Detection Method: Analysis and Implementation
abstract
The Smart Parking System (SPS) has played an important role in solving the parking problem in urban areas recently, resulting in that the battery-powered Wireless Parking Detectors (WPDs) based on Internet of Things (IoT) technology have been widely used in SPS to collect basic data sets. However, the daily turnover rates of parking spaces are generally less than 10 vehicles in Shanghai, China, so that WPDs waste a lot of energy in collecting and processing of redundant sensor data when the parking space is occupied or free. Hence, an environmental Energy Harvesting-driven Parking Detection (EHPD) method, which includes a vehicle-induced vibration energy harvesting and conditioning circuit, as well as a parking detection algorithm based on the signatures of vehicle-induced seismic signals and wireless energy attenuated by vehicles, is proposed in the paper. In EHPD, the signatures of pulse signals converted from vehicle-induced seismic signals are used to drive the wireless energy features to participate in achieving parking detection. Finally, the proposed EHPD has been evaluated in a real testbed, and the experimental results show that it can achieve the same functions as the existing magnetometer-based WPD devices with lower cost and energy consumption.
Liangliang Lou, Miao Zhou, Mingan Lu, Chunlei Zheng, Shiqing Zhang
IEEE Trans. Intell. Transp. Syst.6
2022 Unsupervised Domain Adaptation Integrating Transformer and Mutual Information for Cross-Corpus Speech Emotion Recognition
abstract
This paper focuses on an interesting task, i.e., unsupervised cross-corpus Speech Emotion Recognition (SER), in which the labelled training (source) corpus and the unlabelled testing (target) corpus have different feature distributions, resulting in the discrepancy between the source and target domains. To address this issue, this paper proposes an unsupervised domain adaptation method integrating Transformers and Mutual Information (MI) for cross-corpus SER. Initially, our method employs encoder layers of Transformers to capture long-term temporal dynamics in an utterance from the extracted segment-level log-Mel spectrogram features, thereby producing the corresponding utterance-level features for each utterance in two domains. Then, we propose an unsupervised feature decomposition method with a hybrid Max-Min MI strategy to separately learn domain-invariant features and domain-specific features from the extracted mixed utterance-level features, in which the discrepancy between two domains is eliminated as much as possible and meanwhile their individual characteristic is preserved. Finally, an interactive Multi-Head attention fusion strategy is designed to learn the complementarity between domain-invariant features and domain-specific features so that they can be interactively fused for SER. Extensive experiments on the IEMOCAP and MSP-Improv datasets demonstrate the effectiveness of our proposed method on unsupervised cross-corpus SER tasks, outperforming state-of-the-art unsupervised cross-corpus SER methods.
Shiqing Zhang, Ruixin Liu, Yijiao Yang, Xiaoming Zhao 0002, Jun Yu 0002
ACM Multimedia1
2022 Active contour driven by adaptive-scale local-energy signed pressure force function based on bias correction for medical image segmentation
abstract
Abstract This paper presents an active contour model driven by an adaptive‐scale local‐energy signed pressure force (ALSPF) function based on bias correction for segmenting medical images with intensity inhomogeneity. Firstly, a local energy‐driven signed pressure force (LESPF) function is designed as the driving force to extract local image features, which can effectively deal with intensity inhomogeneity. Secondly, to avid selecting the neighbourhood size in the LESPF function, a multi‐scale adaptive selection schema is put forward to adaptively choose the local window size according to the degree of intensity inhomogeneity. Finally, a novel single‐well potential function proved to be a strictly monotonic function is designed to ensure the stable movement of the neighbourhood points in the vicinity of the evolution curve, which can not only avoid the re‐initialization process but also enhance. Experiments on synthetic and medical images demonstrate that the proposed model is more robust than the popular ACMs for segmenting images with intensity inhomogeneity. The code is available at the website: https://github.com/HuaxiangLiu/ALSPF
Huaxiang Liu, Youyao Fu, Shiqing Zhang, Jiangxiong Fang
IET Image Process.3
2022 Spontaneous Speech Emotion Recognition Using Multiscale Deep Convolutional LSTM
abstract
Recently, emotion recognition in real sceneries such as in the wild has attracted extensive attention in affective computing, because existing spontaneous emotions in real sceneries are more challenging and difficult to identify than other emotions. Motivated by the diverse effects of different lengths of audio spectrograms on emotion identification, this paper proposes a multiscale deep convolutional long short-term memory (LSTM) framework for spontaneous speech emotion recognition. Initially, a deep convolutional neural network (CNN) model is used to learn deep segment-level features on the basis of the created image-like three channels of spectrograms. Then, a deep LSTM model is adopted on the basis of the learned segment-level CNN features to capture the temporal dependency among all divided segments in an utterance for utterance-level emotion recognition. Finally, different emotion recognition results, obtained by combining CNN with LSTM at multiple lengths of segment-level spectrograms, are integrated by using a score-level fusion strategy. Experimental results on two challenging spontaneous emotional datasets, i.e., the AFEW5.0 and BAUM-1s databases, demonstrate the promising performance of the proposed method, outperforming state-of-the-art methods.
Shiqing Zhang, Xiaoming Zhao 0002, Qi Tian 0001
IEEE Trans. Affect. Comput.1
2021 Learning deep multimodal affective features for spontaneous speech emotion recognition
Shiqing Zhang, Yuelong Chuang, Xiaoming Zhao 0002
Speech Commun.1
2020 Transparent partial page migration between CPU and GPU
Shiqing Zhang, Zheng Qin 0002, YaoHua Yang, Li Shen 0007, Zhiying Wang 0003
Frontiers Comput. Sci.1
2019 A Lightweight Method for Handling Control Divergence in GPGPUs
abstract
At present, graphics processing units (GPUs) has been widely used for scientific and high performance acceleration in the general purpose computing area, which is inseparable from the SIMT (Single-Instruction, Multiple-Thread) execution model. With SIMT, GPUs can fully utilize the advantages of SIMD parallel computing. However, when threads in a warp do not follow the same execution path, control divergence generates and affects the hardware utilization. In response to this problem, warp regrouping method has been proposed to combine threads executing the same branch path, which can significantly improve thread-level parallelism. But it is found that not all warps can be regrouped effectively because that may introduce a lot of unnecessary overheads, limiting further performance improvement. In this paper, we analyze the source of overheads and propose a lightweight warp regrouping method --- Partial Warp Regrouping (PWR) that controls the scope of reorganization and avoids most of the unnecessary warp regrouping by setting thresholds. In this method, it also can reduce the complexity of hardware design. Our experimental results show that this mechanism can improve the performance by 12% on average and up to 27% compared with immediate post-dominator.
YaoHua Yang, Shiqing Zhang, Li Shen 0007
HPC Asia2
2018 Merging and Evolution: Improving Convolutional Neural Networks for Mobile Applications
abstract
Compact neural networks are inclined to exploit “sparsely-connected” convolutions such as depthwise convolution and group convolution for employment in mobile applications. Compared with standard “fully-connected” convolutions, these convolutions are more computationally economical. However, “sparsely-connected” convolutions block the inter-group informa-tion exchange, which induces severe performance degradation. To address this issue, we present two novel operations named merging and evolution to leverage the inter-group information. Our key idea is encoding the inter-group information with a narrow feature map, then combining the generated features with the original network for better representation. Taking advantage of the proposed operations, we then introduce the Merging-and- Evolution (ME) module, an architectural unit specifically designed for compact networks. Finally, we propose a family of compact neural networks called MENet based on ME modules. Extensive experiments on ILSVRC 2012 dataset and PASCAL VOC 2007 dataset demonstrate that MENet consistently outperforms other state-of -the-art compact networks under different computational budgets. For instance, under the computational budget of 140 MFLOPs, MENet surpasses ShuffleNet by 1% and MobileNet by 1.95% on ILSVRC 2012 top-l accuracy, while by 2.3% and 4.1% on PASCAL VOC 2007 mAP, respectively.
Zheng Qin 0002, Zhaoning Zhang 0001, Shiqing Zhang, Hao Yu 0010, Jincai Li, Yuxing Peng 0001
IJCNN3
2018 GPU Memory Management Solution Supporting Incomplete Pages
Li Shen 0007, Shiqing Zhang, YaoHua Yang, Zhiying Wang 0003
NPC2
2018 Learning Affective Features With a Hybrid Deep Model for Audio-Visual Emotion Recognition
abstract
Emotion recognition is challenging due to the emotional gap between emotions and audio-visual features. Motivated by the powerful feature learning ability of deep neural networks, this paper proposes to bridge the emotional gap by using a hybrid deep model, which first produces audio-visual segment features with Convolutional Neural Networks (CNNs) and 3D-CNN, then fuses audio-visual segment features in a Deep Belief Networks (DBNs). The proposed method is trained in two stages. First, CNN and 3D-CNN models pre-trained on corresponding large-scale image and video classification tasks are fine-tuned on emotion recognition tasks to learn audio and visual segment features, respectively. Second, the outputs of CNN and 3D-CNN models are combined into a fusion network built with a DBN model. The fusion network is trained to jointly learn a discriminative audio-visual segment feature representation. After average-pooling segment features learned by DBN to form a fixed-length global video feature, a linear Support Vector Machine is used for video emotion classification. Experimental results on three public audio-visual emotional databases, including the acted RML database, the acted eNTERFACE05 database, and the spontaneous BAUM-1s database, demonstrate the promising performance of the proposed method. To the best of our knowledge, this is an early work fusing audio and visual cues with CNN, 3D-CNN, and DBN for audio-visual emotion recognition.
Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 Speech Emotion Recognition Using Deep Convolutional Neural Network and Discriminant Temporal Pyramid Matching
abstract
Speech emotion recognition is challenging because of the affective gap between the subjective emotions and low-level features. Integrating multilevel feature learning and model training, deep convolutional neural networks (DCNN) has exhibited remarkable success in bridging the semantic gap in visual tasks like image classification, object detection. This paper explores how to utilize a DCNN to bridge the affective gap in speech signals. To this end, we first extract three channels of log Mel-spectrograms (static, delta, and delta delta) similar to the red, green, blue (RGB) image representation as the DCNN input. Then, the AlexNet DCNN model pretrained on the large ImageNet dataset is employed to learn high-level feature representations on each segment divided from an utterance. The learned segment-level features are aggregated by a discriminant temporal pyramid matching (DTPM) strategy. DTPM combines temporal pyramid matching and optimal Lp-norm pooling to form a global utterance-level feature representation, followed by the linear support vector machines for emotion classification. Experimental results on four public datasets, that is, EMO-DB, RML, eNTERFACE05, and BAUM-1s, show the promising performance of our DCNN model and the DTPM strategy. Another interesting finding is that the DCNN model pretrained for image applications performs reasonably good in affective speech feature extraction. Further fine tuning on the target emotional speech datasets substantially promotes recognition performance.
Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001
IEEE Trans. Multim.1
2017 Energy efficient uplink transmission for UE-to network relay in heterogeneous networks
abstract
UE-to-Network relay leverages the proximity communication between user equipments (UEs) and allows certain UEs to provide relay assistance for others, which can greatly improve system energy efficiency. In this paper, we consider the scenario where UEs suffering from bad channel condition and low battery level can communicate with the base station directly or via the help of other UEs in heterogeneous networks. The optimal power allocation and connectivity among UEs are studied, which aims at minimizing the system transmission energy while guaranteeing the minimum data rate requirement of each UE. An optimization framework is presented to formulate the system transmission energy minimization problem, which can be converted into a weighted one-to-one matching problem. And a practical joint transmission mode with relay selection and power allocation (JMRP) algorithm is developed to solve it. Simulation results show that the proposed algorithm outperforms the existing works in terms of system transmission energy and throughput.
Shiqing Zhang, Xiaodong Xu 0001, Mengying Sun, Xiaoxuan Tang, Xiaofeng Tao 0001
PIMRC1
2016 Multimodal Deep Convolutional Neural Network for Audio-Visual Emotion Recognition
abstract
Emotion recognition is a challenging task because of the emotional gap between subjective emotion and the low-level audio-visual features. Inspired by the recent success of deep learning in bridging the semantic gap, this paper proposes to bridge the emotional gap based on a multimodal Deep Convolution Neural Network (DCNN), which fuses the audio and visual cues in a deep model. This multimodal DCNN is trained with two stages. First, two DCNN models pre-trained on large-scale image data are fine-tuned to perform audio and visual emotion recognition tasks respectively on the corresponding labeled speech and face data. Second, the outputs of these two DCNNs are integrated in a fusion network constructed by a number of fully-connected layers. The fusion network is trained to obtain a joint audio-visual feature representation for emotion recognition. Experimental results on the RML audio-visual database demonstrates the promising performance of the proposed method. To the best of our knowledge, this is an early work fusing audio and visual cues in DCNN for emotion recognition. Its success guarantees further research in this direction.
Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001
ICMR1
2016 Graph Regularized Sparsity Discriminant Analysis for face recognition
Songjiang Lou, Xiaoming Zhao 0002, Yuelong Chuang, Shiqing Zhang
Neurocomputing5
2015 Spoken emotion recognition via locality-constrained kernel sparse representation
Xiaoming Zhao 0002, Shiqing Zhang
Neural Comput. Appl.2
2014 Locality-sensitive kernel sparse representation classification for face recognition
Shiqing Zhang, Xiaoming Zhao 0002
J. Vis. Commun. Image Represent.1
2014 Robust emotion recognition in noisy speech via sparse representation
Xiaoming Zhao 0002, Shiqing Zhang, Bicheng Lei
Neural Comput. Appl.2
2013 Dimensionality reduction-based spoken emotion recognition
Shiqing Zhang, Xiaoming Zhao 0002
Multim. Tools Appl.1
2012 Phoneme recognition using an adaptive supervised manifold learning algorithm
Xiaoming Zhao 0002, Shiqing Zhang
Neural Comput. Appl.2
2008 Emotion Recognition in Chinese Natural Speech by Combining Prosody and Voice Quality Features
Shiqing Zhang
ISNN (2)1
2002 Tight upper bound on the number of edges in a bipartite K3, 3-free or K5-free graph with an application
Zhi-Zhong Chen, Shiqing Zhang
Inf. Process. Lett.2
1998 Performing Analysis for Dynamic Tree Embedding in k-Partite Networks by a Random Walk
Hong Shen 0001, Keqin Li 0001, Yi Pan 0001, Gilbert H. Young, Shiqing Zhang
J. Parallel Distributed Comput.5