VLDB 2026 Research / reviewers in the wild / expert
Suhong Wang
dblp:19/6246
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Visually multimodal depression assessment based on key questions with weighted multi-task learning
Miaomiao Cao, Xianlin Zhu, Suhong Wang |
Signal Process. Image Commun. | 4 |
| 2025 | A High-Reliability Small-Area Task Offloading Mechanism With Trust Evaluation and Fuzzy Logic in Power IoTsabstractIn order to solve the problem that high-priority tasks can not be processed timely and reliably due to the disorder of multi-task and dynamicity in Power Internet of Things(PIoTs), a high-reliability small-area task offloading mechanism with trust evaluation and fuzzy logic(HRSATF) is proposed. First, considering task priority, preemptive priority queue is introduced to ensure high-priority tasks processed preferentially, and minimum resource allocation coefficients(MRACs) of tasks are solved to ensure the effectiveness of offloading. Second, the trust model between smart device(SD) and edge server(ES) is established, and ESs are divided into three priorities based on trust value and computing power by fast non-dominated sorting. Thirdly, fuzzy logic is applied to select target ES when the priorities of task and ES do not match or the ES is offline, and MRAC is used to schedule tasks between SD and ES. Finally, NSGA2 is modified (MNSGA2) to verify the effectiveness of HRSATF in terms of success rate, time, power consumption and load balancing, where success rate is increased by$102.3\%$, and time and power consumption are decreased by$90.7\%$,$89.3\%$at most, respectively. Suhong Wang, Tuanfa Qin, Yongle Hu, Hongmin Sun |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | MMPF: Multimodal Purification Fusion for Automatic Depression DetectionabstractDepression is a common mental disorder that requires objective and valid assessment tools. However, purely data-driven methods cannot satisfy the clinical diagnostic criteria for automatic depression detection (ADD), and the instability and heterogeneity of multimodal data have not been fully resolved. Therefore, we propose a novel auxiliary tool for ADD based on multimodal purification fusion (MMPF). Initially, a prior constraint gating (PCG) strategy is used to inject doctors’ constraints into depression data to guide and constrain the learning process. Then, we introduce text and audio encoders to extract unpurified features from preprocessed depression data. Afterward, multimodal purification refinement is proposed to extract unintersected common and specific features from unpurified features, generating purified features. Meanwhile, we leverage a multiperspective contrastive learning (MCL) strategy to enhance unpurified and purified features. Finally, modality interaction (MI) based on the transformer is proposed to conduct multimodal fusion. A dynamic corrective learning (DCL) strategy is introduced to tackle modality imbalances and inconsistent sentiment. MMPF is evaluated on the Distress Analysis Interview Corpus Wizard of Oz and performs promisingly in unimodal and multimodal depression detection, indicating its significant role in ADD. Miaomiao Cao, Xianlin Zhu, Suhong Wang, Xiaofeng Liu 0006 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Uncertainty-Aware Label Contrastive Distribution Learning for Automatic Depression DetectionabstractDepression is one of the most common mental illnesses, affecting people’s quality of life and posing a risk to their health. Low-cost and objective automatic depression detection (ADD) is becoming increasingly critical. However, existing ADD methods usually treat depression detection as a regression problem for predicting patient health questionnaire-8 (PHQ-8) scores, ignoring the scores’ ambiguity caused by multiple questionnaire issues. To effectively leverage the score labels, we propose an uncertainty-aware label contrastive and distribution learning (ULCDL) method to estimate PHQ-8 scores, thus detecting depression automatically. ULCDL first simulates the ambiguity within PHQ-8 scores by converting single-valued scores into discrete label distributions. Afterward, it learns to predict the PHQ-8 score distribution by minimizing the Kullback–Leibler (KL) divergence between the score distribution and the discrete label distribution. Finally, the predicted PHQ-8 score distribution outperforms the PHQ-8 score in ADD. Moreover, label-based contrastive learning (LBCL) is introduced to facilitate the model for learning common features related to depression in multimodal data. A multibranch fusion module is proposed to align and fuse multimodal data for better exploring the uncertainty of PHQ-8 labels. The proposed method is evaluated on the publicly available DAIC-WOZ dataset. Experiment results show that ULCDL outperforms regression-based depression detection methods and achieves state-of-the-art performance. The code will be released after acceptance. Miaomiao Cao, Xianlin Zhu, Suhong Wang |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2023 | Towards Next Generation Video Coding: from Neural Network Based Predictive Coding to In-Loop FilteringabstractAudio Video Coding Standard (AVS) Intelligent Coding Group mainly studies video coding tools based on neural network technology and its potential benefit for next generation video coding. Extensive efforts have been dedicated to the research on neural network (NN) based coding tools. In this paper, we present a novel NN based video coding framework by leveraging the supervised trained NN models for multiple modules in the hybrid coding framework, from the predictive coding to the in-loop filtering. Specifically, NN based intra prediction models the non-linear mapping from contextual pixels to the predictions. The inter prediction efficiency is enhanced by introducing a virtual reference frame (VRF) network. The convolutional neural network based loop filtering (CNNLF) with discriminative model selection exploits the texture adaptivity. The experimental results show that the CNNLF, NN Intra, and VRF models can bring 8.60%, 1.02%, and 2.26% luma BD-rate reduction under random access (RA) configuration compared with AVS reference software HPM13.0. Additional experiments with the combined three NN coding tools reveal that around 13% YUV BD-rate reduction could be obtained. The proposed framework opens novel sights for next generation video coding from the intelligent coding perspective. Yanchen Zhao, Suhong Wang, Meng Lei, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001 |
ISCAS | 2 |
| 2023 | Swin-T-NFC CRFs: An encoder-decoder neural model for high-precision UAV positioning via point cloud super resolution and image semantic segmentation
Suhong Wang, Hongqing Wang, Shufeng She, Yanping Zhang 0002, Qingju Qiu, Zhifeng Xiao |
Comput. Commun. | 1 |
| 2023 | Multi-layer task scheduling and resource allocation schemes considering idle resource and task priority in IoT networksabstractAbstract With more and more interconnected smart devices (ISDs) accessing the Internet of Things (IoT), massive and diverse tasks need to be transformed and computed. Mobile edge computing enables the offloading of tasks to nearby servers to enhance processing efficiency, which makes ISDs idle, causing resource waste and failing to satisfy the high real‐time requirements of tasks. Besides, when tasks with different priorities are processed in the order they are generated, it will be difficult for IoT to guarantee a timely response to high‐priority tasks. To address the aforementioned issues, we establish an edge‐terminal‐local architecture by software‐defined networking to centrally manage idle ISD resource (2ISDR). Then the proposed two‐step scheduling mechanism with preemptive priority queue ensures the real‐time responses to high‐priority tasks, and the minimum resource allocation coefficients make offloading effective. Finally, we also propose a modified NSGA‐III algorithm named MNSGA‐III, which is designed to make decisions about offloading and solve resource allocation for tasks, and we correct infeasible solutions by a two‐step correction function to ensure the feasibility of MNSGA‐III. Experimental results show that the method can ensure a timely response to high‐priority tasks and optimize processing time, energy consumption, and economic cost through the utilization of 2ISDR. Suhong Wang, Hongmin Sun, Junyu Ren, Yongle Hu, Tuanfa Qin |
IET Commun. | 1 |
| 2023 | Textural and Directional Information Based Offset In-Loop Filtering in AVS3abstractIn this paper, we propose a novel low-complexity in-loop filtering approach named textural and directional information based offset (TDIO) for the video coding standard AVS3. Different from conventional offset-based filtering methods which partially use contextual samples, the key contribution of TDIO is that it fully utilizes the textural and edge directional features of each sample to comprehensively determine which type of texture characteristics each sample belongs to. The corresponding offsets are generated and signaled to decoder such that sample-level distortion is reduced. Specifically, the multi-directionality and sample-intensity pattern based classifiers are first proposed to extract the directional and textural features, respectively. The classification results are obtained by incorporating these features, and the optimal offset values for each class are derived based on rate-distortion optimization. Since sample-level offset signalling may cause heavy burden to the overhead of TDIO, we subsequently propose a filtering offset sharing mechanism based on historical information between available temporal-adjacent compressed frames. In addition, an iteration-based filter adaptation method is designed to improve the local adaptivity of TDIO for better compression efficiency. Experimental results show that the proposed TDIO achieves 0.64%, 1.29%, 1.86%, and 2.20% bit rate savings for all intra, random access, low delay B, and low delay P configurations, respectively. Moreover, TDIO is helpful to improve subjective quality by leveraging the fine-grained local texture characteristics. It can be observed that the blurring and ringing artifacts could be significantly suppressed by using the proposed method, yielding higher subjective quality. Jiaqi Zhang 0007, Yunrui Jian, Suhong Wang, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | A Pixel-Level Segmentation-Synthesis Framework for Dynamic Texture Video CompressionabstractDynamic textures are time-varying motion patterns that exhibit certain temporal stationarity. The motion of such patterns is usually anisotropic, which is a great challenge for video coding. In this paper, we firstly analyze the predictive and temporal characteristics of dynamic textures and, based on the analysis, a pixel-level segmentation-synthesis framework for dynamic texture compression is proposed to improve predictive-coding efficiency. This framework consists of three sub-modules: dynamic texture segmentation, dynamic texture synthesis and fusion process. A deep-learning-based segmentation method is introduced to obtain a pixel-level mask. Subsequently, optical flow fields are used to generate flowlines and help analyze the motion periodicity of a given sequence. The flowlines and period values are utilized as inputs for flow-based dynamic texture synthesis. The segmented dynamic texture area is then reconstructed using the synthesized results at the decoder side at pixel-level, such that the overhead for transmitting these contents could be saved. The segmented mask is transmitted through an All-Zero-Run-Length coding algorithm and an effective fusion module is further proposed to reduce edge artifacts that occur at the boundary of dynamic texture areas. Coding blocks have been classified into three types and a weighting reconstruction scheme is performed. The proposed framework has been integrated into VTM-10.0. Experimental results on a collected dynamic texture test set demonstrate that 15.1% and 16.1% MOS BD-rate savings can be achieved on average for LDB and LDP configurations, respectively. Suhong Wang, Chuanmin Jia, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Quad-Treea Based Sample Refinement Filter for Video CodingabstractIn-loop filter is a crucial module in video coding, which can improve both subjective and object quality of reconstructed videos. In this paper, a new sample-based classification method is first proposed using features extracted from different stages of the existing in-loop filter process. Based on this method, an adaptive three-layer Quad-tree Based Sample Refinement Filter (QSRF) algorithm is designed to further improve the coding efficiency. Experimental results show that the proposed QSRF algorithm achieves 0.39%, 0.77% and 0.70% BD-rate savings for random access, lowdelay B and lowdelay P configurations compared to AVS3 reference software, respectively. Moreover, the proposed method can also improve visual quality of reconstructed videos significantly. Yunrui Jian, Jiaqi Zhang 0007, Chuanmin Jia, Suhong Wang, Shanshe Wang, Siwei Ma 0001 |
DCC | 4 |
| 2021 | Flow-Grounded Dynamic Texture Synthesis for Video CompressionabstractThe basic ingredients of modern video coding standards are block-based prediction and transforms. However, when dealing with video contents containing dynamic textures (DT), the existing prediction schemes usually failed due to temporal variability and randomness of DT, which results in more bit cost on residual coding compared with other contents. In view of this point, a novel video compression scheme for DT is proposed in this work. In particular, wavelet-based analysis on motion characteristics of DT is firstly presented and based on the analysis, we introduce a flow-grounded texture synthesis method for video compression. Instead of conventional inter prediction, synthesized DT contents are used for reconstruction at the decoder. The proposed scheme has been fully integrated into the test model of Versatile Video Coding standard, VTM-10.0, for validation and a subjective test has also been carried out. Experimental results show that bitrate savings can be achieved by 40% on average at comparable visual quality. Suhong Wang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 1 |
| 2021 | Enhanced Cross Component Sample Adaptive Offset for AVS3abstractCross-component prediction has great potential for removing the redundancy of multi-components. Recently, cross-component sample adaptive offset (CCSAO) was adopted in the third generation of Audio Video coding Standard (AVS3), which utilizes the intensities of co-located luma samples to determine the offsets of chroma sample filters. However, the frame-level based offset is rough for various content, and the edge information of classified samples is ignored. In this paper, we propose an enhanced CCSAO (ECCSAO) method to further improve the coding performance. Firstly, four selectable 1-D directional patterns are added to make the mapping between luma and chroma components more effectively. Secondly, one four-layer quad-tree based structure is designed to improve the filtering flexibility of CCSAO. Experimental results show that the proposed approach achieves 1.51%, 2.33% and 2.68% BD-rate savings for All-Intra (AI), Random-Access (RA) and Low Delay B (LD) configurations compared to AVS3 reference software, respectively. A subset improvement of ECCSAO has been adopted by AVS3. Yunrui Jian, Jiaqi Zhang 0007, Suhong Wang, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 4 |
| 2019 | Adaptive Wavelet Domain Filter for Versatile Video Coding (VVC)abstractOwing to the ability of removing compression artifacts, extensive in-loop filters have been proposed for video coding standards. They are performed after the reconstruction of all coding units (CUs), however, none of them has been taken into account in the mode decision when coding each CU. To address this issue and make the rate-distortion optimization (RDO) more precise for each CU, we introduce a low-pass filter when checking the rate-distortion cost after the reconstruction of each CU. Specifically, based on Haar wavelet, the reconstructed block is transformed to the frequency domain, and then an adaptive wavelet domain filter (AWF) is proposed to suppress the quantization noises in coded blocks. To be adaptive, the filter strength varies from CU to CU according to the texture complexity and quantization parameters (QPs). Experimental results show that the proposed method can reduce the compression artifacts and improve both the objective and subjective quality. Suhong Wang, Xiang Zhang 0004, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
DCC | 1 |