Wei Song 0007

dblp:62/1539-7 · DBLP profile ↗
← Back
49ranked-venue papers
12as first author
32since 2021 · last 2026
0000-0002-0604-5563ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 9 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A noise robust and distribution-adaptive framework for multivariate time series anomaly detection
Yanling Du, Ziliang Yang, Baozeng Chang, Jingxia Gao, Xiaojia Bao, Wei Song 0007
Neural Networks6
2025 A Quality Control Method for Ocean Data Based on Multi-time Scale Downsampling and Dynamic Threshold Strategy
abstract
Ocean data quality control (QC) is important for ocean scientific research and resource management, which aims to identify data errors and ensure the reliability of ocean data. Modern ocean observation data are characterized by large volumes and complex association patterns. This results in low efficiency of manual QC and low accuracy of automatic QC. Therefore, the use of unsupervised deep learning models for ocean data QC has become a major trend in recent years. However, it faces some challenges: the powerful learning ability of deep learning models could learn too well from anomalous data; the determination of anomaly threshold has an important effect on the results of QC. To address these issues, we propose a QC method for ocean data based multi-time scale downsampling and dynamic threshold strategy (MTSDTS-QC). A Multi-time Scale Downsampling module is developed to extract the distribution of normal data at both large and small time scales. This process can reduce the interference of anomalous data in the subsequent process. Then, TimesBlocks are used to reconstruct the ocean data considering their multi-periodic characteristics. The difference between the original and the reconstructed data represents the degree of abnormality of the data. Furthermore, we propose a dynamic threshold strategy that can determine the optimal thresholds for different marine data based on the intrinsic characteristics of anomaly scores. Experimental results on marine meteorological data show that the MTSDTS-QC method outperforms 14 existing baseline methods. The generalizability of our model is also validated on a public industrial control dataset WADI.
Yuhui Lu, Wei Song 0007, Qi He 0003, Yanling Du
IJCNN2
2025 MSFA-RSISR: Multi-scale and Fourier Attention for Remote Sensing Image Super-Resolution
Zexin Xie, Jian Wang 0130, Yanling Du, Wei Song 0007
PRCV (8)5
2025 Breaking the Label Barrier: Underwater Semi-supervised Object Detection with Improved FPN and Adaptive Thresholds
Nana Huang, Qi He 0003, Wei Song 0007, Yinjiang Zhang, Haibin Mei
PRCV (17)4
2025 Observe finer to select better: Learning key frame extraction via semantic coherence for dynamic facial expression recognition in the wild
Shaoqi Yan, Yan Wang 0068, Xinji Mai, Zeng Tao, Wei Song 0007, Qing Zhao 0007, Boyang Wang 0003, Haoran Wang 0006, Shuyong Gao
Inf. Sci.5
2025 Cross-view self-supervised heterogeneous graph representation learning
Danfeng Zhao, Yanhao Chen 0003, Wei Song 0007, Qi He 0003
Neural Networks3
2025 PUTrack: Improved Underwater Object Tracking via Progressive Prompting
abstract
Existing underwater tracking methods can be categorized into two paradigms: first, “enhance-then-track”—first enhancing the quality of the input image, then employing an open-air tracker; second, “track-then-process”—initially using an open-air tracker, followed by calibrating the prediction box. These methods that lack unified objectives among the modules impair tracking performance. To overcome this, we propose a novel end-to-end framework called prompting underwater tracking (PUTrack). It adapts the open-air tracker to the scenario-specific (underwater) tracking task by deploying a set of underwater prompters at the lateral side of an existing open-air tracker, and injecting the generated prompts layer-by-layer into the encoder. Experiments on various underwater tracking datasets demonstrate that the method significantly improves the underwater performance of the tracker by introducing only 0.6 M trainable parameters (0.4% of total parameters). Moreover, to drive the development of underwater tracking, we construct a high quality underwater tracking dataset that is manually annotated with 139 k frames, which exceeds the total number of frames in previous underwater tracking datasets. It provides 90 test sets rich in challenge properties and 200 training sets of diverse kinds.
Qiuyang Zhang 0002, Wei Song 0007
IEEE Trans. Ind. Informatics2
2024 Global Feature Attribution Map based on Optical Flow for Super Resolution Neural Networks
abstract
Research in image super-resolution (SR), which seeks to enhance image quality by producing higher-resolution versions from low-quality inputs, has primarily focused on developing reconstruction algorithms rather than enhancing interpretability. SR networks continue to exhibit the opaque, black-box characteristics typical of deep learning, with limited research dedicated to investigating their internal mechanisms. This study aims to deepen the understanding of SR and conduct attribution analysis of SR networks from a holistic reconstruction perspective. We introduce a novel attribution method based on gradient and optical flow, termed GOFlow. Following verification with five SR models, we demonstrate that (1) GOFlow proves to be an effective tool for analyzing attributed pixels in SR neural networks from a comprehensive perspective; (2) Compared to the Layer Attribution Method (LAM), GOFlow produces a more detailed attribution map from a global perspective; (3) GOFlow is capable of exploring how texture and colorfulness influence the outputs of SR, proving that GOFlow can interpret SR models at the feature level.
Alexander W. Jacob, Wei Song 0007, Antonio Liotta
BDCAT3
2024 Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution
Junxiong Lin, Yan Wang 0068, Zeng Tao, Boyang Wang 0003, Qing Zhao 0007, Haorang Wang, Xuan Tong, Xinji Mai, Yuxuan Lin 0001, Wei Song 0007, Jiawen Yu, Shaoqi Yan
ECCV (52)10
2024 Unsupervised Underwater Image Enhancement Combining Imaging Restoration and Prompt Learning
Wei Song 0007, Chengbing Liu, Mario Di Mauro, Antonio Liotta
PRCV (2)1
2024 EHAT:Enhanced Hybrid Attention Transformer for Remote Sensing Image Super-Resolution
Jian Wang 0130, Zexin Xie, Yanlin Du, Wei Song 0007
PRCV (8)4
2024 Hybrid learning strategies for multivariate time series forecasting of network quality metrics
abstract
This work addresses the challenge of forecasting temporal metrics that characterize cellular traffic behavior. The ultimate goal is to provide network operators with a valuable tool for modeling mobile network traffic and optimizing connected resources. The idea is to estimate beforehand the temporal evolution of some Quality-of-Experience (QoE) and Quality-of-Service (QoS) metrics, which is helpful for accurately tuning the allocation of network resources. Remarkably, these metrics (expressed as time series) are typically correlated, and changes in one time series can affect others in a variety of ways and to different extents. For example, high network delay (a QoS-related metric) is associated with degradation in voice quality over time (a QoE-related metric). Accordingly, we address the problem of cellular traffic forecasting with correlated time series, proposing three innovative hybrid learning strategies designed by combining the advantages of two approaches: (i) a statistical approach, implemented through the Vector Autoregressive (VAR) model, which encodes each metric as a combination of past values of the same metric along with a combination of values of other related metrics, resulting in a multivariate structure; and (ii) an approach based on deep learning techniques (specifically, CNN, LSTM, and GRU) which operate on such a multivariate structure to perform the forecasting. The resulting performance demonstrates the benefits of the proposed hybrid schemes (VAR-CNN, VAR-LSTM, VAR-GRU) over their pure counterparts, with a significant reduction in forecasting errors. The network metrics were gathered in a real urban cellular environment, where the presence of exogenous factors (e.g., interferences, weather conditions, etc.) makes the forecasting assessment particularly challenging.
Mario Di Mauro, Giovanni Galatro, Fabio Postiglione, Wei Song 0007, Antonio Liotta
Comput. Networks4
2024 Empower smart cities with sampling-wise dynamic facial expression recognition via frame-sequence contrastive learning
Shaoqi Yan, Yan Wang 0068, Xinji Mai, Qing Zhao 0007, Wei Song 0007, Zeng Tao, Haoran Wang 0006, Shuyong Gao
Comput. Commun.5
2024 Mixed noise-guided mutual constraint framework for unsupervised anomaly detection in smart industries
Qing Zhao 0007, Yan Wang 0068, Yuxuan Lin 0001, Shaoqi Yan, Wei Song 0007, Boyang Wang 0003, Yang Chang, Lizhe Qi
Comput. Commun.5
2024 A hierarchical probabilistic underwater image enhancement model with reinforcement tuning
Wei Song 0007, Yan Wang 0068, Antonio Liotta
J. Vis. Commun. Image Represent.1
2024 DCENet: A Dense Contextual Ensemble Network for Multiclass Ocean Front Detection
abstract
Ocean fronts are significant mesoscale phenomena in oceanography. Ocean front detection has important impacts on fishery, environmental protection, and other fields. However, in multiclass ocean front detection, existing deep learning models struggle to detect small target fronts (STFs) in remote sensing images effectively. To achieve accurate STF detection, we model multiclass ocean front detection as a semantic segmentation problem and propose the dense contextual ensemble network (DCENet). Specifically, DCENet adopts a novel spatially enhanced contextual ensemble (S-CE) architecture based on encoder-decoder to aggregate rich multiscale context information. The dense aggregation block (DA Block) introduced in the encoder and decoder cascades all convolutional features, allowing the deep layers of the network for precise positioning to use STF detailed information. In addition, the designed hybrid loss consisting of balanced cross-entropy (BCE) loss and Dice loss can balance the losses among the front classes and make the model fit the distribution of fronts. Experimental results on the SCSOF dataset show that DCENet can accurately detect various classes of ocean fronts, with mean intersection over union (mIoU) and mean$F1$-score (mF1) reaching 78.97% and 88%, respectively. DCENet significantly outperforms other representative semantic segmentation models in STF detection.
Qi He 0003, Bo Gong 0006, Wei Song 0007, Yanling Du, Danfeng Zhao, Wenbo Zhang 0004
IEEE Geosci. Remote. Sens. Lett.3
2024 PromptVT: Prompting for Efficient and Accurate Visual Tracking
abstract
While existing lightweight visual trackers can run in real-time at edge devices, they face the difficulty of object appearance changes. An effective solution to this problem is to add an online updatable dynamic template for trackers to learn about changes in target appearance over time. However, existing dynamic template utilization methods are unsuitable for lightweight networks, resulting in limited accuracy improvement and a significant increase in computational workload. In this paper, we propose PromptVT, an efficient and accurate video tracking framework, which consists of two important designs: a plug-and-play dynamic template prompter (DTP) and a hierarchical multi-scale transformer (HMT). The DTP module guides networks to effectively learn changes between initial and dynamic templates through two prompts without additional computational workload. The HMT module combines spatial features of the search area and template at different scales and levels, enabling the tracker to learn a more comprehensive visual representation. Our proposed PromptVT outperforms state-of-the-art real-time trackers on eight benchmarks (VOT2020, LaSOT, GOT-10K, UAV123, AntiUAV, AntiUAV410, TrackingNet, OTB100) while running at 52fps(PyTorch model) and 76fps(ONNX model) on CPUs, with only 2.9G FLOPs and 3M parameters. Code and models are available at https://github.com/faicaiwawa/PromptVT.
Qiuyang Zhang 0002, Wei Song 0007, Dongmei Huang 0001, Qi He 0003
IEEE Trans. Circuits Syst. Video Technol.3
2024 MGR3Net: Multigranularity Region Relation Representation Network for Facial Expression Recognition in Affective Robots
abstract
Automatic facial expression recognition (FER) based on face images is essential for affective robots, which are designed for interactive companions and intelligent healthcare. Although existing DL-based FERs have made significant progress, an accurate FER model in robots is challenging due to the subtle differences in facial expressions across various scenarios. To address this issue, we propose a multigranularity region relation representation network (MGR3Net) to improve the robustness and generalization of FER via attention-guided global-local fusion. The MGR3Net is composed of three modules: multigranularity attention (MGA), holistic-regional feature extractor (HRFE), and hybrid feature fusion. In the MGA module, we first process each holistic cropped face image into three granularity of face regions from coarse to fine, which are four region-cropped faces,$2^{2}$face partitions, and$4^{2}$face partitions. Then, we propose the region attention relation cell to model the relationship between each region and the aggregated representation while preserving the spatial information of the local features. In the HRFE module, we align multigranularity features from the coarse space to the finer space and extract one holistic embedding and multiple region embeddings for each granularity. Finally, we use a hybrid-level fusion strategy to combine global-local features from the three granularities for final classification. Extensive experiments demonstrate that the MGR3Net outperforms the state-of-the-art methods evaluated on the in-the-lab datasets, in-the-wild datasets, and occlusion/pose-based sets.
Yan Wang 0068, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jing Liu 0050, Dingkang Yang, Shuyong Gao
IEEE Trans. Ind. Informatics3
2024 MSC-AD: A Multiscene Unsupervised Anomaly Detection Dataset for Small Defect Detection of Casting Surface
abstract
Intelligent detection of product surface defects in the industrial scene is the key to ensuring product quality. On general benchmarks, current unsupervised anomaly detection techniques have achieved significant success. When used in complex industrial environments (e.g., large industrial components with small defects), the model needs to be able to adapt to different imaging scenarios (e.g., illumination and resolution) and accurately detect and localize anomalies, but its performance is still far from satisfactory. Besides, the complex and unstable optical lighting environment for collecting such data poses major challenges in establishing unified benchmarks for optical lighting and imaging resolution in defect detection. To fill this gap, we build a standard imaging system-based multiscene unsupervised anomaly detection dataset, coined as MSC-AD. In particular, it provides 12 imaging scenes, i.e., a cross combination of low-to-high three illuminations and 150 × 150 to 600 × 600 four resolutions, in which six types of large casting surfaces with different structures include five kinds of small defects with sample-level and pixel-level precise ground truth. We systematically investigate representative baseline methods and empirical analysis on this dataset to obtain a number of interesting findings, e.g., how to detach from distinctly different imaging scenes, and how to distinguish between subtly normal–anomaly classes. To the best of our knowledge, MSC-AD is the first multi-illumination, multiresolution, multisurface, and multidefect dataset built in a standard imaging system.
Qing Zhao 0007, Yan Wang 0068, Boyang Wang 0003, Junxiong Lin, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jiawen Yu, Shuyong Gao
IEEE Trans. Ind. Informatics6
2024 Multivariate Time Series Characterization and Forecasting of VoIP Traffic in Real Mobile Networks
abstract
Predicting the behavior of real-time traffic (e.g., VoIP) in mobility scenarios could help the operators to better plan their network infrastructures and to optimize the allocation of resources. Accordingly, in this work the authors propose a forecasting analysis of crucial QoS/QoE descriptors (some of which neglected in the technical literature) of VoIP traffic in a real mobile environment. The problem is formulated in terms of a multivariate time series analysis. Such a formalization allows to discover and model the temporal relationships among various descriptors and to forecast their behaviors for future periods. Techniques such as Vector Autoregressive models and machine learning (deep-based and tree-based) approaches are employed and compared in terms of performance and time complexity, by reframing the multivariate time series problem into a supervised learning one. Moreover, a series of auxiliary analyses (stationarity, orthogonal impulse responses, etc.) are performed to discover the analytical structure of the time series and to provide deep insights about their relationships. The whole theoretical analysis has an experimental counterpart since a set of trials across a real-world LTE-Advanced environment has been performed to collect, post-process and analyze about 600,000 voice packets, organized per flow and differentiated per codec.
Mario Di Mauro, Giovanni Galatro, Fabio Postiglione, Wei Song 0007, Antonio Liotta
IEEE Trans. Netw. Serv. Manag.4
2023 Few-Shot NER in Marine Ecology Using Deep Learning
Jian Wang 0130, Danfeng Zhao, Wei Song 0007
ICONIP (11)5
2023 A systematic review and analysis of deep learning-based underwater object detection
Shubo Xu, Wei Song 0007, Haibin Mei, Qi He 0003, Antonio Liotta
Neurocomputing3
2023 Ranked Similarity Weighting and Top-nk Sampling in Deep Metric Learning
abstract
Deep metric learning has been widely used in many visual tasks. Its key idea is to increase the similarity of positive samples and decrease the similarity of negative samples through network training. To achieve this purpose, many studies excessively extend the distance between the query sample and hard negative samples. This may compress the distance between similar samples of other classes, causing these samples to cluster together. We call this phenomenon Negative Sample Aggregation. To address this problem, first, we propose a weighting method based on the Ranking Similarity of sample pairs, short for RS. The proposed weighting method can not only enlarge the distance between the query sample and hard negative samples, but also maintain the embedding distribution of proximal negative samples. Second, we propose a Top-nk sampling method, which can dynamically adjust the sampling strategy according to the distribution of a dataset. It solves the problem that the descent direction of the network gradient is inconsistent with the optimization target. The effectiveness of our methods is evaluated by extensive experiments on four public datasets and compared with that of other state-of-the-art methods. The results show that the proposed method obtains excellent performance, reaching 67.8% on CUB-200-2011 and 85.2% on Cars-196 at Recall@1.
Jian Wang 0130, Xinyue Li 0001, Wei Song 0007, Weiqi Guo
IEEE Trans. Multim.4
2022 Multi-Hierarchy Proxy Structure for Deep Metric Learning
abstract
Mainstream methods for deep metric learning can be divided into pair-based and proxy-based methods. In recent years, proxy-based methods have attracted wide attention for their low training complexity and fast network convergence. Most proxy-based studies assign only one proxy per class to capture the features of the class, this leads to ignoring the hidden hierarchy and regular aggregation of features within the class. However, these details are meaningful for capturing features of the class. Therefore, we propose a multi-hierarchy proxy (MHP) structure to extract the hierarchical details and regular features hidden in the embedding space. At the same time, we design a layerwise merging similarity operator to reasonably measure the similarity between samples and classes. Our MHP method maintains the low time complexity of the proxy-based method and can be easily integrated into existing proxy-based losses. The effectiveness of our method is evaluated by extensive experiments on three public datasets and compared with state-of-the-art methods. The results show that the proposed MHP method can significantly improve the performance of proxy-based methods, reaching 69.8% on CUB-2002011 and 87.4% on Cars-196 dataset at Recall@1.
Jian Wang 0130, Xinyue Li 0001, Wei Song 0007, Weiqi Guo
ICASSP3
2022 DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in Videos
abstract
Current works of facial expression learning in video consume significant computational resources to learn spatial channel feature representations and temporal relationships. To mitigate this issue, we propose a Dual Path multi-excitation Collaborative Network (DPCNet) to learn the critical information for facial expression representation from fewer keyframes in videos. Specifically, the DPCNet learns the important regions and keyframes from a tuple of four view-grouped frames by multi-excitation modules and produces dual-path representations of one video with consistency under two regularization strategies. A spatial-frame excitation module and a channel-temporal aggregation module are introduced consecutively to learn spatial-frame representation and generate complementary channel-temporal aggregation, respectively. Moreover, we design a multi-frame regularization loss to enforce the representation of multiple frames in the dual view to be semantically coherent. To obtain consistent prediction probabilities from the dual path, we further propose a dual path regularization loss, aiming to minimize the divergence between the distributions of two-path embeddings. Extensive experiments and ablation studies show that the DPCNet can significantly improve the performance of video-based FER and achieve state-of-the-art results on the large-scale DFEW dataset.
Yan Wang 0068, Yixuan Sun, Wei Song 0007, Shuyong Gao, Zhaoyu Chen 0001, Weifeng Ge
ACM Multimedia3
2022 Efficient Subjective Video Quality Assessment Based on Active Learning and Clustering
Wei Song 0007, Wenbo Zhang 0004, Mario Di Mauro, Antonio Liotta
MoMM2
2022 Regularized lattice Boltzmann method parallel model on heterogeneous platforms
abstract
Abstract As an improved method of lattice Boltzmann method (LBM), regularized lattice Boltzmann method (RLBM) has been applied to simulate fluid flow. Nevertheless, the performance of RLBM needs to be considered when simulating actual problems. The rise of multicore platforms, especially the popularity of graphics processor units (GPUs), has provided possible implementation solutions for parallel computing. In this article, an RLBM parallel model on the CPU/GPU heterogeneous platforms is proposed. To solve the problem of possible GPU memory shortage, the CPU controls the startup of the kernel function in the RLBM algorithm and participates in the calculation. Due to the characteristics of the algorithm, the entire flow field is divided into CPU computing areas and GPU computing areas according to the given flow field division rules. OpenMP and CUDA are, respectively, applied to CPU and GPU for parallel computing. The startup of the kernel function takes a short time, and the CPU and GPU can be approximately regarded as performing calculations simultaneously. Since the data exchange between the CPU and GPU has a significant impact on performance, a buffer is set at the boundary of the CPU and GPU to reduce the frequency of data exchange. The buffer size determines the number of program iterations before exchanging data. The performance of the algorithm is measured by MFLUPS, and the algorithm is applied to the 3D lid‐driven cavity flow. The obtained results show that the MFLUPS of the algorithm is 25 times that of CPU MFLUPS, and it adds an increment equivalent to CPU MFLUPS compared with GPU MFLUPS. And the algorithm is extended to a multi‐GPU version, which is also applied to the 3D lid‐driven cavity flow. The results obtained show that the MFLUPS of the multi‐GPU version is 1.8 times that of the single GPU.
Zhixiang Liu, Wei Song 0007
Concurr. Comput. Pract. Exp.3
2022 Approximation algorithms for the min-max clustered k-traveling salesmen problems
Xiaoguang Bao, Wei Yu 0011, Wei Song 0007
Theor. Comput. Sci.4
2021 A Ranked Similarity Loss Function with pair Weighting for Deep Metric Learning
abstract
Metric learning is a widely-used method for image retrieval. The object of metric learning is to limit the distance between similar samples and increase the distance between samples of different classes through learning. Many studies tend to pay more attention to keep the distance between positive and negative samples, but ignore the distance between different classes of negative samples. In fact, query samples should be separated from negative samples of different classes by different distances. To address these problems, we propose to build a ranked similarity loss function with pair weighting (dubbed RMS loss). The proposed RMS loss can keep a distance between samples of different classes by weighting the negative samples according to the sorting order. Meanwhile, it further widens the distance between positive and negative samples by different processing of similarity of positive pairs and negative pairs. The effectiveness of our method is evaluated by extensive experiments on four public datasets and compared with state-of-the-art methods. The results show the proposed method obtains new performance on four public datasets, e.g., reaching 67.4% on CUB200 at Recall@1.
Jian Wang 0130, Dongmei Huang 0001, Wei Song 0007, Quanmiao Wei, Xinyue Li 0001
ICASSP4
2021 Determining Wave Height from Nearshore Videos Based on Multi-level Spatiotemporal Feature Fusion
abstract
Wave height is one of the most significant parameters of waves, affecting marine navigation, coastal environment, and surfing activities. In this paper, we present a wave height determination network from nearshore surveillance videos. This method consists of three parts. Firstly, video frames and frame differences are extracted to represent static and dynamic characteristics of waves, respectively. Then, by taking the two types of data as different modalities, two independent Network In Network (NIN) networks are constructed to learn spatial and temporal features of wave heights. Meanwhile, the two types of features are fused across multiple feature levels using a central network, where concatenation or attention-based fusion is applied. Finally, the wave height is determined by minimizing a global loss. The proposed network for wave height determination, dubbed NINMf-WH or NINMf-A-WH (with attention), was evaluated by a set of comparative experiments from two aspects: i) the basic network of feature learning: NIN vs. CNN; ii) the spatial-temporal feature fusion strategy: multi-level fusion vs. late fusion vs. intermediate feature fusion. The results show that NIN is suited to learn wave height features due to the nonlinearity of waves in video, and the multi-level spatiotemporal feature fusion can achieve low error and good stability of wave height determination. With attention-based fusion, NINMf-A-WH can further reduce the prediction error. To verify the generalization performance of the proposed network, we applied it for estimating wave heights in new surveillant videos. The results provide a low relative error of 6.43%±4.9, which can well meet the operational requirement for nearshore wave forecasting (<=20%).
Wei Song 0007, Qi-chao Li, Qi He 0003
IJCNN1
2021 Median-Pooling Grad-CAM: An Efficient Inference Level Visual Explanation for CNN Networks in Remote Sensing Image Classification
Wei Song 0007, Shuyuan Dai, Dongmei Huang 0001, Jinling Song, Antonio Liotta
MMM (2)1
2021 Automatic Sea-Ice Classification of SAR Images Based on Spatial and Temporal Features Learning
abstract
Sea ice has a significant effect on climate change and ship navigation. Hence, it is crucial to draw sea-ice maps that reflect the geographical distribution of different types of sea ice. Many automatic sea-ice classification methods using synthetic aperture radar (SAR) images are based on the polarimetric characteristics or image texture features of sea ice. They either require professional knowledge to design the parameters and features or are sensitive to noise and condition changes. Moreover, ice changes over time are often ignored. In this article, we propose a new SAR sea-ice image classification method based on a combined learning of spatial and temporal features, derived from residual convolutional neural networks (ResNet) and long short-term memory (LSTM) networks. In this way, we achieve automatic and refined classification of sea-ice types. First, we construct a seven-type ice data set according to the Canadian Ice Service ice charts. We extract spatial feature vectors of a time series of sea-ice samples using a trained ResNet network. Then, using the feature vectors as inputs, the LSTM network further learns the variation of the set of sea-ice samples with time. Finally, the extracted high-level features are fed into a softmax classifier to output the most recent ice type. Taking both spatial features and time variation into consideration, our method can achieve a high classification accuracy of 95.7% for seven ice types. Our method can automatically produce more objective sea-ice interpretation maps, allowing detailed sea-ice distribution and improving the efficiency of sea-ice monitoring tasks.
Wei Song 0007, Wen Gao 0018, Dongmei Huang 0001, Zhenling Ma, Antonio Liotta, Cristian Perra
IEEE Trans. Geosci. Remote. Sens.1
2018 Multi-modal Remote Sensing Image Classification for Low Sample Size Data
abstract
Recently, multiple and heterogeneous remote sensing images have provided a new development opportunity for Earth observation research. utilizing deep learning to gain the shared representative information between different modalities is important to resolve the problem of geographical region classification. In this paper, a CNN-based multi-modal framework for low-sample-size data classification of remote sensing images is introduced. This method has three main stages. Firstly, features are extracted from high- and low-resolution remote sensing images separately using multiple convolution layers. Then, the two types of features are fused at the fusion algorithm layer. Finally, the fused features are used to train a classifier. The novelty of this method is that not only it considers the complementary relationship between the two modalities, but enhances the value of a small number of samples. Based on our experiments, the proposed model can obtain a state-of-the-art performance, being more accurate than the comparable architectures, such as single-modal LeNet, NanoNets and multi-modal H&L-LeNet that are trained with a double size of samples.
Qi He 0003, Yao Lee, Dongmei Huang 0001, Shengqi He, Wei Song 0007, Yanling Du
IJCNN5
2018 Shallow-Water Image Enhancement Using Relative Global Histogram Stretching Based on Adaptive Parameter Acquisition
Dongmei Huang 0001, Yan Wang 0068, Wei Song 0007, Jean Sequeira, Sébastien Mavromatis
MMM (1)3
2018 Effects of light field subsampling on the quality of experience in refocusing applications
abstract
Light field acquisition devices can capture static and dynamic information of the light field in space that is processed using computational imaging techniques for creating visual representation of the captured scene. The main applications are the creation of visual effects such as perspective change, refocusing, recoloring. Advances in light field technologies will provide the devices and tools needed for the development of new services in the virtual/augmented/merged reality domains. This paper presents an analysis of the effects of light field subsampling on the quality of experience in refocusing applications.
Cristian Perra, Wei Song 0007, Antonio Liotta
QoMEX2
2018 Marine Information System Based on Ocean Data Ontology Construction
abstract
Aiming at the low recall rate of retrieval and the difficulty of sharing due to the non-uniform description of ocean data and their relations, a general framework of marine information system based on ocean data ontology (MDO) is proposed in this paper. The framework consists of four layers: basic data, ontology construction, support and application. The ontology construction layer is based on the basic data layer, and this paper analyzes the characteristics of marine data, and classifies semantic concepts, then establishes the spatial and semantic associations between data. The MDO can support various data analysis and provide the data required in marine applications. Finally, applying the proposed framework into the construction of a demonstration system of polar marine environment monitoring, it is proved that the introduction of MDO for polar environment data management can effectively enhance the accuracy and intelligence of data retrieval, and also provide integrated data support for polar ocean environmental assessments.
Dongmei Huang 0001, Jian Wang 0130, Antonio Liotta, Wei Song 0007, Jiangang Zhu
SMC5
2017 Impact of automatic region-of-interest coding on perceived quality in mobile video
Ivan Himawan, Wei Song 0007, Dian Tjondronegoro
Multim. Tools Appl.2
2015 QoE Modelling for VP9 and H.265 Videos on Mobile Devices
abstract
Current mobile devices and streaming video services support high definition (HD) video, increasing expectation for more contents. HD video streaming generally requires large bandwidth, exerting pressures on existing networks. New generation of video compression codecs, such as VP9 and H.265/HEVC, are expected to be more effective for reducing bandwidth. Existing studies to measure the impact of its compression on users" perceived quality have not been focused on mobile devices. Here we propose new Quality of Experience (QoE) models that consider both subjective and objective assessments of mobile video quality. We introduce novel predictors, such as the correlations between video resolution and size of coding unit, and achieve a high goodness-of-fit to the collected subjective assessment data (adjusted R-square >83%). The performance analysis shows that H.265 can potentially achieve 44% to 59% bit rate saving compared to H.264/AVC, slightly better than VP9 at 33% to 53%, depending on video content and resolution.
Wei Song 0007, Dian Tjondronegoro, Antonio Liotta
ACM Multimedia1
2014 Acceptability-based QoE Management for User-centric Mobile Video Delivery: A Field Study Evaluation
abstract
Effective Quality of Experience (QoE) management for mobile video delivery -- to optimize overall user experience while adapting to heterogeneous use contexts -- is still a big challenge to date. This paper proposes a mobile video delivery system to emphasize the use of acceptability as the main indicator of QoE to manage the end-to-end factors in delivering mobile video services. The first contribution is a novel framework for user-centric mobile video system that is based on acceptability-based QoE (A-QoE) prediction models, which were derived from comprehensive subjective studies. The second contribution is results from a field study that evaluates the user experience of the proposed system during realistic usage circumstances, addressing the impacts of perceived video quality, loading speed, interest in content, viewing locations, network bandwidth, display devices, and different video coding approaches, including region-of-interest (ROI) enhancement and center zooming.
Wei Song 0007, Dian Tjondronegoro, Ivan Himawan
ACM Multimedia1
2014 Acceptability-Based QoE Models for Mobile Video
abstract
Quality of experience (QoE) measures the overall perceived quality of mobile video delivery from subjective user experience and objective system performance. Current QoE prediction models have two main limitations: (1) insufficient consideration of the factors influencing QoE, and (2) limited studies on QoE models for acceptability prediction. In this paper, a set of novel acceptability-based QoE models, denoted as A-QoE, is proposed based on the results of comprehensive user studies on subjective quality acceptance assessments. The models are able to predict users' acceptability and pleasantness in various mobile video usage scenarios. Statistical nonlinear regression analysis has been used to build the models with a group of influencing factors as independent predictors, which include encoding parameters and bitrate, video content characteristics, and mobile device display resolution. The performance of the proposed A-QoE models has been compared with three well-known objective Video Quality Assessment metrics: PSNR, SSIM and VQM. The proposed A-QoE models have high prediction accuracy and usage flexibility. Future user-centred mobile video delivery systems can benefit from applying the proposed QoE-based management to optimize video coding and quality delivery strategies.
Wei Song 0007, Dian Tjondronegoro
IEEE Trans. Multim.1
2013 Automatic region-of-interest detection and prioritisation for visually optimised coding of low bit rate videos
abstract
The increasing popularity of video consumption from mobile devices requires an effective video coding strategy. To overcome diverse communication networks, video services often need to maintain sustainable quality when the available bandwidth is limited. One of the strategy for a visually-optimised video adaptation is by implementing a region-of-interest (ROI) based scalability, whereby important regions can be encoded at a higher quality while maintaining sufficient quality for the rest of the frame. The result is an improved perceived quality at the same bit rate as normal encoding, which is particularly obvious at the range of lower bit rate. However, because of the difficulties of predicting region-of-interest (ROI) accurately, there is a limited research and development of ROI-based video coding for general videos. In this paper, the phase spectrum quaternion of Fourier Transform (PQFT) method is adopted to determine the ROI. To improve the results of ROI detection, the saliency map from the PQFT is augmented with maps created from high level knowledge of factors that are known to attract human attention. Hence, maps that locate faces and emphasise the centre of the screen are used in combination with the saliency map to determine the ROI. The contribution of this paper lies on the automatic ROI detection technique for coding a low bit rate videos which include the ROI prioritisation technique to give different level of encoding qualities for multiple ROIs, and the evaluation of the proposed automatic ROI detection that is shown to have a close performance to human ROI, based on the eye fixation data.
Ivan Himawan, Wei Song 0007, Dian Tjondronegoro
WACV2
2012 Impact of Region-of-Interest Video Coding on Perceived Quality in Mobile Video
abstract
Effective streaming of video can be achieved by providing more bits to the most important region in the frame at the cost of reduced bits in the less important regions. This strategy can be beneficial for delivering high quality videos in mobile devices, especially when the availability of bandwidth is usually low and limited. While the state-of-the-art video codecs such as H.264 may have been optimised for perceived quality, it is hypothesised that users will give more attention to interesting region/object when watching videos. Therefore, giving a higher quality to region of interest (ROI) while reducing quality of other areas may result in improving the overall perceived quality without necessarily increasing the bitrate. In this paper, the impact of ROI-based encoded video on perceived quality is investigated by conducting a user study for various target bit rates. The results from the user study demonstrate that ROI-based video coding has superior perceived quality compared to normal encoded video at the same bitrate in the lower bitrate range.
Ivan Himawan, Wei Song 0007, Dian Tjondronegoro
ICME2
2011 Saving bitrate vs. pleasing users: where is the break-even point in mobile video quality?
abstract
This paper presents a comprehensive study to find the most efficient bitrate requirement to deliver mobile video that optimizes bandwidth, while at the same time maintains good user viewing experience. In the study, forty participants were asked to choose the lowest quality video that would still provide for a comfortable and long-term viewing experience, knowing that higher video quality is more expensive and bandwidth intensive. This paper proposes the lowest pleasing bitrates and corresponding encoding parameters for five different content types: cartoon, movie, music, news and sports. It also explores how the lowest pleasing quality is influenced by content type, image resolution, bitrate, and user gender, prior viewing experience, and preference. In addition, it analyzes the trajectory of users' progression while selecting the lowest pleasing quality. The findings reveal that the lowest bitrate requirement for a pleasing viewing experience is much higher than that of the lowest acceptable quality. Users' criteria for the lowest pleasing video quality are related to the video's content features, as well as its usage purpose and the user's personal preferences. These findings can provide video providers guidance on what quality they should offer to please mobile users.
Wei Song 0007, Dian Tjondronegoro, Michael J. Docherty
ACM Multimedia1
2011 Measuring Bitrate and Quality Trade-Off in a Fast Region-of-Interest Based Video Coding
Salahuddin A. Azad, Wei Song 0007, Dian Tjondronegoro
MMM (2)2
2011 User-driven saliency maps for evaluating Region-of-Interest detection
abstract
Detection of Region of Interest (ROI) in a video leads to more efficient utilization of bandwidth. This is because any ROIs in a given frame can be encoded in higher quality than the rest of that frame, with little or no degradation of quality from the perception of the viewers. Consequently, it is not necessary to uniformly encode the whole video in high quality. One approach to determine ROIs is to use saliency detectors to locate salient regions. This paper proposes a methodology for obtaining ground truth saliency maps to measure the effectiveness of ROI detection by considering the role of user experience during the labelling process of such maps. User perceptions can be captured and incorporated into the definition of salience in a particular video, taking advantage of human visual recall within a given context. Experiments with two state-of-the-art saliency detectors validate the effectiveness of this approach to validating visual saliency in video. This paper will provide the relevant datasets associated with the experiments.
Ivan Himawan, Wei Song 0007, Dian Tjondronegoro
WACV2
2010 Bitrate modeling of scalable videos using quantization parameter, frame rate and spatial resolution
abstract
The quality and bitrate modeling is essential to effectively adapt the bitrate and quality of videos when delivered to multiplatform devices over resource constraint heterogeneous networks. The recent model proposed by Wang et al. estimates the bitrate and quality of videos in terms of the frame rate and quantization parameter. However, to build an effective video adaptation framework, it is crucial to incorporate the spatial resolution in the analytical model for bitrate and perceptual quality adaptation. Hence, this paper proposes an analytical model to estimate the bitrate of videos in terms of quantization parameter, frame rate, and spatial resolution. The model can fit the measured data accurately which is evident from the high Pearson correlation. The proposed model is based on the observation that the relative reduction in bitrate due to decreasing spatial resolution is independent of the quantization parameter and frame rate. This modeling can be used for rate-constrained bit-stream adaptation scheme which selects the scalability parameters to optimize the perceptual quality for a given bandwidth constraint.
Salahuddin A. Azad, Wei Song 0007, Dian Tjondronegoro
ICASSP2
2010 Impact of zooming and enhancing region of interests for optimizing user experience on mobile sports video
abstract
In mobile videos, small viewing size and bitrate limitation often cause unpleasant viewing experiences, which is particularly important for fast-moving sports videos. For optimizing the overall user experience of viewing sports videos on mobile phones, this paper explores the benefits of emphasizing Region of Interest (ROI) by 1) zooming in and 2) enhancing the quality. The main goal is to measure the effectiveness of these two approaches and determine which one is more effective. To obtain a more comprehensive understanding of the overall user experience, the study considers user's interest in video content and user's acceptance of the perceived video quality, and compares the user experience in sports videos with other content types such as talk shows. The results from a user study with 40 subjects demonstrate that zooming and ROI-enhancement are both effective in improving the overall user experience with talk show and mid-shot soccer videos. However, for the full-shot scenes in soccer videos, only zooming is effective while ROI-enhancement has a negative effect. Moreover, user's interest in video content directly affects not only the user experience and the acceptance of video quality, but also the effect of content type on the user experience. Finally, the overall user experience is closely related to the degree of the acceptance of video quality and the degree of the interest in video content. This study is valuable in exploiting effective approaches to improve user experience, especially in mobile sports video streaming contexts, whereby the available bandwidth is usually low or limited. It also provides further understanding of the influencing factors of user experience.
Wei Song 0007, Dian Tjondronegoro, Tony Shu-Hsien Wang, Michael J. Docherty
ACM Multimedia1
2010 User-Centered Video Quality Assessment for Scalable Video Coding of H.264/AVC Standard
Wei Song 0007, Dian Tjondronegoro, Salahuddin A. Azad
MMM1
2010 Exploration and Optimization of User Experience in Viewing Videos on a Mobile Phone
abstract
Compared with viewing videos on PCs or TVs, mobile users have different experiences in viewing videos on a mobile phone due to different device features such as screen size and distinct usage contexts. To understand how mobile user's viewing experience is impacted, we conducted a field user study with 42 participants in two typical usage contexts using a custom-designed iPhone application. With user's acceptance of mobile video quality as the index, the study addresses four influence aspects of user experiences, including context, content type, encoding parameters and user profiles. Accompanying the quantitative method (acceptance assessment), we used a qualitative interview method to obtain a deeper understanding of a user's assessment criteria and to support the quantitative results from a user's perspective. Based on the results from data analysis, we advocate two user-driven strategies to adaptively provide an acceptable quality and to predict a good user experience, respectively. There are two main contributions from this paper. Firstly, the field user study allows a consideration of more influencing factors into the research on user experience of mobile video. And these influences are further demonstrated by user's opinions. Secondly, the proposed strategies — user-driven acceptance threshold adaptation and user experience prediction — will be valuable in mobile video delivery for optimizing user experience.
Wei Song 0007, Dian Tjondronegoro, Michael J. Docherty
Int. J. Softw. Eng. Knowl. Eng.1