EDBT 2026 Demo / reviewers in the wild / expert
Yiping Duan
dblp:158/8256
· DBLP profile ↗
57ranked-venue papers
6as first author
39since 2021 · last 2026
0000-0001-9638-7112ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 1 first-author · 16 since 2021Computer networks · 17 · 2 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Perceptual Audio-Visual Quality Assessment for Compressed Professionally-Generated Content: A New Dataset and An Effective Method
Shuzhan Hu, Yiping Duan, Mingqiang Yuan, Xiaoming Tao 0001 |
ICC | 3 |
| 2026 | Gumbel Sparse Attention Spatio-Temporal Network: A framework for traffic risk prediction
Dongkun Wang, Jieyang Peng, Songsheng Wang, Xiaoming Tao 0001, Yiping Duan |
Adv. Eng. Informatics | 5 |
| 2026 | Semantic encoding for image compression based on semantic segmentation maps
Zhipeng Xie, Yiping Duan, Wei Kun Kong, Qiyuan Du, Qinghua Liang, Xiaoming Tao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | RIDE: Redensification-based intrinsic density estimation for knowledge graphs
Wei Kun Kong, Yiping Duan, Xiaoming Tao 0001, Yueran Zu |
Pattern Recognit. | 2 |
| 2026 | Memory-Enhanced Dynamic Self-Attention Communication Mechanism for Multi-UAV Base Stations Trajectory Planning
Hanxiao Yuan, Yao Shi 0002, Emad Alsusa, Qinyu Zhang 0001, Yiping Duan, Xiaohu You 0001 |
IEEE Trans. Commun. | 5 |
| 2025 | Lite3D: A Lightweight Hybrid CNN-Transformer Framework for 3D Video StabilizationabstractVideo stabilization plays a critical role in IoT, as it enhances the accuracy and reliability of visual data collected from mobile and unstable devices, enabling more effective monitoring and analysis. However, existing stabilization methods frequently lack sufficient accuracy in motion estimation or struggle with expensive computation. In this paper, we introduce Lite3D, a novel lightweight 3D deep learning-based video stabilization framework that combines a hybrid CNN and Transformer architecture for more efficient feature extraction. During training, our method utilizes estimated depth maps and relative camera poses to generate target views, jointly optimizing DepthNet and PoseNet through an unsupervised learning strategy. Compared to existing 3D video stabilization methods, our approach enables more accurate learning of motion information. In the inference phase, we smooth the estimated camera pose trajectory and synthesize stabilized frames using the optimal depth maps. Experimental results on the NUS dataset demonstrate that Lite3D achieves state-of-the-art stability, with additional tests on the DeepStab dataset confirming strong generalization. Moreover, our method contains fewer parameters compared to existing 3D deep learning-based video stabilization, offering both high stabilization quality and computational efficiency. Zejing Shan, Yiping Duan, Yue Wu 0004, Qiyuan Du, Xiaoming Tao 0001 |
ICC | 2 |
| 2025 | Cloud-Edge-End Collaborative Surveillance Video Transmission with Object-Guided Video Super-ResolutionabstractCloud Video Surveillance (CVS) systems, as the backbone of distributed surveillance networks, face increasing challenges in transmitting large volumes of high-resolution video data. While cloud-end collaborative video transmission methods can reduce bitrates beyond conventional compression techniques, the periodic transmission of high-resolution keyframes consumes significant bandwidth due to semantically irrelevant background information. To address this, we propose a cloud-edge-end collaborative video transmission scheme based on object-guided video super-resolution. In this scheme, the edge extracts key objects from sparsely selected keyframes at the end and transmits them along with low-resolution video to the cloud, where our Keyframe-Guided Video Restoration Transformer (KG-VRT) is used to improve the video quality. Experimental results on public datasets show that our network outperforms state-of-the-art keyframe-based baselines with a 1.73 dB PSNR improvement and maintains robust performance even with keyframe intervals of up to 30 frames. A comparative analysis of two transmission strategies—transmitting full keyframes versus transmitting only key object regions—demonstrates a 60% – 80% reduction in keyframe bitrate while maintaining object detection accuracy at a significantly reduced overall system bitrate. This highlights the efficiency and scalability of our approach in bandwidth-constrained surveillance scenarios. Yiping Duan, Xiaoming Tao 0001, Wei Kun Kong, Qiyuan Du |
VTC2025-Fall | 2 |
| 2025 | Joint semi-grant-free NOMA for dual-layer LEO cohesive clustered satellite systems
Yao Shi 0002, Qinyu Zhang 0001, Yiping Duan |
Sci. China Inf. Sci. | 4 |
| 2025 | EEG-Driven Classification of Driver Mental Workload in Diverse Environments: A Dual-Branch Network for Efficient In-Vehicle ApplicationsabstractThe mental load of drivers can profoundly affect their driving performance, to the extent that it affects traffic safety. Therefore, monitoring mental workload has become a crucial aspect of sensor-based driver monitoring systems, especially in the context of the Industrial Internet of Things (IIoT), where driver status information can be exchanged between vehicles to enhance safety. However, the substantial energy consumption and transmission latency associated with traditional central-server-based IoT systems are prominent issues that necessitate the development of lighter algorithms for edge computing in individual vehicles. In this article, we focus on the impact of external traffic events and environmental changes on the mental load of drivers, as well as effective classification algorithms applied in monitoring systems. To analyze the physiological responses of drivers to road events and non driving related tasks under different weather conditions, we proposed a dual branch model, DMW-Net, based on attention mechanism branches and graph attention modules to discriminate the mental load level of drivers from physiological signals. The proposed method was validated on the manD dataset and achieved an accuracy of 90.07% in physiological signals of three different load levels, which is higher than the comparison models. This study provides innovative methods for driver monitoring systems, contributing to advanced driving assistance systems (ADAS) and traffic safety. Yanjun Qin, Shanghang Zhang, Yiping Duan, Xiaoming Tao 0001 |
IEEE Internet Things J. | 4 |
| 2025 | Object-Attribute-Relation Representation-Based Video Semantic CommunicationabstractWith the rapid growth of multimedia data volume, there is an increasing need for efficient video transmission in applications such as virtual reality and future video streaming services. Semantic communication is emerging as a vital technique for ensuring efficient and reliable transmission in low-bandwidth, high-noise settings. However, most current approaches focus on joint source-channel coding (JSCC) that depends on end-to-end training. These methods often lack an interpretable semantic representation and struggle with adaptability to various downstream tasks. In this paper, we introduce the use of object-attribute-relation (OAR) as a semantic framework for videos to facilitate low bit-rate coding and enhance the JSCC process for more effective video transmission. We utilize OAR sequences for both low bit-rate representation and generative video reconstruction. Additionally, we incorporate OAR into the image JSCC model to prioritize communication resources for areas more critical to downstream tasks. Our experiments on traffic surveillance video datasets assess the effectiveness of our approach in terms of video transmission performance. The empirical findings demonstrate that our OAR-based video coding method not only outperforms H.265 coding at lower bit-rates but also synergizes with JSCC to deliver robust and efficient video transmission. Qiyuan Du, Yiping Duan, Qianqian Yang 0002, Xiaoming Tao 0001, Mérouane Debbah |
IEEE J. Sel. Areas Commun. | 2 |
| 2025 | Data-Free Cloud-Edge Distillation for Safe and Efficient Intelligent CommunicationsabstractEfficiency and security are the core challenges in the intelligent communication field. Lightweight neural networks have accelerated the information interpretation and communication efficiency, thereby fostering the rapid development of the Internet of Things. Enhancing the recognition capability of lightweight neural networks remains challenging. Knowledge distillation, a technique that transfers knowledge from a complex model to a smaller one, is often used to improve the recognition performance of lightweight networks. However, practical issues such as transmission constraints and user privacy make the original data required for knowledge distillation difficult to access directly. To tackle this issue, this paper proposes a Data-Free Cloud-Edge Knowledge Distillation (DF-CEKD) model, which uses a complex network in the cloud to provide training guidance for lightweight networks deployed on mobile devices. Specifically, DF-CEKD employs a novel Deep Inversion Diffusion Generation (DIDG) module to provide proxy data as input for the distillation process, thereby transferring the feature learning capability from the cloud network to the edge network. Meanwhile, a Multi-Layer Feature Joint Supervision Distillation (MLF-JSD) module is designed to further enhance the feature selection guidance provided by the teacher network in the cloud for training the lightweight student network. The simulation results demonstrate that the proposed DF-CEKD reduces the number of parameters to 1/20 and the floating-point operations to 1/12 when distilling from WRN40-2 to WRN16-1, resulting in only a 0.27% decrease in accuracy. Xiufang Li, Yiping Duan, Xiaoming Tao 0001, Qigong Sun, Qiyuan Du, Qianqian Yang 0002, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 2 |
| 2025 | Brain-Inspired Video Quality Assessment via Visual-EEG Feature AlignmentabstractVideo quality assessment (VQA) is crucial in applications such as video calls, real-time meetings, and surveillance, where video quality directly impacts user experience greatly. Traditional objective methods like SSIM and PSNR fail to capture the subjective perception of video quality, while subjective Quality of Experience (QoE) assessment metrics like Mean Opinion Score (MOS) are not scalable for large-scale automated VQA tasks. To overcome these limitations, deep learning approaches have emerged, but mostly focusing only on a single video modality, extracting low-level visual features such as color and texture. Recently, electroencephalography (EEG) has been shown to align with users' subjective experiences, offering valuable insights into neural responses to visual content. Hence, in this letter, we propose a brain-inspired deep learning framework for VQA that aligns EEG and video features. We build a video distortion dataset annotated with both MOS and EEG signals to analyze the impact of video distortions on EEG responses and subjective ratings. We then employ an adaptive EEG feature learning network to extract EEG features linked to video distortions, and propose a video quality prediction network that aligns both video and EEG features using a three-stage training strategy. Our method outperforms existing techniques, showing strong alignment with human subjective ratings. Experimental results validate the effectiveness of EEG in enhancing VQA with a more human-centric approach. Shuzhan Hu, Chenxing Li, Yiping Duan, Xiaoming Tao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Capacity Optimizing Resource Allocation in Joint Source-Channel Coding Systems With QoS ConstraintsabstractBenefited from the advances of deep learning (DL) techniques, deep joint source-channel coding (JSCC) has shown its great potential to improve the performance of wireless transmission. However, most of the existing works focus on the DL-based transceiver design of the JSCC model, while ignoring the resource allocation problem in wireless systems. In this paper, we consider a downlink resource allocation problem, where a base station (BS) jointly optimizes the compression ratio (CR) and power allocation as well as resource block (RB) assignment of each user according to the latency and performance constraints to maximize the number of users that successfully receive their requested content with desired quality. To solve this problem, we first decompose it into two subproblems without loss of optimality. The first subproblem is to minimize the required transmission power for each user under given RB allocation. We derive the closed-form expression of the optimal transmit power by searching the maximum feasible compression ratio. The second one aims at maximizing the number of supported users through optimal user-RB pairing, which is solved by utilizing bisection search as well as interior-point algorithm. To reduce the computational complexity, we propose a heuristic greedy algorithm to obtain a simplified problem. Then the Bregman alternating direction method of multipliers (BADMM) based algorithm is adopted to decompose the simplified problem into several subproblems that can be computed in parallel. Simulation results validate the effectiveness of the proposed resource allocation methods in terms of the number of satisfied users with given resources. It is also shown that the BADMM-based algorithm can significantly reduce the computational complexity and retain high performance. Kaiyi Chi, Qianqian Yang 0002, Zhaohui Yang 0001, Yiping Duan, Zhaoyang Zhang 0001 |
IEEE Trans. Commun. | 4 |
| 2025 | Group Image Compression for Dual Use of Machine and Human VisionabstractFaces in a scene of human group, if coded with sufficient precision, can be computer analyzed for machine vision tasks involving faces. But this requires storing and communicating them at a very high bit rate. Traditional ROI-based image compression methods are ill suited to code many faces at high precision against a complex background. In this work, we propose a novel group image compression neural network (GICNet) of two layers: 1) the face layer dedicated to machine analysis, in which face bounding boxes are first cropped out of the background and converted to a compression-friendly canonical sketch-guided representation of fixed resolution for compact coding and facilitating downstream tasks without additional preprocessing; 2) the background layer dedicated to overall human vision perceptual quality, in which face residuals and background elements are coded and appended to the code stream. Experimental results demonstrate the effectiveness of our proposed GICNet, conserving up to 13%-57% bitrate for machine vision applications while maintaining competitive perceptual quality. Xiaolin Wu 0001, Fan Li 0003, Yiping Duan, Xiaoming Tao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Brain-Inspired VR Video Quality Assessment Based on ElectroencephalographyabstractWith the rapid development of virtual reality (VR) technology, users are able to access a large number of new applications in their daily lives. VR expands users’ perceptual dimensions, bringing them entirely new experience. However, the user experience assessment for VR videos is still under exploration, which remains an unresolved issue. In such immersive scenarios, the methods based on user scoring require active feedback from users, which will interrupt the immersion experience. Besides, it is difficult to monitor the user experience status in real-time by user scoring. With the development of psychophysiological research, electroencephalographic (EEG) signal measurement is considered to have the potential to non-intrusively obtain the user experience. Hence, this paper employs EEG measurements to capture users’ EEG signals while watching VR videos with varying levels of stuttering, constructing a VR-EEG dataset. Subsequently, we analyze the dataset using time-frequency analysis methods to validate the feasibility of EEG signals reflecting user experience. Finally, we utilize machine learning methods to construct a QoE measurement network capable of analyzing users’ perceptual experience from single-trial EEG signals. Experimental results demonstrate that the proposed method establishes a relationship between brain activities and user experience and can effectively predict QoE scores from EEG signals. It provides a technical means for real-time, non-disturbing measurement of user experience in VR video playback. Shuzhan Hu, Jian Chu, Yiping Duan, Xiaoming Tao 0001, Jianhua Lu |
GLOBECOM | 3 |
| 2024 | Semantic Security: A Digital Watermark Method for Image Semantic PreservationabstractDigital watermarking has long been used to protect digital images from abuse. However, applying digital watermarking to semantic communication remains a challenge. This work introduces a secure coding method that combines semantic coding and digital watermarking techniques. The proposed method selects points in the target image with high semantic importance and embeds watermark information into their position vectors, perpendicular to the embedding domains of previous works. The experiments conducted on the Cityscapes dataset demonstrate that our method integrates well with semantic communication systems. Compared to the previous approach, our proposed method can more completely preserve the structural features of the target image and is better suited for Machine Type Communication (MTC) tasks, such as target detection and semantic segmentation. Tianwei Zuo, Yiping Duan, Qiyuan Du, Xiaoming Tao 0001 |
ICASSP | 2 |
| 2024 | Pilot-Free Semantic Communication Over Multi-User Mimo Fading ChannelsabstractWireless communication systems operating in fading channels often demand pilots for channel estimation and data recovery, leading to substantial transmission overhead. In this paper, we propose a novel pilot-free semantic communication system designed for transmitting images over multi-user MIMO (MU-MIMO) fading channels. Specifically, our method involves extracting multi-scale semantic features from the source image at the transmitter, effectively embedding pilot-like information. At the receiver, we extract channel features from these semantic features at each scale, enabling the reconstruction of the source image without requiring explicit channel estimation and signal detection. To enhance the image reconstruction process, we introduce a novel module, called Resnet Transformer, which combines multi-head self-attention (MHSA) with Resnet block. Our experimental results demonstrate that this pilot-free system outperforms existing pilot-aided semantic communication methods in terms of perceptual quality and transmission efficiency. Weixuan 'Vincent' Chen, Qianqian Yang 0002, Zhaohui Yang 0001, Yiping Duan, Zhaoyang Zhang 0001 |
ICIP | 4 |
| 2024 | Brain-Inspired Image Perceptual Quality Assessment Based on EEG: A QoE PerspectiveabstractHuman-oriented image communication should take the quality of experience (QoE) as an optimization goal, which requires effective image perceptual quality metrics. However, traditional user-based assessment metrics are limited by the deviation caused by human high-level cognitive activities. To tackle this issue, in this paper, we construct a brain response-based image perceptual quality metric and develop a brain-inspired network to assess the image perceptual quality based on it. Our method aims to establish the relationship between image quality changes and underlying brain responses in image compression scenarios using the electroencephalography (EEG) approach. We first establish EEG datasets by collecting the corresponding EEG signals when subjects watch distorted images. Then, we design a measurement model to extract EEG features that reflect human perception to establish a new image perceptual quality metric: EEG perceptual score (EPS). To use this metric in practical scenarios, we embed the brain perception process into a prediction model to generate the EPS directly from the input images. Experimental results show that our proposed measurement model and prediction model can achieve better performance. The proposed brain response-based image perceptual quality metric can measure the human brain's perceptual state more accurately, thus performing a better assessment of image perceptual quality. Shuzhan Hu, Yiping Duan, Xiaoming Tao 0001, Geoffrey Ye Li, Jianhua Lu, Guangyi Liu 0001, Zhimin Zheng, Chengkang Pan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Optical Flow-Based Spatiotemporal Sketch for Video Representation: A Novel FrameworkabstractWith the rapid development of multimedia services and the dramatic growth of video data volume, efficient video representation and AI-generated content (AIGC) become critical parts of future multimedia communication systems. Sketch graph is a structured abstraction of key textures in an image, and video sketch graph further exploits the temporal continuity of videos to achieve a sparse representation. Sketch-based representation has potential applications in communication systems for both human subjective perception and machine vision tasks, and provides a new idea for AIGC. However, current video sketch extraction methods rely on human assistance and correction, and cannot be applied to end-to-end communication systems. We design a novel framework for spatiotemporal sketch extraction based on deep learning methods. In the proposed framework, sketch extraction and sparse coding are performed at the sender side using structural and temporal features of the video. The original videos are generatively reconstructed at the receiver side or applied to downstream machine vision tasks. We validate the performance of the proposed method on Cityscapes dataset with different metrics. Experiments show that our proposed framework can be end-to-end adapted to video communication tasks in different scenarios and can achieve efficient video characterization and transmission. Moreover, our proposed method enables sketch-based end-to-end AIGC for video generation. Qiyuan Du, Yiping Duan, Zhipeng Xie, Xiaoming Tao 0001, Linsu Shi, Zhijuan Jin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | OARNet: Object-Attribute-Relation Network for Predicting Soccer EventsabstractEvent prediction involves analyzing and forecasting events that occur at a specific time and location to inform decision-making and take the next actions. Current event prediction approaches primarily employ deep learning methods to analyze regular patterns from large amounts of historical data. However, predicting adversarial soccer events remains a significant challenge due to strong antagonism and complex relationships between players. With this consideration, we propose an objectattribute-relation (OAR) network for predicting soccer events using multimodal data, including spatiotemporal trajectory data and video data. The proposed scheme aims to enhance prediction performance by transforming multimodal data into an OAR space that integrates global and local relationships (adversarial information and multi-objective information). In particular, the scheme consists mainly of a relation module, an object attribute module, and a graph prediction module. We first use ConvLSTM to extract the visual features of players from video data and use LSTM to extract the movement features of players from spatiotemporal data. Additionally, we apply a multihead GRU attention mechanism to calculate the relation weights. These three components are then combined into an OAR graph of a clip in a soccer game. Finally, an OAR GNN is designed to determine the influence of different objects and predict events. The entire process constitutes an end-to-end event prediction learning framework. Extensive experimental results on the two challenging datasets, namely, soccER and SkillCorner, verify the effectiveness of the proposed framework. Yiping Duan, Xiaoming Tao 0001, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2023 | Visual Attention Measurement Based on Electroencephalogram Feature LearningabstractWith the continuous development of multimedia technology, there is a growing demand for video applications. Among the huge number of videos, some video clips attract higher visual attention from viewers. Measuring this visual attention is not only crucial for evaluating the quality of experience (QoE) but also holds the potential for guiding video compression techniques. Therefore, there is an urgent need to propose an effective method to evaluate the changes of viewers' attention. To address this challenge, we employ electroencephalography (EEG) as a bridge to establish the relationship between video clips and visual attention. We introduce a visual attention measurement (VAM) method by combining EEG and machine learning approaches. Specifically, we design an EEG experiment to collect brain responses from viewers while watching various video clips, thus constructing a visual attention EEG dataset. Using these EEG signals, we propose a VAM network to identify whether viewers pay attention to the video clips. The experimental results show that our EEG experiment can capture the physiological signals related to visual attention. Moreover, the proposed VAM network can accurately determine viewers' attention levels to video clips. These results provide new insights into the evolution of human-oriented video communications. Jian Chu, Shuzhan Hu, Yiping Duan, Xiaoming Tao 0001 |
GLOBECOM | 3 |
| 2023 | Generative Model based Highly Efficient Semantic Communication Approach for Image TransmissionabstractDeep learning (DL) based semantic communication methods have been explored to transmit images efficiently in recent years. In this paper, we propose a generative model based semantic communication to further improve the efficiency of image transmission and protect private information. In particular, the transmitter extracts the interpretable latent representation from the original image by a generative model exploiting the GAN inversion method. We also employ a privacy filter and a knowledge base to erase private information and replace it with natural features in the knowledge base. The simulation results indicate that our proposed method achieves comparable quality of received images while significantly reducing communication costs compared to the existing methods. Tianxiao Han, Jiancheng Tang, Qianqian Yang 0002, Yiping Duan, Zhaoyang Zhang 0001, Zhiguo Shi 0001 |
ICASSP | 4 |
| 2023 | Resource Allocation for Capacity Optimization in Joint Source-Channel Coding SystemsabstractBenefited from the advances of deep learning (DL) techniques, deep joint source-channel coding (JSCC) has shown its great potential to improve the performance of wireless transmission. However, most of the existing works focus on the DL-based transceiver design of the JSCC model, while ignoring the resource allocation problem in wireless systems. In this paper, we consider a downlink resource allocation problem, where a base station (BS) jointly optimizes the compression ratio (CR) and power allocation as well as resource block (RB) assignment of each user according to the latency and performance constraints to maximize the number of users that successfully receive their requested content with desired quality. To solve this problem, we first decompose it into two subproblems without loss of optimality. The first subproblem is to minimize the required transmission power for each user under given RB allocation. We derive the closed-form expression of the optimal transmit power by searching the maximum feasible compression ratio. The second one aims at maximizing the number of supported users through optimal user-RB pairing, which we solve by utilizing bisection search as well as Karmarkar's algorithm. Simulation results validate the effectiveness of the proposed resource allocation method in terms of the number of satisfied users with given resources. Kaiyi Chi, Qianqian Yang 0002, Zhaohui Yang 0001, Yiping Duan, Zhaoyang Zhang 0001 |
ICC | 4 |
| 2023 | Sketch Graph Representation for Multimedia Computational Communications: A Learning-Based MethodabstractMultimedia computational communications towards 6G can improve the transmission efficiency significantly by introducing intelligent computation in the communication process. This intelligent/smart communication architecture includes multimedia representation, coding, transmission and other parts from the perspective of semantics, where multimedia semantic representation is the core part and is mainly utilized to reduce the amount of multimedia data. In this paper, sketch graph is proposed as an effective representation of images to describe the pixel variations, geometric feature distribution and structural information and has potential applications in multimedia computational communications. Specifically, we developed a learning-based method to extract sketch graphs with edge detection, sketch point detection and sketch line detection by deep neural networks (DNNs). Moreover, we designed an end-to-end extraction method and achieved real-time processing. The experimental results on several datasets demonstrated the advanced performance in terms of classification and generation tasks. Image classification results on the HumanSketch, ImageNet, and Caltech datasets showed that sketch graphs extracted by our method had better describing ability than those extracted by traditional methods and other traditional image compression methods. On the other hand, image generation results on the Cityscapes dataset indicated the potential of the sketch-graph-based image compression codec. Qiyuan Du, Yiping Duan, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001 |
ICC | 2 |
| 2023 | Video Reconstruction with Multimodal InformationabstractVideo reconstruction refers to generate videos through the high-level representations (edge map, labels and so on), while the reconstruction quality is always unsatisfactory due to sparse high-level representations, especially on video data. In order to improve the video reconstruction quality, we proposed a novel approach that generates realistic video from its multimodal information including structure features and color features. To extract color features, we mainly apply the k-means algorithm to segment labels and the structure features are extracted by an edge detection network. Video generation is regarded as learning the mapping from multimodal representations to the original videos. So, a conditional GAN is applied with a learning objective that models the temporal video dynamics. We use a spatio-temporal generator with attention to model the inter-frame dynamics and video consistency is improved in this way. Moreover, we use a multiscale discriminator to improve the improve the intra-frame quality of the video. Experimental results on Cityscapes, Apolloscape datasets demonstrate that our proposed approach performs better in both traditional and generative evaluating indicators. Zhipeng Xie, Yiping Duan, Qiyuan Du, Xiaoming Tao 0001, Jiazhong Yu |
VTC Fall | 2 |
| 2023 | Adjustable Dielectric Resonator Antenna With Parasitic Elements for 5G SAGOI-IoT ApplicationsabstractWith the development of modern communication and the Internet of Things (IoT) in the need to provide seamless interconnection between heterogeneous devices, we design an adjustable-distance resonant antenna suitable for 5G communication. The antenna meets the multiangle radiation requirements of the antenna for flexible positioning of IoT devices and conforms to the concept of green Internet basic hardware, with low energy and small size. In our design, three parasitic elements will couple with the higher order modes through the slot-hole excitation of a higher order mode dielectric resonator antenna with a dielectric constant of 10. By controlling the distance between the three parasitic elements and changing the capacitor at their terminals, radiation variation in multiple directions can be achieved. The proposed model focuses on the relationships among the three element distances and the effect of the third parasitic element on the radiation angle. Good results were obtained for gain, bandwidth, and radiation angle. Through simulation, the dielectric resonator antenna works successfully in the 15-GHz frequency band. Thus, the antenna array can be from −34° to 34° in the horizontal direction and from 0° to −36° in the vertical direction. At the same time, the gains are kept at a good level and the bandwidths are greater than 2 GHz. Compared with other dielectric resonant antennas, the parameters of this antenna are not reduced, and the flexibility of the radiation direction angles is increased. These evaluation parameters are considered ideal conditions for device-to-device communication in 5G IoT applications. Chenxing Li, Yiping Duan, Xiaoming Tao 0001 |
IEEE Internet Things J. | 2 |
| 2023 | Sketch Assisted Face Image Coding for Human and Machine Vision: A Joint Training ApproachabstractImage coding is one of the most fundamental techniques and is widely used in image/video processing and multimedia communications. Current image coding methods are mainly human-oriented, and the visual quality is always unsatisfactory, especially at low bitrates. Moreover, the recent emergence of machine vision goes beyond the scope of current coding. With these considerations, we proposed a sketch assisted face image coding for human and machine vision by a joint training approach. In the proposed approach, we design a new feature representation: a color sketch, which aims to satisfy both low-frequency features of human vision and high-frequency features of machine analysis. Then, we present a novel end-to-end image codec framework with joint training that consists of three models: an image-to-image translation module, a coding module, and a two-stage reconstruction module. Specifically, the input image is first translated into the edge map with the Canny edge as the auxiliary label to merely preserve the structure information. Afterward, the backpropagation from reconstruction module guides the edge map to increase or decrease the information through joint training, which results in the generation of color sketch. Then, the generated sketch is compressed into the bitstream and decompressed back to a sketch in the coding module. Finally, the decompressed sketch is reconstructed to support the machine and human tasks, respectively. In this way, the color sketch is designed to bridge the gap between human and machine vision, and the joint training strategy helps to adjust the low-frequency information in the sketch. The experimental results on challenge datasets demonstrate that our proposed algorithm offers 40.9%-86.6% bitrate savings on machine vision and is comparable to state-of-the-art image coding methods on human vision. Yiping Duan, Qiyuan Du, Xiaoming Tao 0001, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Interpretable Multi-Modal Image Registration Network Based on Disentangled Convolutional Sparse CodingabstractMulti-modal image registration aims to spatially align two images from different modalities to make their feature points match with each other. Captured by different sensors, the images from different modalities often contain many distinct features, which makes it challenging to find their accurate correspondences. With the success of deep learning, many deep networks have been proposed to align multi-modal images, however, they are mostly lack of interpretability. In this paper, we first model the multi-modal image registration problem as a disentangled convolutional sparse coding (DCSC) model. In this model, the multi-modal features that are responsible for alignment (RA features) are well separated from the features that are not responsible for alignment (nRA features). By only allowing the RA features to participate in the deformation field prediction, we can eliminate the interference of the nRA features to improve the registration accuracy and efficiency. The optimization process of the DCSC model to separate the RA and nRA features is then turned into a deep network, namely Interpretable Multi-modal Image Registration Network (InMIR-Net). To ensure the accurate separation of RA and nRA features, we further design an accompanying guidance network (AG-Net) to supervise the extraction of RA features in InMIR-Net. The advantage of InMIR-Net is that it provides a universal framework to tackle both rigid and non-rigid multi-modal image registration tasks. Extensive experimental results verify the effectiveness of our method on both rigid and non-rigid registrations on various multi-modal image datasets, including RGB/depth images, RGB/near-infrared (NIR) images, RGB/multi-spectral images, T1/T2 weighted magnetic resonance (MR) images and computed tomography (CT)/MR images. The codes are available at https://github.com/lep990816/Interpretable-Multi-modal-Image-Registration. Xin Deng 0002, Enpeng Liu, Shengxi Li, Yiping Duan, Mai Xu |
IEEE Trans. Image Process. | 4 |
| 2022 | Measuring Human Perception of Audiovisual Errors using EEGabstractAudiovisual synchronization is an essential indicator of video quality. The degradation in the quality of experience caused by such synchronization errors is often measured using the mean opinion score (MOS). However, this method is susceptible to emotion and bias. Electroencephalography (EEG), as an objective tool, overcomes the drawbacks of subjective testing for evaluating human perception. In this work, we measure the human perception of audio and video synchronization based on the EEG approach through a series of experiments. A scoring model was developed for audio and video under multilevel synchronization errors by extracting the power spectral density (PSD) of EEG signals as features. The use of EEG signals for perceptual ability assessment of audiovisual distortion provides a potential neurally informed approach that is more objective than the high-level cognitive activity in subjective data. Dingcheng Gao, Bingrui Geng, Yiping Duan, Xiaoming Tao 0001, Chengkang Pan |
VTC Fall | 3 |
| 2022 | Image Generation from Scene Graph with Object EdgesabstractSignificant progress has been made on methods for generating images from structured semantic descriptions, but the generated images only retain semantic information, and the appearance of objects cannot be constrained and effectively represented. Therefore, we propose a scene graph structure image generation method assisted by object edge information. Our model uses two graph convolution neural networks(GCN) to process scene graphs and obtains object features as well as relation features which aggregate related information. The object bounding boxes are predicted by a method a decoupling the size and position. Where auxiliary models are added to coordinate with segmentation mask network training. Our experiments show that the introduction of object edges provides clearer object appearance information for image generation, which can constrain object shapes and improve image quality greatly. Finally, the cascaded refinement network is used to generate images. Additionally, compared with other appearance features, such as object slices, edge information occupies a smaller quantity of data, which greatly improves the image quality with less increase in the input information. This feature also benefits semantic communication systems. A large number of experiments show that our method is significantly superior to the latest Sg2im method when evaluated on Visual Genome datasets. Chenxing Li, Yiping Duan, Qiyuan Du, Chengkang Pan, Guangyi Liu 0001, Xiaoming Tao 0001 |
VTC Fall | 2 |
| 2022 | Multi-Relational Pedestrian Trajectory Prediction in Complex ScenesabstractPedestrian trajectory prediction has an important impact on the construction of smart cities and the popularization of autonomous vehicles. Pedestrian trajectory prediction in complex scenes is challenging because the trajectories are largely disturbed by the surrounding social environment. To better model the relationship between pedestrians and the social environment, we propose a novel Multi-Relation Network based on Long Short-Term Memory(LSTM). We first use Convolution Neural Network(CNN) and LSTM for feature extraction of scenes and pedestrians respectively. We realize the interaction between pedestrians and social environment by introducing attention mechanism. The experiment results on two public datasets, i.e. ETH and UCY, demonstrate that the prediction accuracy can be effectively improved by considering the social interactions. Wenshuo Peng, Zhoujuan Cui, Yiping Duan, Xiaoming Tao 0001 |
VTC Fall | 3 |
| 2022 | Path-based Multimodal Trajectories PredictionabstractPredicting the future movement of other traffic participants in complex traffic is one of the key tasks to realizing automatic driving. However, it is a challenge to accurately predict the future movement of the target, because the driving behavior is inherently random and multimodal, and the behavior of the agent is also determined by the complex interaction of other agents. The target-based method has a good performance in generating multi-modal trajectories. In this paper, a path-based multimodal trajectories prediction method is proposed, which takes the path as the target and performs two tasks path classification and trajectory regression. In this method, different paths will produce different embedding, and the sparse attention method based on path is adopted to deal with the interaction between agents, which reduces the complexity of the model and the risk of mode collapse. Experiments on the Argoverse dataset show that the proposed method has achieved good performance in all metrics, and its outstanding performance in the MR proves that the method can ensure the multimodality of output. Yiping Duan, Xiaoming Tao 0001 |
VTC Fall | 2 |
| 2022 | Human Perception Measurement by Electroencephalography for Facial Image CompressionabstractFacial images are the main focused contents in video conferences and many other applications. Therefore, it turns to a critical issue to measure and maintain the perceptual quality of facial images if transmitted over a bandwidth-limited communication system. In this letter, we propose a regional distortion perceptual threshold measurement model based on electroencephalography (EEG) to establish the relationship between image quality and human perception. Then, a facial image compression method is presented based on the model to improve the perceptual quality. Specifically, we construct a facial image dataset with regional distortion using the better portable graphics (BPG) compression. With the dataset, we design an EEG experiment and collect the brain responses to measure the human perception on regional distortions. By this method, a regional distortion perceptual threshold map (RDPTM) is constructed to guide the data rate allocation process for different regions of facial images. The experimental results show that our method can measure the human perception of regional distortion using EEG and improve the image perceptual quality by data rate allocation based on the RDPTM. Shuzhan Hu, Yiping Duan, Xiaoming Tao 0001, Geoffrey Ye Li, Jianhua Lu |
IEEE Signal Process. Lett. | 2 |
| 2021 | Brain-Inspired Image Quality Assessment Method based on Electroencephalography Feature LearningabstractWith the explosion of multimedia data, quality of experience (QoE) has become a critical metric in multimedia transmission, and therefore, QoE-oriented image quality assess-ment (IQA) turns more important and urgent. However, the performance of the traditional user-based assessment methods is limited by the deviation caused by human cognitive activities. In this paper, we propose a brain-inspired IQA method based on electroencephalography (EEG) feature learning, which is a psychophysiological method for studying human perception for IQA. We first establish the EEG dataset by collecting the corresponding EEG signals when subjects watch distorted facial images and then design a siamese network to extract the EEG features that can distinguish image quality levels and measure user scores. The siamese network establishes the relationship between image quality and QoE that is reflected by the EEG scores. The relationship is then embedded into a prediction network that directly obtains the EEG scores from images with different qualities. In this way, EEG scores can be predicted through end-to-end learning. Experiment results show that our proposed method can not only better evaluate the perceptual quality of facial images and reflect real human perceptions but also achieve better score prediction performance on the facial image datasets. Shuzhan Hu, Yiping Duan, Xiaoming Tao 0001, Geoffrey Ye Li, Jianhua Lu |
GLOBECOM | 2 |
| 2021 | Video Quality Measurement For Buffering Time Based On EEG Frequency FeatureabstractCurrently, user-based quality of experience (QoE) measurement methods (e.g., mean opinion score, MOS) are often employed. However, their results might be affected by human subjective experience and thoughts. Physiological measurement methods can overcome these disadvantages. In the field of video quality models, the video buffering problem caused by poor network conditions is an important factor that affects QoE. In this paper, a reasonable psychophysiological measurement method, electroencephalography (EEG), is proposed to quantitatively analyze QoE changes when users face different levels of video buffering time. By extracting the band power of EEG signals as the feature and analyzing the correlation and variance, a more objective buffering time EEG (BT-EEG) score model of video quality under the single-factor video buffering time is established, which solves the problem of uneven subjective data quality. Bingrui Geng, Xiaoming Tao 0001, Yiping Duan, Dingcheng Gao, Shuzhan Hu |
ICIP | 4 |
| 2021 | Eeg Based Visual Classification With Multi-Feature Joint LearningabstractWith a significant boost in neuroscience and artificial intelligence, decoding the process of human vision has become a hot topic in the last few decades. Although many existing deep learning models are employed to explore and solve mysteries of human brain activity, the accuracy and reliability of the visual classification task based on electroencephalography (EEG) still have space for promotion. In our research, we design the experiments to collect the subjects’ EEG data when they are watching the different types of images. In this way, an image-EEG dataset corresponding to 80 ImageNet object classes was constructed. Afterward, we proposed a dual-EEGNet for joint feature learning for multi-category visual classification. Especially, one branch EEGNet is used to extract the spatio-temporal embeddings of EEG signals, and the other branch is used to extract the time-frequency embeddings of EEG signals. The experimental results demonstrate that EEG signals can reflect the human brain activity and distinguish the different types of images. Moreover, the proposed model with joint features has a better classification performance in terms of accuracy compared with other methods. Yiping Duan, Shuzhan Hu, Xiaoming Tao 0001, Ning Ge 0001 |
ICIP | 2 |
| 2021 | Deep Coupled Feedback Network for Joint Exposure Fusion and Image Super-ResolutionabstractNowadays, people are getting used to taking photos to record their daily life, however, the photos are actually not consistent with the real natural scenes. The two main differences are that the photos tend to have low dynamic range (LDR) and low resolution (LR), due to the inherent imaging limitations of cameras. The multi-exposure image fusion (MEF) and image super-resolution (SR) are two widely-used techniques to address these two issues. However, they are usually treated as independent researches. In this paper, we propose a deep Coupled Feedback Network (CF-Net) to achieve MEF and SR simultaneously. Given a pair of extremely over-exposed and under-exposed LDR images with low-resolution, our CF-Net is able to generate an image with both high dynamic range (HDR) and high-resolution. Specifically, the CF-Net is composed of two coupled recursive sub-networks, with LR over-exposed and under-exposed images as inputs, respectively. Each sub-network consists of one feature extraction block (FEB), one super-resolution block (SRB) and several coupled feedback blocks (CFB). The FEB and SRB are to extract high-level features from the input LDR image, which are required to be helpful for resolution enhancement. The CFB is arranged after SRB, and its role is to absorb the learned features from the SRBs of the two sub-networks, so that it can produce a high-resolution HDR image. We have a series of CFBs in order to progressively refine the fused high-resolution HDR image. Extensive experimental results show that our CF-Net drastically outperforms other state-of-the-art methods in terms of both SR accuracy and fusion performance. The software code is available here https://github.com/ytZhang99/CF-Net. Xin Deng 0002, Mai Xu, Shuhang Gu, Yiping Duan |
IEEE Trans. Image Process. | 5 |
| 2021 | Semantic Perceptual Image Compression With a Laplacian Pyramid of Convolutional NetworksabstractThe existing image compression methods usually choose or optimize low-level representation manually. Actually, these methods struggle for the texture restoration at low bit rates. Recently, deep neural network (DNN)-based image compression methods have achieved impressive results. To achieve better perceptual quality, generative models are widely used, especially generative adversarial networks (GAN). However, training GAN is intractable, especially for high-resolution images, with the challenges of unconvincing reconstructions and unstable training. To overcome these problems, we propose a novel DNN-based image compression framework in this paper. The key point is decomposing an image into multi-scale sub-images using the proposed Laplacian pyramid based multi-scale networks. For each pyramid scale, we train a specific DNN to exploit the compressive representation. Meanwhile, each scale is optimized with different aspects, including pixel, semantics, distribution and entropy, for a good "rate-distortion-perception" trade-off. By independently optimizing each pyramid scale, we make each stage manageable and make each sub-image plausible. Experimental results demonstrate that our method achieves state-of-the-art performance, with advantages over existing methods in providing improved visual quality. Additionally, a better performance in the down-stream visual analysis tasks which are conducted on the reconstructed images, validates the excellent semantics-preserving ability of the proposed method. Juan Wang 0012, Yiping Duan, Xiaoming Tao 0001, Mai Xu, Jianhua Lu |
IEEE Trans. Image Process. | 2 |
| 2021 | Saliency Prediction on Omnidirectional Image With Generative Adversarial Imitation LearningabstractWhen watching omnidirectional images (ODIs), subjects can access different viewports by moving their heads. Therefore, it is necessary to predict subjects' head fixations on ODIs. Inspired by generative adversarial imitation learning (GAIL), this paper proposes a novel approach to predict saliency of head fixations on ODIs, named SalGAIL. First, we establish a dataset for attention on ODIs (AOI). In contrast to traditional datasets, our AOI dataset is large-scale, which contains the head fixations of 30 subjects viewing 600 ODIs. Next, we mine our AOI dataset and discover three findings: (1) the consistency of head fixations are consistent among subjects, and it grows alongside the increased subject number; (2) the head fixations exist with a front center bias (FCB); and (3) the magnitude of head movement is similar across the subjects. According to these findings, our SalGAIL approach applies deep reinforcement learning (DRL) to predict the head fixations of one subject, in which GAIL learns the reward of DRL, rather than the traditional human-designed reward. Then, multi-stream DRL is developed to yield the head fixations of different subjects, and the saliency map of an ODI is generated via convoluting predicted head fixations. Finally, experiments validate the effectiveness of our approach in predicting saliency maps of ODIs, significantly better than 11 state-of-the-art approaches. Our AOI dataset and code of SalGAIL are available online at https://github.com/yanglixiaoshen/SalGAIL. Mai Xu, Li Yang 0014, Xiaoming Tao 0001, Yiping Duan, Zulin Wang |
IEEE Trans. Image Process. | 4 |
| 2020 | Local-to-Global Semantic Supervised Learning for Image CaptioningabstractImage captioning is a challenging problem owing to the complexity of image content and the diverse ways of describing the content in natural language. Although current methods have made substantial progress in terms of objective metrics (such as BLEU, METEOR, ROUGE-L and CIDEr), there still exist some problems. Specifically, most of these methods are trained to maximize the log-likelihood or objective metrics. As a result, these methods often generate rigid and semantically incomplete captions. In this paper, we develop a new model that aims to generate captions conforming to human evaluation. The core idea is to use local-to-global semantic supervised learning by introducing the two-level optimization objective functions. At the word level, we match each word to the image regions using the local attention objective function; at the sentence level, we align the entire sentence and the image using the global semantic objective function. Experimentally, we compare the proposed model with current methods on MSCOCO dataset. We show that either local attention supervision or global semantic supervision is the necessary component for the success of our model through ablation studies. Furthermore, combining these two supervision objective functions achieves state-of-the-art performance in terms of both standard evaluation metrics and human judgment. Juan Wang 0012, Yiping Duan, Xiaoming Tao 0001, Jianhua Lu |
ICC | 2 |
| 2020 | Deep Learning Based Classification Using Semantic Information for Polsar ImageabstractThe deep learning has been applied to PolSAR image classification tasks in many researches. But they are pixel-based or spatial-based. None of them take the polarimetric semantic information into consideration. In this paper, a hierarchical classification method based on the polarimetric semantic information and deep learning is proposed. The method combines deep learning and semantic model to benefit from both the discriminative deep features and the polarimetric semantic information. The semantic information can provide the priori knowledge to improve the classification result of deep learning. Experimental results show that the proposed method can well preserve image details meanwhile suppress the speckle noises. Lu Zhang 0028, Wen Xie 0007, Feng Zhao 0005, Hanqiang Liu 0001, Yiping Duan |
IGARSS | 5 |
| 2020 | A Novel EEG Based Directed Transfer Function for Investigating Human Perception to Audio NoiseabstractAudio quality greatly affects users evaluation of multimedia communication, especially when the communication signal is disturbed, the noise in audio and video will decrease the quality of user experience. Psychophysiological indicators have high time resolution and precision, which can be used as important quality of experience characteristics. In this paper, electroencephalography is used as a psychophysiological method to assess brain connectivity in response to perceive the noise under different scenario. Specifically, we first record the response of the subjects' brainwaves to the audio quality using a high resolution electroencephalogram. Then, directed transfer function is used to analyze the directional information flow intensity between channels in the frequency domain, and 10% directed transfer function value are selected to construct the edge set of the directed graph to obtain the brain connectivity graph. Finally, the human perception to audio noise is obtained by using the weighted degree clustering method. In addition, the effectiveness of above results is verified by the small-world network coefficients experiment. Bingrui Geng, Yiping Duan, Qiwei Song, Xiaoming Tao 0001, Jianhua Lu, Jincheng Shi |
IWCMC | 3 |
| 2020 | EEG-Based Maritime Object Detection for IoT-Driven Surveillance Systems in Smart OceanabstractAutomated maritime object detection is a significant research challenge in intelligent marine surveillance systems for the Internet of Things (IoT) and smart ocean applications. In particular, ship detection is recognized as one of the core research issues of these IoT-driven intelligent marine surveillance systems. Traditional methods based on machine learning have made some achievements in detection tasks for specific objects. However, the ship objects are relatively small, and they are usually not accurately detected. In this article, we propose an electroencephalography (EEG)-based maritime object detection algorithm for IoT-driven surveillance systems in the smart ocean. For this purpose, we conduct experiments to record the EEG signals of subjects when they are watching the maritime image scenes. With the feature analysis of EEG signals, the event-related potential (ERP) components associated with detecting objects are induced, such as the$P3$and$N2$components. Employing classification based on linear discriminant analysis (LDA), the area under curve (AUC) of the receiver operating characteristic (ROC) is used to evaluate the detection accuracy. We use this novel method to determine and identify essential objects and areas from IoT devices, such as digital camera imaging sensors. Our proposed method can not only help to detect small objects accurately using fewer samples but can also be used to reduce the data volume needed to be stored and transmitted in IoT-driven marine surveillance systems. Yiping Duan, Xiaoming Tao 0001, Qiang Li 0035, Shuzhan Hu, Jianhua Lu |
IEEE Internet Things J. | 1 |
| 2020 | Toward Variable-Rate Generative Compression by Reducing the Channel RedundancyabstractCompressing large images with a generative model goes beyond typical image encoding standards under a notably low bitrate. In this paper, we step toward practical generative compression systems based on recent advances. Specifically, we show that the channel redundancy of the latent representation produced by an autoencoder network can be effectively compressed via mask compression. The mask compression performs quantization on the channel variance of latent representation instead of original values. Instead of training multiple models, changing the mask leads to a simple and efficient variable rate compression scheme. Then, we estimate the relative bitrate by measuring the L1 norm of the channel variance and hence obtain the rate-distortion formulation. The L1 regularizer assumes a Laplacian prior on the channel variance, through which model we develop corresponding methods to produce approximate images at a target bitrate. This eliminates the need for manually searching hyperparameters for our variable-rate compression. We conduct exhaustive experiments to demonstrate the advanced performance of the proposed method in preserving image quality and semantics. Chaoyi Han, Yiping Duan, Xiaoming Tao 0001, Mai Xu, Jianhua Lu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | PolSAR image classification based on multi-scale stacked sparse autoencoder
Lu Zhang 0028, Licheng Jiao, Wenping Ma 0001, Yiping Duan |
Neurocomputing | 4 |
| 2019 | Learning QoE of Mobile Video Transmission With Deep Neural Network: A Data-Driven ApproachabstractQuality of experience (QoE) serves as a direct evaluation of users' experiences in mobile video transmission and thus essential for network management, such as network optimization. In this paper, we propose a deep learning-based QoE prediction approach with a large-scale QoE dataset for mobile video transmission. Specifically, we develop a mobile phone application for collecting user QoE data when viewing videos transmitted over the mobile internet in a practical environment. Then, we construct a large-scale dataset by collecting over 80000 piece of data with four kinds of subjective scores and 89 network parameters. Each QoE metric is related to only some of the 89 network parameters. Therefore, we apply the feature selection method to find the feature parameters related to user scores. Additionally, the boxplot method is used to clean the raw data by removing outliers. Finally, a deep neural network (DNN) is developed to learn the relationships between the network parameters and the subjective QoE scores. The proposed DNN can also be seen as a data-driven objective QoE prediction approach for mobile video transmission, which can be used to predict the user QoE scores. The experimental results show that the proposed approach can effectively remove most features irrelevant to QoE prediction. Moreover, the performance of QoE prediction by the proposed model outperforms other state-of-the-art approaches. Xiaoming Tao 0001, Yiping Duan, Mai Xu, Zhishen Meng, Jianhua Lu |
IEEE J. Sel. Areas Commun. | 2 |
| 2018 | Multi-Scale Convolutional Neural Network for SAR Image Semantic SegmentationabstractAlthough the recent success of convolutional neural networks (CNNs) greatly advance the semantic segmentation of the natural images, few work has focused on the remote sensing images, especially the synthetic aperture radar (SAR) images. Specifically, the existing methods do not consider the speckle noise of the SAR images and the multi-scale characteristics contained in the SAR images. In this paper, we propose a multiscale convolutional neural network (CNN) model for SAR image semantic segmentation. The multi-scale CNN model includes noise removal stage, convolutional stage, feature concatenation stage and classification stage. In particular, we construct a sparse representation loss function to obtain a clear SAR image in noise removal stage. Then, the multi-scale convolutional stage is employed to learn the multi-scale deep features. The concatenation stage is used to connect the features with different scales and depths. Finally, softmax classifier is developed to obtain the labels of the SAR images with the multi-scale CNN model being trained in an end-to-end way. The experimental results on synthetic and real SAR images demonstrate the effectiveness of the proposed method. Yiping Duan, Xiaoming Tao 0001, Chaoyi Han, Xiaowci Qin, Jianhua Lu |
GLOBECOM | 1 |
| 2018 | Semantic Conditional Random Field for Object Based SAR Image SegmentationabstractConditional random filed (CRF) model relaxes the conditional independence of the observed data and simultaneously captures the spatial contextual information. However, the single spatial contextual model is difficult to describe the heterogeneous structures of the synthetic aperture radar (SAR) images. This paper propose an semantic conditional random field (SCRF), which integrate the semantic space and pixel space for object-based SAR image segmentation. Specifically, the SAR image is divided into aggregated, structural and homogeneous subspaces by using the hierarchical semantic model. Then, we design gaussian kernel function, geometric kernel function and uniform kernel function to adaptively describe the spatial contextual constraints in the different subspaces. These kernel functions are incorporated into the pairwise potential of CRF model to improve the ability of the model. Afterwards, the piecewise training and Bayesian inference are proposed to achieve the object-based segmentation. Experiments on the synthetic and real SAR images demonstrate the effectiveness of the proposed method in the semantic consistency and detail preservations. Yiping Duan, Xiaoming Tao 0001, Chaoyi Han, Jianhua Lu |
ICIP | 1 |
| 2018 | Dense Convolution for Semantic SegmentationabstractState-of-the-art semantic segmentation methods adopt fully convolutional neural networks (FCNs) to solve this dense prediction problem. However, replacing fully connected layers with the standard 2D convolution layer is straightforward yet not optimal in generating segmentation results. In this paper we develop a dense convolution scheme that is more suitable for semantic segmentation. Instead of generating a single output, dense convolution produces the same number of output as its input and introduces spatial overlaps into current convolutions. Then each activation is obtained from multiple overlapped dense convolutions with learnable weights. Such dense convolution helps to reinforce local connections between activations and provide more flexible receptive fields for predictions. Experiments on benchmark dataset demonstrate the effectiveness of the proposed approach in semantic segmentation tasks. Chaoyi Han, Xiaoming Tao 0001, Yiping Duan, Jianhua Lu |
ICIP | 3 |
| 2018 | Hyperspectral Image Denoising via Nonnegative Matrix Factorization and Convolutional Neural NetworksabstractHyperspectral image (HSI) denoising plays an important role to enhance the image quality for subsequent applications. This paper proposes a novel denoising framework for the HSI, in which the denoising issue is modeled as a convolutional neural network (CNN) constrained non-negative matrix factorization problem. Then, by adopting the proximal alternating linearized minimization, the proposed approach can be decomposed into two iterative steps: In the first step, the spectral matrix is updated using a proximal operator to guarantee the non-negativity; in the second step, the abundance matrix is updated using a designed CNN. Finally, we propose a transfer learning scheme to train the design CNN with a natural gray image dataset. Exhaustive experiments show that the proposed approach outperforms the comparison state-of-the-art methods on two different test HSI datasets. Baihong Lin, Xiaoming Tao 0001, Xiaowei Qin, Yiping Duan, Jianhua Lu |
IGARSS | 4 |
| 2018 | Visual information assisted UAV positioning using priori remote-sensing information
Xijia Liu, Xiaoming Tao 0001, Yiping Duan, Ning Ge 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Hierarchical objectness network for region proposal generation and object detection
Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Yiping Duan, Jianhua Lu |
Pattern Recognit. | 4 |
| 2018 | Adaptive Hierarchical Multinomial Latent Model With Hybrid Kernel Function for SAR Image Semantic SegmentationabstractSynthetic aperture radar (SAR) images have been one of the important tools to support earth observations and topographic measurements. It means that SAR images are essentially rich in structures. However, the single spatial relationship is difficult to deal with the heterogeneous structures of the SAR images. In this paper, we propose an adaptive hierarchical multinomial latent model with hybrid kernel function for SAR image semantic segmentation. In the proposed approach, we design a hybrid kernel function combing Gaussian radial basis function (GRBF) and ridgelet kernel function to adaptively describe the spatial relationships between the central pixel and the surrounding pixels. Then, based on the hybrid kernel function, adaptive methods are proposed for semantic segmentation. Specifically, an SAR image is divided into different characteristics subspaces, homogeneous, structural, and aggregated subspaces, by SAR hierarchical semantic model. For the homogeneous subspace, GRBF is used to describe the isotropic spatial relationships. Then, multilayer multinomial latent model with GRBF is used for segmentation to improve the labeling consistency and reduce the wrong segmentation. For the structural subspace, the ridgelet kernel function is used to describe the anisotropic spatial relationships. Then, we adopt the single-layer multinomial latent model with ridgelet kernel function for segmentation to preserve the details (such as edge, lines, and small objects). For aggregated subspace, bag-of-words model is used to extract the features of the aggregated portions, and then affinity propagation cluster is used for segmentation. Finally, the segmentation results of different subspaces are integrated together to obtain the final segmentation result. Comprehensive experiments on both synthetic and real SAR images demonstrate that the segmentation results by our proposed approach achieve the semantic consistency, labeling consistency, and detail preservation simultaneously. Yiping Duan, Fang Liu 0001, Licheng Jiao, Xiaoming Tao 0001, Jie Wu 0016, Cheng Shi 0002, Martin O. Wimmers |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | SAR Image segmentation based on convolutional-wavelet neural network and markov random field
Yiping Duan, Fang Liu 0001, Licheng Jiao, Lu Zhang 0028 |
Pattern Recognit. | 1 |
| 2016 | Sketching Model and Higher Order Neighborhood Markov Random Field-Based SAR Image SegmentationabstractThe Markov random field (MRF) model has been successfully applied to synthetic aperture radar (SAR) image segmentation because of its excellent ability of capturing the local contextual information in the prior model. However, the geometric structures of the SAR image are always ignored when capturing the contextual information in the prior model. Therefore, this letter presents a new SAR image segmentation method based on the sketching model and higher order neighborhood MRF. In this approach, the sketching model is utilized to represent the geometric structures of the SAR image. Meanwhile, a higher order neighborhood is constructed to capture the complex priors. Then, according to the structure fluctuation in the higher order neighborhood, the homogeneous and heterogeneous neighborhoods are distinguished. Finally, the local energy function in the prior model is constructed in the higher order neighborhood with different characteristics. Specifically, the energy functions considering the labeling consistency and focusing on the structure preservations are designed for the homogeneous and heterogeneous neighborhoods, respectively. In this way, the ability of the prior model is improved by adding the geometric structures into the energy functions. Experiments on the real SAR images demonstrate the effectiveness of the proposed method in labeling consistency and structure preservations. Yiping Duan, Fang Liu 0001, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2016 | SAR Image Segmentation Based on Hierarchical Visual Semantic and Adaptive Neighborhood Multinomial Latent ModelabstractA synthetic aperture radar (SAR) imaging system usually produces pairs of bright area and dark area when depicting the ground objects, such as a building or tree and its shadow. Many buildings (trees) are aggregated together to form urban areas (forests). It means that the pairs of bright and dark areas often exist in the aggregated scenes. Conventional unsupervised segmentation approaches usually segment the scenes (e.g., urban areas and forests) into different regions simply according to the gray values of the image. However, a more convincing way is to regard them as the consistent regions. In this paper, we aim at addressing this issue and propose a new SAR image segmentation approach via a hierarchical visual semantic and adaptive neighborhood multinomial latent model. In this approach, the hierarchical visual semantic of SAR images is proposed, which divides SAR images into aggregated, structural, and homogeneous regions. Based on the division, different segmentation methods are chosen for these regions with different characteristics. For the aggregated region, locality-constrained linear coding-based hierarchical clustering is used for segmentation. For the structural region, visual semantic rules are designed for line object location, and a geometric structure window-based multinomial latent model is proposed for segmentation. For the homogeneous region, a multinomial latent model with adaptive window selection is proposed for segmentation. Finally, these results are integrated together to obtain the final segmentation. Experiments on both synthetic and real SAR images indicate that the proposed method achieves promising performances in terms of the consistencies of the regions and the preservations of the edges and line objects. Fang Liu 0001, Yiping Duan, Lingling Li 0002, Licheng Jiao, Jie Wu 0016, Shuyuan Yang 0001, Xiangrong Zhang, Jialing Yuan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | Learning Interpolation via Regional Map for Pan-SharpeningabstractAlthough the bandwidth of the high-resolution panchromatic (HR PAN) image is wide, it is narrow in each band of the low-resolution multispectral (LR MS) image. Hence, the spatial resolution of the HR PAN image is much higher than that of the LR MS image. However, HR PAN image only has a single band. The purpose of the Pan-sharpening algorithm is to make the Pan-sharpened image with both high spatial resolution and good spectral information. In this paper, a novel learning interpolation method for Pan-sharpening is proposed by expanding the sketch information in the HR PAN image. The sketch information contains the edges and lines features of the image, and each segment of the sketch information has its own direction. According to the primal sketch graph of the HR PAN image, a regional map is obtained by a designed geometrical template. Since the size of the HR PAN image is different from that of the LR MS image, the LR MS image is interpolated into an interpolated multispectral (IMS) image by the nearest interpolation method. In addition, the IMS image can be mapped into the structure and the nonstructure regions by this regional map. The nonstructure regions are divided into the smooth and the texture regions by a variance value. For the structure and texture regions, the interpolated pixels in the IMS image are relearned and readjusted by the proposed structure and texture learning interpolation method, respectively. Experimental results show that the proposed Pan-sharpening method can provide superior performance in both visual effect and quality metrics, particularly for the images with a large spectral difference. Cheng Shi 0002, Fang Liu 0001, Lingling Li 0002, Licheng Jiao, Yiping Duan, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |