VLDB 2026 Research / reviewers in the wild / expert
Qiyuan Du
dblp:326/1302
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-3966-080XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic encoding for image compression based on semantic segmentation maps
Zhipeng Xie, Yiping Duan, Wei Kun Kong, Qiyuan Du, Qinghua Liang, Xiaoming Tao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Lite3D: A Lightweight Hybrid CNN-Transformer Framework for 3D Video StabilizationabstractVideo stabilization plays a critical role in IoT, as it enhances the accuracy and reliability of visual data collected from mobile and unstable devices, enabling more effective monitoring and analysis. However, existing stabilization methods frequently lack sufficient accuracy in motion estimation or struggle with expensive computation. In this paper, we introduce Lite3D, a novel lightweight 3D deep learning-based video stabilization framework that combines a hybrid CNN and Transformer architecture for more efficient feature extraction. During training, our method utilizes estimated depth maps and relative camera poses to generate target views, jointly optimizing DepthNet and PoseNet through an unsupervised learning strategy. Compared to existing 3D video stabilization methods, our approach enables more accurate learning of motion information. In the inference phase, we smooth the estimated camera pose trajectory and synthesize stabilized frames using the optimal depth maps. Experimental results on the NUS dataset demonstrate that Lite3D achieves state-of-the-art stability, with additional tests on the DeepStab dataset confirming strong generalization. Moreover, our method contains fewer parameters compared to existing 3D deep learning-based video stabilization, offering both high stabilization quality and computational efficiency. Zejing Shan, Yiping Duan, Yue Wu 0004, Qiyuan Du, Xiaoming Tao 0001 |
ICC | 4 |
| 2025 | Cloud-Edge-End Collaborative Surveillance Video Transmission with Object-Guided Video Super-ResolutionabstractCloud Video Surveillance (CVS) systems, as the backbone of distributed surveillance networks, face increasing challenges in transmitting large volumes of high-resolution video data. While cloud-end collaborative video transmission methods can reduce bitrates beyond conventional compression techniques, the periodic transmission of high-resolution keyframes consumes significant bandwidth due to semantically irrelevant background information. To address this, we propose a cloud-edge-end collaborative video transmission scheme based on object-guided video super-resolution. In this scheme, the edge extracts key objects from sparsely selected keyframes at the end and transmits them along with low-resolution video to the cloud, where our Keyframe-Guided Video Restoration Transformer (KG-VRT) is used to improve the video quality. Experimental results on public datasets show that our network outperforms state-of-the-art keyframe-based baselines with a 1.73 dB PSNR improvement and maintains robust performance even with keyframe intervals of up to 30 frames. A comparative analysis of two transmission strategies—transmitting full keyframes versus transmitting only key object regions—demonstrates a 60% – 80% reduction in keyframe bitrate while maintaining object detection accuracy at a significantly reduced overall system bitrate. This highlights the efficiency and scalability of our approach in bandwidth-constrained surveillance scenarios. Yiping Duan, Xiaoming Tao 0001, Wei Kun Kong, Qiyuan Du |
VTC2025-Fall | 5 |
| 2025 | Object-Attribute-Relation Representation-Based Video Semantic CommunicationabstractWith the rapid growth of multimedia data volume, there is an increasing need for efficient video transmission in applications such as virtual reality and future video streaming services. Semantic communication is emerging as a vital technique for ensuring efficient and reliable transmission in low-bandwidth, high-noise settings. However, most current approaches focus on joint source-channel coding (JSCC) that depends on end-to-end training. These methods often lack an interpretable semantic representation and struggle with adaptability to various downstream tasks. In this paper, we introduce the use of object-attribute-relation (OAR) as a semantic framework for videos to facilitate low bit-rate coding and enhance the JSCC process for more effective video transmission. We utilize OAR sequences for both low bit-rate representation and generative video reconstruction. Additionally, we incorporate OAR into the image JSCC model to prioritize communication resources for areas more critical to downstream tasks. Our experiments on traffic surveillance video datasets assess the effectiveness of our approach in terms of video transmission performance. The empirical findings demonstrate that our OAR-based video coding method not only outperforms H.265 coding at lower bit-rates but also synergizes with JSCC to deliver robust and efficient video transmission. Qiyuan Du, Yiping Duan, Qianqian Yang 0002, Xiaoming Tao 0001, Mérouane Debbah |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | Data-Free Cloud-Edge Distillation for Safe and Efficient Intelligent CommunicationsabstractEfficiency and security are the core challenges in the intelligent communication field. Lightweight neural networks have accelerated the information interpretation and communication efficiency, thereby fostering the rapid development of the Internet of Things. Enhancing the recognition capability of lightweight neural networks remains challenging. Knowledge distillation, a technique that transfers knowledge from a complex model to a smaller one, is often used to improve the recognition performance of lightweight networks. However, practical issues such as transmission constraints and user privacy make the original data required for knowledge distillation difficult to access directly. To tackle this issue, this paper proposes a Data-Free Cloud-Edge Knowledge Distillation (DF-CEKD) model, which uses a complex network in the cloud to provide training guidance for lightweight networks deployed on mobile devices. Specifically, DF-CEKD employs a novel Deep Inversion Diffusion Generation (DIDG) module to provide proxy data as input for the distillation process, thereby transferring the feature learning capability from the cloud network to the edge network. Meanwhile, a Multi-Layer Feature Joint Supervision Distillation (MLF-JSD) module is designed to further enhance the feature selection guidance provided by the teacher network in the cloud for training the lightweight student network. The simulation results demonstrate that the proposed DF-CEKD reduces the number of parameters to 1/20 and the floating-point operations to 1/12 when distilling from WRN40-2 to WRN16-1, resulting in only a 0.27% decrease in accuracy. Xiufang Li, Yiping Duan, Xiaoming Tao 0001, Qigong Sun, Qiyuan Du, Qianqian Yang 0002, Tony Q. S. Quek |
IEEE J. Sel. Areas Commun. | 5 |
| 2024 | Semantic Security: A Digital Watermark Method for Image Semantic PreservationabstractDigital watermarking has long been used to protect digital images from abuse. However, applying digital watermarking to semantic communication remains a challenge. This work introduces a secure coding method that combines semantic coding and digital watermarking techniques. The proposed method selects points in the target image with high semantic importance and embeds watermark information into their position vectors, perpendicular to the embedding domains of previous works. The experiments conducted on the Cityscapes dataset demonstrate that our method integrates well with semantic communication systems. Compared to the previous approach, our proposed method can more completely preserve the structural features of the target image and is better suited for Machine Type Communication (MTC) tasks, such as target detection and semantic segmentation. Tianwei Zuo, Yiping Duan, Qiyuan Du, Xiaoming Tao 0001 |
ICASSP | 3 |
| 2024 | Optical Flow-Based Spatiotemporal Sketch for Video Representation: A Novel FrameworkabstractWith the rapid development of multimedia services and the dramatic growth of video data volume, efficient video representation and AI-generated content (AIGC) become critical parts of future multimedia communication systems. Sketch graph is a structured abstraction of key textures in an image, and video sketch graph further exploits the temporal continuity of videos to achieve a sparse representation. Sketch-based representation has potential applications in communication systems for both human subjective perception and machine vision tasks, and provides a new idea for AIGC. However, current video sketch extraction methods rely on human assistance and correction, and cannot be applied to end-to-end communication systems. We design a novel framework for spatiotemporal sketch extraction based on deep learning methods. In the proposed framework, sketch extraction and sparse coding are performed at the sender side using structural and temporal features of the video. The original videos are generatively reconstructed at the receiver side or applied to downstream machine vision tasks. We validate the performance of the proposed method on Cityscapes dataset with different metrics. Experiments show that our proposed framework can be end-to-end adapted to video communication tasks in different scenarios and can achieve efficient video characterization and transmission. Moreover, our proposed method enables sketch-based end-to-end AIGC for video generation. Qiyuan Du, Yiping Duan, Zhipeng Xie, Xiaoming Tao 0001, Linsu Shi, Zhijuan Jin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Sketch Graph Representation for Multimedia Computational Communications: A Learning-Based MethodabstractMultimedia computational communications towards 6G can improve the transmission efficiency significantly by introducing intelligent computation in the communication process. This intelligent/smart communication architecture includes multimedia representation, coding, transmission and other parts from the perspective of semantics, where multimedia semantic representation is the core part and is mainly utilized to reduce the amount of multimedia data. In this paper, sketch graph is proposed as an effective representation of images to describe the pixel variations, geometric feature distribution and structural information and has potential applications in multimedia computational communications. Specifically, we developed a learning-based method to extract sketch graphs with edge detection, sketch point detection and sketch line detection by deep neural networks (DNNs). Moreover, we designed an end-to-end extraction method and achieved real-time processing. The experimental results on several datasets demonstrated the advanced performance in terms of classification and generation tasks. Image classification results on the HumanSketch, ImageNet, and Caltech datasets showed that sketch graphs extracted by our method had better describing ability than those extracted by traditional methods and other traditional image compression methods. On the other hand, image generation results on the Cityscapes dataset indicated the potential of the sketch-graph-based image compression codec. Qiyuan Du, Yiping Duan, Xiaoming Tao 0001, Chengkang Pan, Guangyi Liu 0001 |
ICC | 1 |
| 2023 | Video Reconstruction with Multimodal InformationabstractVideo reconstruction refers to generate videos through the high-level representations (edge map, labels and so on), while the reconstruction quality is always unsatisfactory due to sparse high-level representations, especially on video data. In order to improve the video reconstruction quality, we proposed a novel approach that generates realistic video from its multimodal information including structure features and color features. To extract color features, we mainly apply the k-means algorithm to segment labels and the structure features are extracted by an edge detection network. Video generation is regarded as learning the mapping from multimodal representations to the original videos. So, a conditional GAN is applied with a learning objective that models the temporal video dynamics. We use a spatio-temporal generator with attention to model the inter-frame dynamics and video consistency is improved in this way. Moreover, we use a multiscale discriminator to improve the improve the intra-frame quality of the video. Experimental results on Cityscapes, Apolloscape datasets demonstrate that our proposed approach performs better in both traditional and generative evaluating indicators. Zhipeng Xie, Yiping Duan, Qiyuan Du, Xiaoming Tao 0001, Jiazhong Yu |
VTC Fall | 3 |
| 2023 | Toward Semantic Communications: Deep Learning-Based Image Semantic CodingabstractSemantic communications has received growing interest since it can remarkably reduce the amount of data to be transmitted without missing critical information. Most existing works explore the semantic encoding and transmission for text and apply techniques in Natural Language Processing (NLP) to interpret the meaning of the text. In this paper, we conceive the semantic communications for image data that is much more richer in semantics and bandwidth sensitive. We propose an reinforcement learning based adaptive semantic coding (RL-ASC) approach that encodes images beyond pixel level. Firstly, we define the semantic concept of image data that includes the category, spatial arrangement, and visual feature as the representation unit, and propose a convolutional semantic encoder to extract semantic concepts. Secondly, we propose the image reconstruction criterion that evolves from the traditional pixel similarity to semantic similarity and perceptual performance. Thirdly, we design a novel RL-based semantic bit allocation model, whose reward is the increase in rate-semantic-perceptual performance after encoding a certain semantic concept with adaptive quantization level. Thus, the task-related information is preserved and reconstructed properly while less important data is discarded. Finally, we propose the Generative Adversarial Nets (GANs) based semantic decoder that fuses both locally and globally features via an attention module. Experimental results demonstrate that the proposed RL-ASC is noise robust and could reconstruct visually pleasant and semantic consistent image in low bit rate condition. Danlan Huang, Feifei Gao 0001, Xiaoming Tao 0001, Qiyuan Du, Jianhua Lu |
IEEE J. Sel. Areas Commun. | 4 |
| 2023 | Sketch Assisted Face Image Coding for Human and Machine Vision: A Joint Training ApproachabstractImage coding is one of the most fundamental techniques and is widely used in image/video processing and multimedia communications. Current image coding methods are mainly human-oriented, and the visual quality is always unsatisfactory, especially at low bitrates. Moreover, the recent emergence of machine vision goes beyond the scope of current coding. With these considerations, we proposed a sketch assisted face image coding for human and machine vision by a joint training approach. In the proposed approach, we design a new feature representation: a color sketch, which aims to satisfy both low-frequency features of human vision and high-frequency features of machine analysis. Then, we present a novel end-to-end image codec framework with joint training that consists of three models: an image-to-image translation module, a coding module, and a two-stage reconstruction module. Specifically, the input image is first translated into the edge map with the Canny edge as the auxiliary label to merely preserve the structure information. Afterward, the backpropagation from reconstruction module guides the edge map to increase or decrease the information through joint training, which results in the generation of color sketch. Then, the generated sketch is compressed into the bitstream and decompressed back to a sketch in the coding module. Finally, the decompressed sketch is reconstructed to support the machine and human tasks, respectively. In this way, the color sketch is designed to bridge the gap between human and machine vision, and the joint training strategy helps to adjust the low-frequency information in the sketch. The experimental results on challenge datasets demonstrate that our proposed algorithm offers 40.9%-86.6% bitrate savings on machine vision and is comparable to state-of-the-art image coding methods on human vision. Yiping Duan, Qiyuan Du, Xiaoming Tao 0001, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Image Generation from Scene Graph with Object EdgesabstractSignificant progress has been made on methods for generating images from structured semantic descriptions, but the generated images only retain semantic information, and the appearance of objects cannot be constrained and effectively represented. Therefore, we propose a scene graph structure image generation method assisted by object edge information. Our model uses two graph convolution neural networks(GCN) to process scene graphs and obtains object features as well as relation features which aggregate related information. The object bounding boxes are predicted by a method a decoupling the size and position. Where auxiliary models are added to coordinate with segmentation mask network training. Our experiments show that the introduction of object edges provides clearer object appearance information for image generation, which can constrain object shapes and improve image quality greatly. Finally, the cascaded refinement network is used to generate images. Additionally, compared with other appearance features, such as object slices, edge information occupies a smaller quantity of data, which greatly improves the image quality with less increase in the input information. This feature also benefits semantic communication systems. A large number of experiments show that our method is significantly superior to the latest Sg2im method when evaluated on Visual Genome datasets. Chenxing Li, Yiping Duan, Qiyuan Du, Chengkang Pan, Guangyi Liu 0001, Xiaoming Tao 0001 |
VTC Fall | 3 |