Yong Du 0003

dblp:73/3980-3 · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0003-1864-8326ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 16 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Invert Your Prompt: Editing-Aware Diffusion Inversion
Yangyang Xu 0003, Wenqi Shao, Yong Du 0003, Haiming Zhu, Yang Zhou 0038, Jiayuan Xie, Ping Luo 0002, Shengfeng He
Int. J. Comput. Vis.3
2026 AvatarVTON: 4D Virtual Try-On for Animatable Avatars
abstract
We propose AvatarVTON, the first 4D virtual try-on framework that generates realistic try-on results from a single in-shop garment image, enabling free pose control, novel-view rendering, and diverse garment choices. Unlike existing methods, AvatarVTON supports dynamic garment interactions under single-view supervision, without relying on multi-view garment captures or physics priors. The framework consists of two key modules: (1) a Reciprocal Flow Rectifier, an optical-flow-based correction strategy without external priors that stabilizes avatar fitting and ensures temporal coherence; and (2) a Non-Linear Deformer, which decomposes Gaussian maps into view-pose-invariant and view-pose-specific components, enabling adaptive, non-linear garment deformations. To establish a benchmark for 4D virtual try-on, we extend existing baselines with unified modules for fair qualitative and quantitative comparisons. Extensive experiments show that AvatarVTON achieves high fidelity, diversity, and dynamic garment realism, making it well-suited for AR/VR, gaming, and digital-human applications.
Zicheng Jiang, Jixin Gao, Shengfeng He, Xinzhe Li 0003, Yulong Zheng, Zhaotong Yang, Junyu Dong, Yong Du 0003
IEEE Trans. Vis. Comput. Graph.8
2025 PersonaMagic: Stage-Regulated High-Fidelity Face Customization with Tandem Equilibrium
abstract
Personalized image generation has made significant strides in adapting content to novel concepts. However, a persistent challenge remains: balancing the accurate reconstruction of unseen concepts with the need for editability according to the prompt, especially when dealing with the complex nuances of facial features. In this study, we delve into the temporal dynamics of the text-to-image conditioning process, emphasizing the crucial role of stage partitioning in introducing new concepts. We present PersonaMagic, a stage-regulated generative technique designed for high-fidelity face customization. Using a simple MLP network, our method learns a series of embeddings within a specific timestep interval to capture face concepts. Additionally, we develop a Tandem Equilibrium mechanism that adjusts self-attention responses in the text encoder, balancing text description and identity preservation, improving both areas. Extensive experiments confirm the superiority of PersonaMagic over state-of-the-art methods in both qualitative and quantitative evaluations. Moreover, its robustness and flexibility are validated in non-facial domains, and it can also serve as a valuable plug-in for enhancing the performance of pretrained personalization models.
Xinzhe Li 0003, Jiahui Zhan, Shengfeng He, Yangyang Xu 0003, Junyu Dong, Huaidong Zhang, Yong Du 0003
AAAI7
2025 NexusGS: Sparse View Synthesis with Epipolar Depth Priors in 3D Gaussian Splatting
abstract
Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-based approach that enhances novel view synthesis from sparse-view images by directly embedding depth information into point clouds, without relying on complex manual regularizations. Exploiting the inherent epipolar geometry of 3DGS, our method introduces a novel point cloud densification strategy that initializes 3DGS with a dense point cloud, reducing randomness in point placement while preventing over-smoothing and overfitting. Specifically, NexusGS comprises three key steps: Epipolar Depth Nexus, Flow-Resilient Depth Blending, and Flow-Filtered Depth Pruning. These steps leverage optical flow and camera poses to compute accurate depth maps, while mitigating the inaccuracies often associated with optical flow. By incorporating epipolar depth priors, NexusGS ensures reliable dense point cloud coverage and supports stable 3DGS training under sparse-view conditions. Experiments demonstrate that NexusGS significantly enhances depth accuracy and rendering quality, surpassing state-of-the-art methods by a considerable margin. Furthermore, we validate the superiority of our generated point clouds by substantially boosting the performance of competing methods. Project page: https://usmizuki.github.io/NexusGS/.
Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du 0003
CVPR7
2025 Cross-Subject Mind Decoding from Inaccurate Representations
Yangyang Xu 0003, Bangzhen Liu, Wenqi Shao, Yong Du 0003, Shengfeng He
ICCV4
2025 OmniVTON: Training-Free Universal Virtual Try-On
abstract
Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve fine-grained texture retention. For precise pose alignment, we utilize DDIM inversion to capture structural cues while suppressing texture interference, ensuring accurate body alignment independent of the original image textures. By disentangling garment and pose constraints, OmniVTON eliminates the bias inherent in diffusion models when handling multiple conditions simultaneously. Experimental results demonstrate that OmniVTON achieves superior performance across diverse datasets, garment types, and application scenarios. Notably, it is the first framework capable of multi-human VTON, enabling realistic garment transfer across multiple individuals in a single scene. Code is available at https://github.com/Jerome-Young/OmniVTON
Zhaotong Yang, Shengfeng He, Xinzhe Li 0003, Yangyang Xu 0003, Junyu Dong, Yong Du 0003
ICCV7
2025 Stable Score Distillation
abstract
Text-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieves cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's structure and ensures that editing trajectories are closely aligned with the source prompt, enabling smooth, prompt-specific modifications while maintaining coherence in surrounding regions. Additionally, SSD incorporates a prompt enhancement branch to boost editing strength, particularly for style transformations. Our method achieves state-of-the-art results in 2D and 3D editing tasks, including NeRF and text-driven style edits, with faster convergence and reduced complexity, providing a robust and efficient solution for text-guided editing.
Haiming Zhu, Yangyang Xu 0003, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du 0003, Jun Yu 0002, Shengfeng He
ICCV6
2025 DiffusionMat: Alpha Matting as Deterministic Sequential Refinement Learning
Yangyang Xu 0003, Shengfeng He, Wenqi Shao, Yong Du 0003, Kwan-Yee Kenneth Wong, Yu Qiao 0001, Jun Yu 0002, Ping Luo 0002
ACM Multimedia4
2025 One-for-All: Towards Universal Domain Translation With a Single StyleGAN
abstract
In this paper, we propose a novel translation model, UniTranslator, for transforming representations between visually distinct domains under conditions of limited training data and significant visual differences. The main idea behind our approach is leveraging the domain-neutral capabilities of CLIP as a bridging mechanism, while utilizing a separate module to extract abstract, domain-agnostic semantics from the embeddings of both the source and target realms. Fusing these abstract semantics with target-specific semantics results in a transformed embedding within the CLIP space. To bridge the gap between the disparate worlds of CLIP and StyleGAN, we introduce a new non-linear mapper, the CLIP2P mapper. Utilizing CLIP embeddings, this module is tailored to approximate the latent distribution in the StyleGAN's latent space, effectively acting as a connector between these two spaces. The proposed UniTranslator is versatile and capable of performing various tasks, including style mixing, stylization, and translations, even in visually challenging scenarios across different visual domains. Notably, UniTranslator generates high-quality translations that showcase domain relevance, diversity, and improved image quality. UniTranslator surpasses the performance of existing general-purpose models and performs well against specialized models in representative tasks.
Yong Du 0003, Jiahui Zhan, Xinzhe Li 0003, Junyu Dong, Sheng Chen 0001, Ming-Hsuan Yang 0001, Shengfeng He
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 SITA: Structurally Imperceptible and Transferable Adversarial Attacks for Stylized Image Generation
Jingdan Kang, Haoxin Yang, Yan Cai 0021, Huaidong Zhang, Xuemiao Xu, Yong Du 0003, Shengfeng He
IEEE Trans. Inf. Forensics Secur.6
2025 HG-SFDA: HyperGraph Learning Meets Source-Free Unsupervised Domain Adaptation
abstract
Source-Free unsupervised Domain Adaptation (SFDA) aims to classify target samples by only accessing a pre-trained source model and unlabelled target samples. Since no source data is available, transferring the knowledge from the source domain to the target domain is challenging. Existing methods normally exploit the pair-wise relation among target samples and attempt to discover their correlations by clustering these samples based on semantic features. The drawbacks of these methods include: 1) the pair-wise relation is limited to exposing the underlying correlations of two more samples, hindering the exploration of the structural information embedded in the target domain; and 2) the clustering process only relies on the semantic feature, while overlooking the critical effect of domain shift, i.e., the distribution differences between the source and target domains. To address these issues, we propose a new SFDA method that exploits the high-order neighborhood relation and explicitly takes the domain shift effect into account. Specifically, we formulate the SFDA as a hypergraph learning problem and construct hyperedges to explore the deep structural and context information among multiple samples. Moreover, we integrate a self-loop strategy into the constructed hypergraph to elegantly introduce the domain uncertainty of each sample. By clustering these samples based on hyperedges, both the semantic feature and domain shift effects are considered. We then describe an adaptive relation-based objective to tune the model with soft attention levels for all samples. Extensive experiments are conducted on Office-31, Office-Home, VisDA, DomainNet-126 and PointDA-10 datasets. The results demonstrate the superiority of our method over state-of-the-art counterparts. Our code is avaliable at https://github.com/OUC-POVA/HG-SFDA.
Jinkun Jiang, Qingxuan Lv, Yuezun Li, Yong Du 0003, Junyu Dong, Sheng Chen 0001, Hui Yu 0001
IEEE Trans. Image Process.4
2025 StyleGAN-$\infty$∞: Extending StyleGAN to Arbitrary-Ratio Translation With StyleBook
abstract
Although pre-trained large-scale generative models StyleGAN series have proven to be effective in various editing and translation tasks, they are limited to pre-defined fixed aspect ratio. To overcome this limitation, we propose StyleGAN-$\infty$∞, a model that enables pre-trained StyleGAN to perform arbitrary-ratio conditional synthesis. Our key insight is to distill the expressive StyleGAN features into a StyleBook, such that an arbitrary-ratio condition can be translated to other forms by properly assembling pre-defined StyleBook vectors. To learn and leverage the StyleBook, we employ a network with three distinct stages, each corresponding to StyleBook extraction, StyleBook correspondence learning, and arbitrary-ratio synthesis. Extensive experiments on various conditional synthesis tasks, like super-resolution, sketch synthesis, and semantic synthesis, demonstrate superior performances over state-of-the-art image-to-image translation methods. Moreover, our model can easily generate megapixel images in diverse modalities by taking advantage of different pre-trained StyleGAN models.
Yihua Dai, Tianyi Xiang, Bailin Deng, Yong Du 0003, Hongmin Cai, Harry Qin, Shengfeng He
IEEE Trans. Vis. Comput. Graph.4
2024 D3still: Decoupled Differential Distillation for Asymmetric Image Retrieval
abstract
Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However, these one-to-one constraint approaches often fail to maintain retrieval order consistency, especially when the query network has limited representational capacity. To overcome this problem, we introduce the Decoupled Differential Distillation (D3still) framework. This framework shifts from absolute one-to-one supervision to optimizing the relational differences in pairwise similarities produced by the query and gallery networks, thereby preserving a consistent retrieval order across both networks. Our method involves computing a pairwise similarity differential matrix within the gallery domain, which is then decomposed into three components: feature representation knowledge, inconsistent pairwise similarity differential knowledge, and consistent pairwise similarity differential knowledge. This strategic decomposition aligns the retrieval ranking of the query network with the gallery network effectively. Extensive experiments on various bench-mark datasets reveal that D3still surpasses state-of-the-art methods in asymmetric image retrieval. Code is available at https://github.com/SCY-X/D3still.
Yihong Lin, Xuemiao Xu, Huaidong Zhang, Yong Du 0003, Shengfeng He
CVPR6
2024 rmD4-VTON: Dynamic Semantics Disentangling for Differential Diffusion Based Virtual Try-On
Zhaotong Yang, Zicheng Jiang, Xinzhe Li 0003, Huiyu Zhou 0001, Junyu Dong, Huaidong Zhang, Yong Du 0003
ECCV (46)7
2024 Enhancing Generalized Zero-Shot Learning with Dynamic Selective Knowledge Distillation
Weihua Lv, Yulong Zheng, Chao Liu 0008, Yong Du 0003
WASA (1)4
2023 Curricular Contrastive Regularization for Physics-Aware Single Image Dehazing
abstract
Considering the ill-posed nature, contrastive regularization has been developed for single image dehazing, introducing the information from negative images as a lower bound. However, the contrastive samples are non-consensual, as the negatives are usually represented distantly from the clear (i.e., positive) image, leaving the solution space still under-constricted. Moreover, the interpretability of deep dehazing models is underexplored towards the physics of the hazing process. In this paper, we propose a novel curricular contrastive regularization targeted at a consensual contrastive space as opposed to a non-consensual one. Our negatives, which provide better lower-bound constraints, can be assembled from 1) the hazy image, and 2) corresponding restorations by other existing methods. Further, due to the different similarities between the embeddings of the clear image and negatives, the learning difficulty of the multiple components is intrinsically imbalanced. To tackle this issue, we customize a curriculum learning strategy to reweight the importance of different negatives. In addition, to improve the interpretability in the feature space, we build a physics-aware dual-branch unit according to the atmospheric scattering model. With the unit, as well as curricular contrastive regularization, we establish our dehazing network, named C2PNet. Extensive experiments demonstrate that our C2PNet significantly outperforms state-of-the-art methods, with extreme PSNR boosts of 3.94dB and 1.50dB, respectively, on SOTS-indoor and SOTS-outdoor datasets. Code is available at https://github.com/YuZheng9/C2PNet.
Yu Zheng 0036, Jiahui Zhan, Shengfeng He, Junyu Dong, Yong Du 0003
CVPR5
2023 DSDNet: Toward single image deraining with self-paced curricular dual stimulations
Yong Du 0003, Junjie Deng, Yulong Zheng, Junyu Dong, Shengfeng He
Comput. Vis. Image Underst.1
2022 Editing Out-of-Domain GAN Inversion via Differential Activations
Haorui Song, Yong Du 0003, Tianyi Xiang, Junyu Dong, Harry Qin, Shengfeng He
ECCV (17)2
2022 Delving deep into pixelized face recovery and defense
Zhixuan Zhong, Yong Du 0003, Yang Zhou 0007, Jiang-Zhong Cao, Shengfeng He
Neurocomputing2
2022 Pro-PULSE: Learning Progressive Encoders of Latent Semantics in GANs for Photo Upsampling
abstract
The state-of-the-art photo upsampling method, PULSE, demonstrates that a sharp, high-resolution (HR) version of a given low-resolution (LR) input can be obtained by exploring the latent space of generative models. However, mapping an extreme LR input (162) directly to an HR image (10242) is too ambiguous to preserve faithful local facial semantics. In this paper, we propose an enhanced upsampling approach, Pro-PULSE, that addresses the issues of semantic inconsistency and optimization complexity. Our idea is to learn an encoder that progressively constructs the HR latent codes in the extended$\mathcal {W}+$latent space of StyleGAN. This design divides the complex$64\times $upsampling problem into several steps, and therefore small-scale facial semantics can be inherited from one end to the other. In particular, we train two encoders, the base encoder maps latent vectors in$\mathcal {W}$space and serves as a foundation of the HR latent vector, while the second scale-specific encoder performed in$\mathcal {W}+$space gradually replaces the previous vector produced by the base encoder at each scale. This process produces intermediate side-outputs, which injects deep supervision into the training of encoder. Extensive experiments demonstrate superiorities over the latest latent space exploration methods, in terms of efficiency, quantitative quality metrics, and qualitative visual results.
Yang Zhou 0038, Yangyang Xu 0003, Yong Du 0003, Shengfeng He
IEEE Trans. Image Process.3
2021 Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip Reading
abstract
Lip reading aims to predict the spoken sentences from silent lip videos. Due to the fact that such a vision task usually performs worse than its counterpart speech recognition, one potential scheme is to distill knowledge from a teacher pretrained by audio signals. However, the latent domain gap between the cross-modal data could lead to a learning ambiguity and thus limits the performance of lip reading. In this paper, we propose a novel collaborative framework for lip reading, and two aspects of issues are considered: 1) the teacher should understand bi-modal knowledge to possibly bridge the inherent cross-modal gap; 2) the teacher should adjust teaching contents adaptively with the evolution of the student. To these ends, we introduce a trainable "master" network which ingests both audio signals and silent lip videos instead of a pretrained teacher. The master produces logits from three modalities of features: audio modality, video modality, and their combination. To further provide an interactive strategy to fuse these knowledge organically, we regularize the master with the task-specific feedback from the student, in which the requirement of the student is implicitly embedded. Meanwhile, we involve a couple of "tutor" networks into our system as guidance for emphasizing the fruitful knowledge flexibly. In addition, we incorporate a curriculum learning design to ensure a better convergence. Extensive experiments demonstrate that the proposed network outperforms the state-of-the-art methods on several benchmarks, including in both word-level and sentence-level scenarios.
Sucheng Ren, Yong Du 0003, Jianming Lv, Guoqiang Han 0002, Shengfeng He
CVPR2
2021 From Continuity to Editability: Inverting GANs with Consecutive Images
abstract
Existing GAN inversion methods are stuck in a paradox that the inverted codes can either achieve high-fidelity reconstruction, or retain the editing capability. Having only one of them clearly cannot realize real image editing. In this paper, we resolve this paradox by introducing consecutive images (e.g., video frames or the same person with different poses) into the inversion process. The rationale behind our solution is that the continuity of consecutive images leads to inherent editable directions. This inborn property is used for two unique purposes: 1) regularizing the joint inversion process, such that each of the inverted codes is semantically accessible from one of the other and fastened in an editable domain; 2) enforcing inter-image coherence, such that the fidelity of each inverted code can be maximized with the complement of other images. Extensive experiments demonstrate that our alternative significantly outperforms state-of-the-art methods in terms of reconstruction fidelity and editability on both the real image dataset and synthesis dataset. Furthermore, our method provides the first support of video-based GAN inversion and an interesting application of unsupervised semantic transfer from consecutive images. Source code can be found at: https://github.com/cnnlstm/InvertingGANs_with_ConsecutiveImgs.
Yangyang Xu 0003, Yong Du 0003, Wenpeng Xiao, Xuemiao Xu, Shengfeng He
ICCV2
2021 Fast scene labeling via structural inference
Huaidong Zhang, Chu Han, Xiaodan Zhang 0003, Yong Du 0003, Xuemiao Xu, Guoqiang Han 0002, Harry Qin, Shengfeng He
Neurocomputing4
2021 Mask-ShadowNet: Toward Shadow Removal via Masked Adaptive Instance Normalization
abstract
Shadow removal is an important yet challenging task in image processing and computer vision. Existing methods are limited in extracting good global features due to the interference of shadow. And also, most of them ignore a fact that features inside and outside the shaded area should be treated disparately because of different semantics or materials. In this letter, we propose a novel deep neural network Mask-ShadowNet for shadow removal. The core of our approach is a well-designed masked adaptive instance normalization (MAdaIN) mechanism with embedded aligners that serves two goals: 1) producing hidden features that considering an illumination consistency of different regions. 2) treating the feature statistics of shadow and non-shadow areas discriminately based on the shadow mask. Experimental results demonstrate that the proposed model outperforms the state-of-the-art on the ISTD benchmark. Our code is available inhttps://github.com/penguinbing/Mask-ShadowNet.
Shengfeng He, Bing Peng, Junyu Dong, Yong Du 0003
IEEE Signal Process. Lett.4
2021 Blind Image Denoising via Dynamic Dual Learning
abstract
Existing discriminative learning methods for image denoising use either a single residual learning or a nonresidual learning design. However, we observe that these two schemes perform differently with the same noise level, and yet, there have been no explorations regarding whether residual or nonresidual designs are better suited for denoising. Additionally, many discriminative denoisers are designed to learn a model that corresponds to a fixed noise level, which means that multiple models are required to recover corrupted images with noise at different levels. In this paper, we propose a dynamic dual learning network for blind image denoising, namely, DualBDNet. Instead of modeling a sole task prediction network, the proposed DualBDNet investigates the inherent relations between the residual estimation and the nonresidual estimation. In particular, DualBDNet produces task-dependent feature maps, and each part of the features is devoted to one specific task (residual/nonresidual mapping). To address different noise levels with a single network or even cases where the statistics of noise are unknown, we further introduce an embedded subnetwork into DualBDNet. One output of the subnetwork is the learning of a dynamic compositional attention to highlight the more significant task-dependent feature maps, adaptively coinciding with the extent of corruption. The other output is the learning of a weight used for fusion of the results to ensure an end-to-end manner. Extensive experiments demonstrate that the proposed DualBDNet outperforms the state-of-the-art methods on both synthetic and real noisy images without estimating the noise levels as input.
Yong Du 0003, Guoqiang Han 0002, Yinjie Tan, Chu-Feng Xiao 0001, Shengfeng He
IEEE Trans. Multim.1
2021 Invertible Grayscale with Sparsity Enforcing Priors
abstract
Color dimensionality reduction is believed as a non-invertible process, as re-colorization results in perceptually noticeable and unrecoverable distortion. In this article, we propose to convert a color image into a grayscale image that can fully recover its original colors, and more importantly, the encoded information is discriminative and sparse, which saves storage capacity. Particularly, we design an invertible deep neural network for color encoding and decoding purposes. This network learns to generate a residual image that encodes color information, and it is then combined with a base grayscale image for color recovering. In this way, the non-differentiable compression process (e.g., JPEG) of the base grayscale image can be integrated into the network in an end-to-end manner. To further reduce the size of the residual image, we present a specific layer to enhance Sparsity Enforcing Priors (SEP), thus leading to negligible storage space. The proposed method allows color embedding on a sparse residual image while keeping a high, 35dB PSNR on average. Extensive experiments demonstrate that the proposed method outperforms state-of-the-arts in terms of image quality and tolerability to compression.
Yong Du 0003, Yangyang Xu 0003, Taizhong Ye, Chu-Feng Xiao 0001, Junyu Dong, Guoqiang Han 0002, Shengfeng He
ACM Trans. Multim. Comput. Commun. Appl.1
2021 TPR-DTVN: A Routing Algorithm in Delay Tolerant Vessel Network Based on Long-Term Trajectory Prediction
abstract
An efficient and low‐cost communication system has great significance in maritime communication, but it faces enormous challenges because of high communication costs, incomplete communication infrastructure, and inefficient routing algorithms. Delay Tolerant Vessel Networks (DTVNs), which can create low‐cost communication opportunities among vessels, have recently attracted considerable attention in the academic community. Most existing maritime ad hoc routing algorithms focus on predicting vessels’ future contacts by mining coarse‐grained social relations or spatial distribution, which has led to poor performance. In this paper, we analyze 3‐year trajectory data of 5123 fishery vessels in the China East Sea. Using entropy theory, we observe that the trajectory of the vessel has strongly spatial‐temporal distribution regularity, especially when previous states were given. To predict accurate future trajectories, we develop a long‐term accurate trajectory prediction model by improving the Bidirectional Long‐Short Term Memory (Bi‐LSTM) model. Based on predicted trajectories and the confident degree of each prediction step, we propose a series of routing algorithms called TPR‐DTVN to achieve efficient communication performance. Finally, we carry out simulation experiments with extensive real data. Compared with existing algorithms, the simulation results show that TPR‐DTVN can achieve a higher delivery ratio with lower cost and transmission delay.
Chao Liu 0008, Yingbin Li, Ruobing Jiang, Yong Du 0003, Zhongwen Guo
Wirel. Commun. Mob. Comput.4
2020 Real-Time Hierarchical Supervoxel Segmentation via a Minimum Spanning Tree
abstract
Supervoxel segmentation algorithm has been applied as a preprocessing step for many vision tasks. However, existing supervoxel segmentation algorithms cannot generate hierarchical supervoxel segmentation well preserving the spatiotemporal boundaries in real time, which prevents the downstream applications from accurate and efficient processing. In this paper, we propose a real-time hierarchical supervoxel segmentation algorithm based on the minimum spanning tree (MST), which achieves state-of-the-art accuracy meanwhile at least 11× faster than existing methods. In particular, we present a dynamic graph updating operation into the iterative construction process of the MST, which can geometrically decrease the numbers of vertices and edges. In this way, the proposed method is able to generate arbitrary scales of supervoxels on the fly. We prove the efficiency of our algorithm that can produce hierarchical supervoxels in the time complexity of O(n) , where n denotes the number of voxels in the input video. Quantitative and qualitative evaluations on public benchmarks demonstrate that our proposed algorithm significantly outperforms the state-of-the-art algorithms in terms of supervoxel segmentation accuracy and computational efficiency. Furthermore, we demonstrate the effectiveness of the proposed method on a downstream application of video object segmentation.
Bo Wang 0057, Yiliang Chen, Wenxi Liu, Harry Qin, Yong Du 0003, Guoqiang Han 0002, Shengfeng He
IEEE Trans. Image Process.5
2019 Exploiting Global Low-Rank Structure and Local Sparsity Nature for Tensor Completion
abstract
In the era of data science, a huge amount of data has emerged in the form of tensors. In many applications, the collected tensor data are incomplete with missing entries, which affects the analysis process. In this paper, we investigate a new method for tensor completion, in which a low-rank tensor approximation is used to exploit the global structure of data, and sparse coding is used for elucidating the local patterns of data. Regarding the characterization of low-rank structures, a weighted nuclear norm for the tensor is introduced. Meanwhile, an orthogonal dictionary learning process is incorporated into sparse coding for more effective discovery of the local details of data. By simultaneously using the global patterns and local cues, the proposed method can effectively and efficiently recover the lost information of incomplete tensor data. The capability of the proposed method is demonstrated with several experiments on recovering MRI data and visual data, and the experimental results have shown the excellent performance of the proposed method in comparison with recent related methods.
Yong Du 0003, Guoqiang Han 0002, Yuhui Quan, Zhiwen Yu 0002, Hau-San Wong, C. L. Philip Chen, Jun Zhang 0003
IEEE Trans. Cybern.1