Dingguo Yu

dblp:37/7487 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Learning Multi-Grained Interpretable Latent Representation for 3D Face Manipulation
abstract
Representing 3D faces using generative models has been investigated for several years for its numerous applications in computer vision and graphics. However, general 3D face manipulation is often limited by the lack of multi-level interpretability of the latent space in 3D generative models. To address this problem, we propose a novel generative approach dubbed hierarchically semantic regularized variational auto-encoders (HSR-VAE), which explicitly endows latent variables with multi-grained semantics of the synthesized 3D face shapes. Specifically, to accommodate the hierarchical structure of the human face, we decompose the latent space to represent variations in facial features at different scales, from local facial segments to fine-grained attributes. Moreover, part-aware and attribute-aware semantic regularizers are introduced to establish a linkage between hierarchically organized latent variables and multi-grained facial semantics, allowing more interpretable and meaningful representations of the 3D face. Extensive quantitative and qualitative experiments show the effectiveness of HSR-VAE and demonstrate that it can provide a more interpretable, manipulable, and generalizable latent representation than current approaches, facilitating a wide range of 3D face shape manipulation tasks.
Wen Gao 0001, Naye Ji, Dingguo Yu
Comput. Vis. Media4
2025 Image Quality Difference Perception Ability: A BIQA model effectiveness metric based on model falsification method
abstract
Blind Image Quality Assessment (BIQA) methods refer to algorithms that predict image quality scores without reference images. This challenging area is crucial for preprocessing and optimizing visual tasks. Since BIQA is a small-sample task, it has become a common practice to apply transfer learning by using models pre-trained on other visual tasks, such as semantic recognition, to BIQA. This paper is interested in whether a backbone that performs better in semantic recognition also improves predictive accuracy in BIQA tasks. Comparative experiments showed that different semantic backbones with varying precision exhibit minimal differences in PLCC (Pearson Linear Correlation Coefficient) and SRCC (Spearman Rank Order Correlation Coefficient). However, further comparison shows that the semantic backbone does enhance the model’s IQA abilities, which traditional metrics fail to capture due to their oversight of quality differences between images. We term this ability as Image Quality Difference Perception Ability (IQDP Ability). Based on this, we propose an IQDP comparison method and a new metric, which can effectively compare and measure a model’s IQDP Ability, supplementing traditional metrics and providing an effective means of identifying superior models. • Semantic backbones offer insignificant PLCC/SRCC improvement in IQA performance. • IQDP highlights ‘Image Quality Difference Perception Ability’ that PLCC/SRCC miss. • A new metric describes the quality difference perception ability of BIQA models.
Jinchi Zhu, Dingguo Yu, Yidan Zhao
Expert Syst. Appl.3
2025 Disentangled text-driven stylization of 3D faces via directional CLIP losses
abstract
3D face stylization remains challenging due to limited training samples, diverse style domains, and the complex mapping between ambiguous style features and 3D face structures. To address these issues, we propose ClipStyleFace, a text-driven approach for 3D face stylization that leverages CLIP (Contrastive Language-Image Pre-training) knowledge to create style variations in both geometric and texture structures. ClipStyleFace comprises three components. For geometry deformation, a deformable surface is designed to model stylized geometric residuals on the initial mesh. For texture transformation, we construct a compact parameter space enabling style transfer using a pre-trained albedo generator. Both modules are optimized consistently by distilling semantic alignment and domain correction knowledge from the CLIP model. Extensive experiments demonstrate the effectiveness of our approach in generating stylized 3D faces that match target style prompts while preserving identity characteristics and facial details. Our model also holds promise for applications such as animation and image-driven 3D stylized face generation. Our code is released on https://github.com/cutegao715/ClipStyleFace .
Wenjing Gao, JiaoJiao Wang, Dingguo Yu
Vis. Comput.5
2024 AMSC-Net: Anatomy and multi-label semantic consistency network for semi-supervised fluid segmentation in retinal OCT
abstract
Automated segmentation of pathological fluid regions is crucial for digital diagnosis and individualized therapy under optical coherence tomography (OCT) images. Nonetheless, methods that rely on enormous annotations impede their clinical applications, as pixel-wise fine-grained labels are extraordinarily costly and require the expertise of ophthalmology. Though there exist works based on consistency-based and pseudo-label-based methods for reducing annotation reliance, the underutilization of consistency mechanisms and ignorance of fluid anatomical structures result in sub-optimal segmentation performance. In this work, we proposed a novel AMSC-Net devoted to semi-supervised fluid segmentation and achieves a 73.95% Dice score with 5% labeled data. Concretely, we developed a Heterogeneous Architecture Consistency (HAC) strategy based on our dual decoders enjoying different inductive biases. Moreover, we invented a Multi-label Semantic Consistency Loss (MSC-Loss) module in hierarchical semantic features and an Anatomy Contour Consistency Loss (ACC-Loss) module integrating anatomical constraint. These modules complementarily reinforce the quality of pseudo labels and boost semi-supervised training for robust segmentation results. To verify the superiority and bedside potential of our proposed AMSC-Net, we collected a large-scale fluid segmentation dataset composed of 22517 OCT images. Extensive quantitative and qualitative experiments validated the efficacy of our AMSC-Net with multiple novel techniques. Also, experimental results in a public fluid segmentation dataset demonstrated that our method achieves state-of-the-art performance. Code will be available at: https://github.com/ZeroOneGame/S4_Fluid_OCT.
Yaqi Wang 0002, Ruilong Dan, Shan Luo 0003, Lingling Sun, Qicen Wu, Kangming Yan, Xin Ye 0006, Dingguo Yu
Expert Syst. Appl.10
2024 Learning degradation priors for reliable no-reference image quality assessment
Zhuonan Shen, Bolun Zheng, Dingguo Yu, Chenggang Yan 0001
J. Vis. Commun. Image Represent.5
2023 A Model-Agnostic Semantic-Quality Compatible Framework based on Self-Supervised Semantic Decoupling
abstract
Blind Image Quality Assessment (BIQA) is a challenging research topic that is critical for preprocessing and optimizing downstream vision tasks such as semantic recognition and image restoration. However, there has been a significant disconnect between BIQA research and other vision tasks. The primary cause of such disconnect is the incompatibility of existing BIQA models with other vision tasks, resulting in significant computational complexity. To address this issue, we propose a model-agnostic semantic-quality compatible framework that can simultaneously generate quality and semantic predictions. By incorporating a lightweight learning architecture, we demonstrate that a parameter-fixed semantic-oriented backbone can predict the perceptual quality of images as accurately as models trained end-to-end. We systematically study the major components of our framework, and our experimental results demonstrate the superiority of our model in terms of both complexity and accuracy. The source code of this work is available at https://github.com/MaxiaoyuHehe/SQCFNet.
Chenxi Feng, Suiyu Zhang, Jinchi Zhu, Chang Liu 0114, Dingguo Yu
ACM Multimedia9
2023 ASCAM-Former: Blind image quality assessment based on adaptive spatial & channel attention merging transformer and image to patch weights sharing
abstract
Blind Image Quality Assessment (BIQA) is a challenging, unsolved research topic which is crucial for analyzing, understanding, and improving visual experience. Recently, transformer-based BIQA models are drawing increasing attention due to their powerful capacity in modeling global dependencies amongst tokens. However, existing works tend to apply self-attention mechanism for exploring the spatial dependencies whilst neglecting the impact of channel-wise self-attention. In this paper, we explore the feasibility of incorporating attention mechanism in a channel-wise manner for BIQA. By systematically studying the interactions between channel-wise and spatial-wise attention, an adaptive spatial and channel attention merging Transformer (ASCAM-Former) is then proposed for aggregating both the spatial-wise and channel-wise attention information. In addition, to accommodate IQA datasets containing both image and patch quality labels, an image to patch weights sharing (I2PWS) scheme is designed to take advantage of local quality learning tasks for reinforcing the learning of global quality, and vice versa. The experimental results indicate that channel-wise attention mechanism is as competitive as spatial-wise for IQA tasks, and the proposed ASCAM-Former yield accurate prediction on both authentically and synthetically distorted image quality datasets.
Suiyu Zhang, Yaqi Wang 0002, Dingguo Yu
Expert Syst. Appl.6
2022 ADGNet: Attention Discrepancy Guided Deep Neural Network for Blind Image Quality Assessment
abstract
This work explores how to efficiently incorporate semantic knowledge for blind image quality assessment and proposes an end-to-end attention discrepancy guided deep neural network for perceptual quality assessment. Our method is established on a multi-task learning framework in which two sub-tasks including semantic recognition and image quality prediction are jointly optimized with a shared feature-extracting branch and independent spatial-attention branch. The discrepancy between semantic-aware attention and quality-aware attention is leveraged to refine the quality predictions. The proposed ADGNet is based on the observation that human visual systems exhibit different mechanisms when viewing images with different amounts of distortion. Such a manner would result in the variation of attention discrepancy between the quality branch and semantic branch, which are therefore employed to enhance the accuracy and generalization ability of our method. We systematically study the major components of our framework, and experimental results on both authentically and synthetically distorted image quality datasets demonstrate the superiority of our model as compared to the state-of-the-art approaches.
Chang Liu 0114, Suiyu Zhang, Dingguo Yu
ACM Multimedia5
2022 An embedding and interactions learning approach for ID feature in deep recommender system
abstract
The embedding of features plays a critical role in the neural network-based recommendation model. The ID feature, which is discrete and high-dimensional sparse, contains the significant information for relevance reasoning in recommender system. However, most models use the same embedding method to address the encoding of all features. This results in inaccurate expression of ID features, as well as reduces the generalization ability of the model. Based on the assumption that learning higher-order information from the mixture of ID and other features by crossover and extraction operations provides limited benefits, or negative results, we propose an ID feature learning model (IFLM) to embed the ID tags independently and learn their interactions efficiently. IFLM separates the ID tags from the complete inputs as its exclusive input, then outputs the vectors with general data structure, thus, making it easily parallel to other deep recommendation models. Besides, IFLM generates the second-order cross feature from pairwise IDs and captures higher-order feature through the specific hidden layer, which is more efficient in learning the knowledge of ID features. We used our approach on three typical deep recommendation models. The experimental results showed that the optimized model produced better results than the corresponding baseline, in terms of prediction accuracy and convergence time.
Suiyu Zhang, Dingguo Yu
Expert Syst. Appl.5
2022 Conformance-oriented Predictive Process Monitoring in BPaaS Based on Combination of Neural Networks
abstract
Abstract As a new cloud service for delivering complex business applications, Business Process as a Service (BPaaS) is another challenge faced by cloud service platforms recently. To effectively reduce the security risk caused by business process execution load in BPaaS, it is necessary to detect the non-compliant process executions (instances) from tenants in advance by checking and monitoring the conformance of the executing process instances in real-time. However, the vast majority of existing conformance checking techniques can only be applied to the process instances that have been executed completely offline and only focus on the conformance from the single control-flow perspective. We develop an extensible multi-perspective conformance measurement method to address these issues first and then investigate the predictive conformance monitoring approach by automatically constructing an online multi-perspective conformance prediction model based on deep learning techniques. In addition, to capture more decisive features in the model from both local information and long-distance dependency within an executed process instance, we propose an approach, called CNN-BiGRU, by combining Convolutional Neural Network (CNN) with a variant and enhancement of Recurrent Neural Network (RNN). Extensive experiments on two data sets demonstrate the effectiveness and efficiency of the proposed CNN-BiGRU.
Victor Chang 0001, Dongjin Yu, Dingguo Yu
J. Grid Comput.6
2021 NBA Basketball Video Summarization for News Report via Hierarchical-Grained Deep Reinforcement Learning
Naye Ji, Dingguo Yu, Youbing Zhao
ICIG (3)4
2021 Neurophysiological Assessment of Image Quality from EEG Using Persistent Homology of Brain Network
abstract
The neuropsychological characteristics inside the brain are still not sufficiently understood during the conventional image quality assessment procedure. In this paper, we extract the physiologically meaningful features of brain responses to different distortion levels images in conventional image assessment by combining persistent homology analysis with electroencephalogram (EEG). The experimental results show that more brain regions in the frontal lobe are involved when the subject perceives an unclear image compared to a clear one, which indicates that human perception of image quality might be related to the advanced cognitive processes to some extent. Meanwhile, a statistically significantly higher persistent entropy of EEG data evoked by a clear image compared to that of an unclear image is observed in several frequency bands. In general, this paper evaluates and quantifies quality-related neural correlates by persistent homology features of EEG signal, which provides an approach for a utility neurophysiological assessment of image quality.
Chang Liu 0114, Jiefang Zhang, Honggang Zhang 0001, Songyun Xie, Dingguo Yu
ICME7
2021 DFIAM: deep factorization integrated attention mechanism for smart TV recommendation
Xuewen Shen, Suiyu Zhang, Dingguo Yu, Guandong Xu
World Wide Web4
2020 Video Preview Generation Based on Playback Records
abstract
A video preview is formed by joining a subset of video snippets from a source video. It is expected to be expressive to help enrich users' experience and hence increase the views of the full video. However, it is challenging to identify suitable video snippets to compose such video previews in support of user recommendation of full-length videos. In this paper, we formulate this problem as an optimization problem to maximize the total playback popularity of video segments based on the analysis of a large amount of users' playback records. We design an algorithm for this problem and provide proof of optimality. Furthermore, we conduct an experiment using real industrial data and recruit volunteers to assess the quality of generated video previews. The experimental results show the feasibility and applicability of our methods.
Xuewen Shen, Songlin He, Dingguo Yu, Zhiyan Tang
AICCSA3
2020 Online Predicting Conformance of Business Process with Recurrent Neural Networks
abstract
Conformance Checking is a problem to detect and describe the differences between a given process model representing the expected behaviour of a business process and an event log recording its actual execution by the Process-aware Information System (PAIS). However, such existing conformance checking techniques are offline and mainly applied for the completely executed process instances, which cannot provide the real-time conformance-oriented process monitoring for an on-going process instance. Therefore, in this paper, we propose three approaches for online conformance prediction by constructing a classification model automatically based on the historical event log and the existing reference process model. By utilizing Recurrent Neural Networks, these approaches can capture the features that have a decisive effect on the conformance for an executed case to build a prediction model and then use this model to predict the conformance of a running case. The experimental results on two real datasets show that our approaches outperform the state-of-the-art ones in terms of prediction accuracy and time performance.
Dingguo Yu, Victor Chang 0001, Xuewen Shen
IoTBDS2
2017 Constrained NMF-based semi-supervised learning for social media spammer detection
Dingguo Yu, Frank Jiang 0001, Aihong Qin
Knowl. Based Syst.1
2009 The Improving of IKE with PSK for Using in Mobile Computing Environments
abstract
The rapid increase in using mobile communication networks for transmitting confidential data and conducting commercial transactions such as mobile e-commerce is creating large demands in designing secure mobile business systems. However, the mobile devices and mobile communication network have some weakness. It can cause some problems using traditional VPN technologies in mobile computing environments immediately. Currently, mobile users' authentication in IKE is being done using certificates or PSK with aggressive mode commonly. They have serious security related issues (for PSK with aggressive mode) and need high deployment and maintain cost (for certificates). In this paper, we propose a new approach that is based on PSK where the IKE negotiation phase is modified for using in mobile computing environments. The modified IKE consists of four messages, and the responder doesn't need to store any state while receiving message 1. It uses strong cookies and pre-calculated DHpp stack, etc technologies to counter IP flooding attacks and man-in-the-middle DoS attacks, because it does not require the responder to perform heavy computations before the initiator has authenticated itself. Otherwise, for one mobile user, it has a group of PSKs to be random selected, and the initiator and responder exchange identity info and agree on PSK with Hash (PSK-ID|IDi) or Hash (PSK-ID|IDr) info. Therefore, it provides the initiator and responder's identity protection and prevention of passive dictionary based attacks on pre-shared keys.
Dingguo Yu
IAS1