VLDB 2026 Research / reviewers in the wild / expert
Guiduo Duan
dblp:162/1088
· DBLP profile ↗
29ranked-venue papers
2as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TiCAL: Typicality-Based Consistency-Aware Learning for Multimodal Emotion RecognitionabstractMultimodal Emotion Recognition (MER) aims to accurately identify human emotional states by integrating heterogeneous modalities such as visual, auditory, and textual data. Existing approaches predominantly rely on unified emotion labels to supervise model training, often overlooking a critical challenge: inter-modal emotion conflicts, wherein different modalities within the same sample may express divergent emotional tendencies. In this work, we address this overlooked issue by proposing a novel framework, Typicality-based Consistent-aware Multimodal Emotion Recognition (TiCAL), inspired by the stage-wise nature of human emotion perception. TiCAL dynamically assesses the consistency of each training sample by leveraging pseudo unimodal emotion labels alongside a typicality estimation. To further enhance emotion representation, we embed features in a hyperbolic space, enabling the capture of fine-grained distinctions among emotional categories. By incorporating consistency estimates into the learning process, our method improves model performance, particularly on samples exhibiting high modality inconsistency. Extensive experiments on benchmark datasets, e.g, MOSEI and MER2023, validate the effectiveness of TiCAL in mitigating inter-modal emotional conflicts and enhancing overall recognition accuracy, e.g., with about 2.6% improvements over the state-of-the-art DMD. Siyu Zhan, Cencen Liu, Guiduo Duan, Xiurui Xie, Yuan-Fang Li, Tao He 0007 |
AAAI | 5 |
| 2026 | FreSCo: Joint Frequency-Aware and Spatial Control for Image Zero-Shot Style TransferabstractLatent Diffusion Models (LDMs) have become a cornerstone for zero-shot style transfer in multimedia content creation, but they frequently struggle with a critical trade-off between artistic stylization fidelity and semantic structural preservation. A key oversight in existing methods is the neglect of frequency-domain distinctions in visual signals, which leads to prevalent issues like content drift and style leakage. To address these limitations, we propose FreSCo, a novel training-free framework that explicitly decouples content and style through dual-domain control mechanisms. First, the Dynamic Wavelet Latent Fusion (DWLF) module decomposes latent features via Discrete Wavelet Transform (DWT), injecting style exclusively into high-frequency texture sub-bands while boosting spectral energy to counteract VAE-induced smoothing. Second, the VAE-Compressed Masking strategy encodes edge maps directly into the latent space, resolving pixel-latent misalignment for precise spatial control. We construct a comprehensive benchmark with 1,280 content-style pairs to rigorously evaluate performance. Extensive experiments demonstrate that FreSCo achieves state-of-the-art results, generating high-fidelity artistic textures while maintaining superior structural consistency across diverse multimedia content creation scenarios compared to existing baselines. Tingrun Chen, Xudong Ling, Shicai Wei, Guiduo Duan, Yue Zhang 0042 |
ICMR | 4 |
| 2026 | Towards Conflict-aware Selective Knowledge Unlearning for Continual Few-shot Knowledge Graph Completion
Junlin Zhu 0001, Bo Fu 0007, Guiduo Duan |
SIGIR | 3 |
| 2026 | Anchor Drift No More: Hierarchical Consistency-Guided Prompt Distillation for Incomplete Multimodal LearningabstractWeb-scale content is rich in modalities yet frequently incomplete due to device limits, transmission errors, or privacy controls, making learning with missing modalities a core challenge. Prior reconstruction and alignment strategies often fail to preserve a stable class geometry when inputs are partial, leading to anchor drift -- a shift of class prototypes between complete and incomplete views that distorts the shared representation space and degrades generalization. We introduce HiCoD (Hierarchical Consistency-Guided Pro mpt Distillation), which learns a robust, class-anchored semantic space. HiCoD combines: (1) a modality-aware semantic graph that restores cross-modal structure under partial observations; (2) dual-level anchoring that unifies large-language-model–derived global category prototypes with top-K local exemplars to balance cross-modal coherence and modality-specific detail; and (3) multi-level distillation that aligns unimodal features, fused embeddings, and prompt-completed signals within a single anchor space. Across CMU-MOSI, CMU-MOSEI, and additional benchmarks, HiCoD sets a new state of the art under both fixed-pattern and random missingness, improving Acc-2 by up to 6.4 points over MPLMM and remaining robust when key modalities are absent. Ruiting Dai, Zesen Cai, Lisi Mo, Guiduo Duan, Keren Shi, Tao He 0007 |
WWW | 4 |
| 2026 | Modular knowledge distillation for MIM pre-trained models
Dongyang Zhang 0001, Guiduo Duan |
Neurocomputing | 5 |
| 2026 | Efficient Dataset Distillation via Generative PruningabstractDataset distillation (DD) has demonstrated the promise of synthesizing smaller datasets that enable competitive performance. While most of DD methods operate in pixel space, they often suffer from poor scalability on high-resolution datasets. Recent works have thus shifted towards parameterizing synthetic data using deep generative priors. However, these approaches apply latent updates to all generator layers, leading to substantial computational overhead. To address this limitation, we propose Generative Lightweight Distillation (GLiD), a unified framework that jointly compresses the generator and accelerates latent optimization for efficient distillation. GLiD introduces two key components: (1) We introduce a Classification-Diversity Sensitivity Pruning mechanism that quantifies both discriminative utility and semantic diversity of each output channel to guide structural pruning and selective layer-wise optimization. (2) We also present a Layer-Adaptive Scheduling strategy that dynamically allocates latent update steps across generator stages based on convergence behavior. Extensive experiments on CIFAR-10 and ImageNet-1K and its subsets demonstrate that our GLiD achieves up to 10× acceleration across different datasets, while maintaining performance competitive with state-of-the-art methods. Yingyi Ma, Muquan Li, Guiduo Duan, Ke Qin, Shuang Liang 0002, Dongyang Zhang 0001 |
IEEE Trans. Big Data | 3 |
| 2025 | DebiasedKGE: Towards Mitigating Spurious Forgetting in Continual Knowledge Graph EmbeddingabstractTo maintain an effective memory of old knowledge in a dynamically growing knowledge environment, continual knowledge graph embedding (CKGE) focuses on alleviating catastrophic forgetting. However, existing CKGE methods still suffer substantial performance degradation in dynamic knowledge graphs (DKG). We have found this challenge is mainly posed by spurious forgetting, a previously overlooked phenomenon that arises from the inherent interference effects in the continual learning (CL) process. In this paper, we deeply explore spurious forgetting in CKGE. First, we reveal two primary causes of spurious forgetting, knowledge interference and knowledge misalignment, and how to affect knowledge biasing within dynamic learning scenarios. Second, to fill this research gap, we propose a robust and efficient CKGE method (DebiasedKGE) for mitigating spurious forgetting. Specifically, to alleviate knowledge interference, we propose a mutual information-guided disentangled learning mechanism, which identifies latent features of different knowledge types and learns independent semantic representations for each, thereby reducing interference in knowledge embedding. Furthermore, to mitigate the deviation of new knowledge from previously learned knowledge, we design a dual-view regularized knowledge alignment mechanism that jointly constrains both the magnitude and direction of embedding transitions. Finally, we evaluate DebiasedKGE on four public CKGE datasets and two additional datasets constructed to contain knowledge perturbations of different dimensions. The results show that DebiasedKGE effectively alleviates spurious forgetting and achieves significant performance improvements. Our codes and datasets are available at https://anonymous.4open.science/r/DebiasedKGE. Junlin Zhu 0001, Bo Fu 0007, Guiduo Duan |
CIKM | 3 |
| 2025 | ROLL: Robust Noisy Pseudo-label Learning for Multi-View Clustering with Noisy CorrespondenceabstractMulti-view clustering (MVC) aims to exploit complementary information from diverse views to enhance clustering performance. Since pseudo-labels can provide additional semantic information, many MVC methods have been proposed to guide unsupervised multi-view learning through pseudo-labels. These methods implicitly assume that the predicted pseudo-labels are predicted correctly. However, due to the challenges in training a flawless unsupervised model, this assumption can be easily violated, thereby leading to the Noisy Pseudo-label Problem (NPP). Moreover, these existing approaches typically rely on the assumption of perfect cross-view alignment. In practice, it is frequently compromised due to noise or sensor differences, thereby resulting in the Noisy Correspondence Problem (NCP). Based on the above observations, we reveal and study unsupervised multi-view learning under NPP and NCP. To this end, we propose Robust Noisy Pseudo-label Learning (ROLL) to prevent the overfitting problem caused by both NPP and NCP. Specifically, we first adopt traditional contrastive learning to warm up the model, thereby generating the pseudo-labels in a self-supervised manner. Afterward, we propose noise-tolerance pseudo-label learning to deal with the noise in the predicted pseudo-labels, thereby embracing the robustness against NPP. To further mitigate the overfitting problem, we present robust multi-view contrastive learning to mitigate the negative impact of NCP. Extensive experiments on five multi-view datasets demonstrate the superior clustering performance of our ROLL compared to 11 state-of-the-art methods. Yuan Sun 0016, Zhenwen Ren, Guiduo Duan, Dezhong Peng, Peng Hu 0002 |
CVPR | 4 |
| 2025 | Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion RecognitionabstractVisual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stickers, limiting VER models’ cross-domain generalizability. To fill this gap, we introduce an Unsupervised Cross-Domain Visual Emotion Recognition (UCDVER) task, which aims to generalize visual emotion recognition from the source domain (e.g., realistic images) to the low-resource target domain (e.g., stickers) in an unsupervised manner. Compared to the conventional unsupervised domain adaptation problems, UCDVER presents two key challenges: a significant emotional expression variability and an affective distribution shift. To mitigate these issues, we propose the Knowledge-aligned Counterfactual-enhancement Diffusion Perception (KCDP) framework. Specifically, KCDP leverages a VLM to align emotional representations in a shared knowledge space and guides diffusion models for improved visual affective perception. Furthermore, a Counterfactual-Enhanced Language-image Emotional Alignment (CLIEA) method generates high-quality pseudo-labels for the target domain. Extensive experiments demonstrate that our model surpasses SOTA models in both perceptibility and generalization, e.g., gaining 12% improvements over SOTA VER model TGCA-PVT. The project page is at https://yinwen2019.github.io/ucdver/. Guiduo Duan, Dongyang Zhang 0001, Yuan-Fang Li, Tao He 0007 |
CVPR | 3 |
| 2025 | WMAJL: Watcher-Mediated Attention Joint Learning Model for Multimodal Relation ExtractionabstractIn the domain of Multimodal Relation Extraction (MRE), we present the $\color{Red}{\text{W}}$atcher-$\color{Red}{\text{M}}$ediated $\color{Red}{\text{A}}$ttention $\color{Red}{\text{J}}$oint $\color{Red}{\text{L}}$earning Model ($\color{Red}{\text{WMAJL}}$), a novel approach addressing the challenges of modality alignment noise, cross-modal fusion disparity, preservation of textual relative position information, and the distinctiveness of classification labels. WMAJL employs an integrative framework leveraging contrastive learning and variational autoencoder constraints to mitigate modality alignment noise by prioritizing relevant semantic data and effectively reducing extraneous noise that does not contribute to the task. The model’s innovative architecture includes a mediator watcher, which facilitates enhanced cross-modal fusion by enabling nuanced information exchange between textual and visual modalities while preserving the unique characteristics of each modality. Additionally, the design of auxiliary tasks, such as Named Entity Recognition (NER), and output supervision constructs loss functions that preserve relative position information, ensuring a precise depiction of entity relationships throughout the multilayer encoding processes. A key differentiator of WMAJL is its label-centric self-information loss technique, inspired by InfoNCE, which trains the model to cluster similar relation labels in semantically coherent areas, thereby optimizing classification label uniqueness by discerning subtle differences among relation types. The synergistic application of these strategies has led to a significant enhancement of WMAJL’s performance, as evidenced by its state-of-the-art F1 score of $\color{Red}{84.93\%}$ on the MNRE dataset. This achievement surpasses existing benchmarks and sets a new standard for multimodal knowledge extraction, underscoring WMAJL’s potential to revolutionize the MRE landscape. Yunrui Dong, Guiduo Duan, Tianxi Huang |
ICASSP | 2 |
| 2025 | Unbiased Multimodal Audio-to-Intent RecognitionabstractAudio-to-intent recognition is a critical task focused on identifying a speaker’s intent from spoken language. Recently, multimodal audio-to-intent approaches have emerged as the predominant strategy for enhancing audio-based intent recognition. In this study, we conduct a series of empirical experiments that reveal a significant modality bias in current multimodal audio-to-intent recognition methods. Specifically, these methods disproportionately rely on the textual modality to determine intent, often neglecting the audio data. To address this issue, we propose a context-enhanced contrastive learning framework designed to capture rich regional- and global-audio context information, thereby enabling more balanced audio-to-intent recognition. Additionally, we introduce a prototype-based intent classification strategy that encourages different intent classes and modalities to converge toward unified prototypes, leading to smoother classification boundaries as opposed to the traditionally skewed boundaries. Extensive experiments demonstrate that our approach effectively mitigates modality bias, e.g., a performance improvement of 2.12% in intent classification compared to the state-of-the-art method GZAIR on the dataset MintRec. Yuezhou Dong, Ke Qin, Guiduo Duan, Tao He 0007 |
ICASSP | 4 |
| 2025 | DFMA: Adaptive Dual Fusion for Multimodal Relation Extraction with Mutual AttentionabstractMultimodal relation extraction (MRE) is an emerging research field that combines techniques from natural language processing, computer vision, and machine learning, helping us better understand and interpret data. However, current methods are faced with two main issues. The first issue is that the auxiliary images may focus on irrelevant information, leading the model to place too much attention on unrelated entities when extracting visual features from the global image. The second issue is modality imbalance during the multimodal fusion stage, where the encoder might overly focus on one modality’s information over the other. Therefore, we propose a MRE method using Dual Fusion with Mutual Attention (DFMA). To address the first issue, we introduce the Pyramid Feature Extraction (PFE) module. To tackle the second issue, we propose the Multimodal Mutual Fusion (MMF) module. PFE dynamically generates prompt vectors to mitigate the impact of irrelevant information, while MMF, designed with mutual attention, can balance the fusion of textual and visual information simultaneously. Experiments show that our approach achieves the state-of-the-art performances on the multimodal neural relation extraction (MNRE) dataset. Tianxi Huang, Guiduo Duan, Yunrui Dong |
ICASSP | 3 |
| 2025 | SPADE: Spatial-Aware Denoising Network for Open-Vocabulary Panoptic Scene Graph Generation with Long- and Local-Range Context ReasoningabstractPanoptic Scene Graph Generation (PSG) integrates instance segmentation with relation understanding to capture pixel-level structural relationships in complex scenes. Although recent approaches leveraging pre-trained vision-language models (VLMs) have significantly improved performance in the open-vocabulary setting, they commonly ignore the inherent limitations of VLMs in spatial relation reasoning, such as difficulty in distinguishing object relative positions, which results in suboptimal relation prediction. Motivated by the denoising diffusion model's inversion process in preserving the spatial structure of input images, we propose SPADE (SPatial-Aware Denoising-nEtwork) framework -- a novel approach for open-vocabulary PSG. SPADE consists of two key steps: (1) inversion-guided calibration for the UNet adaptation, and (2) spatial-aware context reasoning. In the first step, we calibrate a general pre-trained teacher diffusion model into a PSG-specific denoising network with cross-attention maps derived during inversion through a lightweight LoRA-based fine-tuning strategy. In the second step, we develop a spatial-aware relation graph transformer that captures both local and long-range contextual information, facilitating the generation of high-quality relation queries. Extensive experiments on benchmark PSG and Visual Genome datasets demonstrate that SPADE outperforms state-of-the-art methods in both closed- and open-set scenarios, particularly for spatial relationship prediction. Ke Qin, Guiduo Duan, Ming Li 0065, Yuan-Fang Li, Tao He 0007 |
ICCV | 3 |
| 2025 | Deep Fuzzy Multi-view Learning for Reliable ClassificationabstractMulti-view learning methods primarily focus on enhancing decision accuracy but often neglect the uncertainty arising from the intrinsic drawbacks of data, such as noise, conflicts, etc. To address this issue, several trusted multi-view learning approaches based on the Evidential Theory have been proposed to capture uncertainty in multi-view data. However, their performance is highly sensitive to conflicting views, and their uncertainty estimates, which depend on the total evidence and the number of categories, often underestimate uncertainty for conflicting multi-view instances due to the neglect of inherent conflicts between belief masses. To accurately classify conflicting multi-view instances and precisely estimate their intrinsic uncertainty, we present a novel Deep Fuzzy Multi-View Learning (FUML) method. Specifically, FUML leverages Fuzzy Set Theory to model the outputs of a classification neural network as fuzzy memberships, incorporating both possibility and necessity measures to quantify category credibility. A tailored loss function is then proposed to optimize the category credibility. To further enhance uncertainty estimation, we propose an entropy-based uncertainty estimation method leveraging category credibility. Additionally, we develop a Dual Reliable Multi-view Fusion (DRF) strategy that accounts for both view-specific uncertainty and inter-view conflict to mitigate the influence of conflicting views in multi-view fusion. Extensive experiments demonstrate that our FUML achieves state-of-the-art performance in terms of both accuracy and reliability. Siyuan Duan, Yuan Sun 0016, Dezhong Peng, Guiduo Duan, Xi Peng 0001, Peng Hu 0002 |
ICML | 4 |
| 2025 | Reliable Disentanglement Multi-view Learning Against View Adversarial AttacksabstractTrustworthy multi-view learning has attracted extensive attention because evidence learning can provide reliable uncertainty estimation to enhance the credibility of multi-view predictions. Existing trusted multi-view learning methods implicitly assume that multi-view data is secure. However, in safety-sensitive applications such as autonomous driving and security monitoring, multi-view data often faces threats from adversarial perturbations, thereby deceiving or disrupting multi-view models. This inevitably leads to the adversarial unreliability problem (AUP) in trusted multi-view learning. To overcome this tricky problem, we propose a novel multi-view learning framework, namely Reliable Disentanglement Multi-view Learning (RDML). Specifically, we first propose evidential disentanglement learning to decompose each view into clean and adversarial parts under the guidance of corresponding evidences, which is extracted by a pretrained evidence extractor. Then, we employ the feature recalibration module to mitigate the negative impact of adversarial perturbations and extract potential informative features from them. Finally, to further ignore the irreparable adversarial interferences, a view-level evidential attention mechanism is designed. Extensive experiments on multi-view classification tasks with adversarial attacks show that RDML outperforms the state-of-the-art methods by a relatively large margin. Our code is available at https://github.com/Willy1005/2025-IJCAI-RDML. Siyuan Duan, Qizhi Li, Guiduo Duan, Yuan Sun 0016, Dezhong Peng |
IJCAI | 4 |
| 2025 | Noise-Robust Cross-modal Learning for Reliable 2D-3D RetrievalabstractWith the rapid proliferation of 2D and 3D data, driven by advances in virtual environments and AI-generated content, cross-modal 2D-3D retrieval has attracted growing attention. However, it is easy to introduce noisy labels due to the spatial complexity of 3D content. Although various methods have been proposed to address this issue, they still struggle to handle or effectively re-exploit noisy samples. Moreover, existing approaches are prone to error accumulation due to the self-reinforcement of the model during training. To address these issues, we propose a Noise-Robust Cross-modal Learning (NRCL) framework based on the hybrid strategy. Specifically, NRCL introduces a Robust Cross-modal Co-separator (RCC), which separates noisy samples from clean ones by leveraging modality complementarity and adopting a co-teaching paradigm to mitigate potential error accumulation of the single model during training. Besides, a Reliable Soft Rectification (RSR) method is adopted to correct noisy labels by aggregating historical and dual-model predictions, exploiting the discriminative information from noisy samples. Finally, a Robust Cross-modal Prototype Learning (RCPL) is proposed to improve the discriminability of inter-class and alleviate the inherent gaps across modalities in the shared common space, which jointly leverages clean and rectified labels, thereby mitigating the detrimental impact of noisy samples. Extensive experiments are conducted on three 3D multimodal datasets to verify the effectiveness of our method by comparing it with 10 state-of-the-art methods. The code is available at https://github.com/yangaonidaye123/NRCL. Yanglin Feng, Yuan Sun 0016, Dezhong Peng, Guiduo Duan |
ACM Multimedia | 5 |
| 2025 | Multi-Perspective Dialogue Non-Quota Selection with loss monitoring for dialogue state tracking
Jinyu Guo, Zhaokun Wang, Jingwen Pu, Wenhong Tian, Guiduo Duan, Guangchun Luo |
Expert Syst. Appl. | 5 |
| 2025 | Outlier detection in mixed-attribute data: A semi-supervised approach with fuzzy approximations and relative entropy
Baiyang Chen, Zhong Yuan, Dezhong Peng, Chang Liu 0088, Guiduo Duan |
Int. J. Approx. Reason. | 7 |
| 2025 | Unbiased multimodal intent recognition with auxiliary rationale generation
Ruiting Dai, Guiduo Duan, Ke Qin, Tao He 0007 |
Neurocomputing | 3 |
| 2025 | Label as Equilibrium: A performance booster for Graph Neural Networks on node classification
Guangchun Luo, Guiduo Duan |
Neural Networks | 3 |
| 2024 | Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware QueriesabstractThe open-vocabulary human-object interaction (Ov-HOI) detection aims to identify both base and novel categories of human-object interactions while only base categories are available during training. Existing Ov-HOI methods commonly leverage knowledge distilled from CLIP to extend their ability to detect previously unseen interaction categories. However, our empirical observations indicate that the inherent noise present in CLIP has a detrimental effect on HOI prediction. Moreover, the absence of novel human-object position distributions often leads to overfitting on the base categories within their learned queries. To address these issues, we propose a two-step framework named, CaM-LQ, Calibrating visual-language Models, (e.g., CLIP) for open-vocabulary HOI detection with Locality-aware Queries. By injecting the fine-grained HOI supervision from the calibrated CLIP into the HOI decoder, our model can achieve the goal of predicting novel interactions. Extensive experimental results demonstrate that our approach performs well in open-vocabulary human-object interaction detection, surpassing state-of-the-art methods across multiple metrics on mainstream datasets and showing superior open-vocabulary HOI detection performance, e.g., with 4.54 points improvement on the HICO-DET dataset over the SoTA CLIP4HOI on the UV task with the same backbone ResNet-50. Zhenhao Yang, Deqiang Ouyang, Guiduo Duan, Dongyang Zhang 0001, Tao He 0007, Yuan-Fang Li |
ACM Multimedia | 4 |
| 2024 | Towards Elastic Image Super-Resolution Network via Progressive Self-distillation
Xin'an Yu, Dongyang Zhang 0001, Cencen Liu, Qiang Dong, Guiduo Duan |
PRCV (8) | 5 |
| 2022 | Improving finger vein discriminant representation using dynamic margin softmax loss
Huachuan Li, Guiduo Duan |
Neural Comput. Appl. | 3 |
| 2022 | Multi-local feature relation network for few-shot learning
Guiduo Duan, Tianxi Huang, Zhao Kang 0001 |
Neural Comput. Appl. | 2 |
| 2021 | Robust Visual Relationship Detection towards Sparse Images in Internet-of-ThingsabstractVisual relationship can capture essential information for images, like the interactions between pairs of objects. Such relationships have become one prominent component of knowledge within sparse image data collected by multimedia sensing devices. Both the latent information and potential privacy can be included in the relationships. However, due to the high combinatorial complexity in modeling all potential relation triplets, previous studies on visual relationship detection have used the mixed visual and semantic features separately for each object, which is incapable for sparse data in IoT systems. Therefore, this paper proposes a new deep learning model for visual relationship detection, which is a novel attempt for cooperating computational intelligence (CI) methods with IoTs. The model imports the knowledge graph and adopts features for both entities and connections among them as extra information. It maps the visual features extracted from images into the knowledge‐based embedding vector space, so as to benefit from information in the background knowledge domain and alleviate the impacts of data sparsity. This is the first time that visual features are projected and combined with prior knowledge for visual relationship detection. Moreover, the complexity of the network is reduced by avoiding the learning of redundant features from images. Finally, we show the superiority of our model by evaluating on two datasets. Guiduo Duan, Guangchun Luo |
Wirel. Commun. Mob. Comput. | 2 |
| 2020 | An Improved Group Similarity-Based Association Rule Mining Algorithm in Complex ScenesabstractAssociation rule (AR) mining in complex scene has attracted extensive attention of researchers in recent years. Typically, many researchers focused on an algorithm itself and ignored a generalization method to improve the performance of AR mining. Tuna et al., presented a general data structure Speeding-Up AR Structure with Inverted Index Compression (SAII) which could be utilized in most of the existing algorithms to improve their performance IEEE Trans. Cybern. 46(12) (2016) 3059–3072. However, we found that this algorithm consumes a lot of time in re-ordering data because a one-to-one comparison method is used in this process, which is the main reason that the speeding-up structure is difficult to establish when coping with much more large amount of data. To overcome these problems, this paper aims to propose an improved speeding-up AR algorithm based on group similarity and Apache Spark framework to further reduce the memory requirements and runtime. Our simulation results on the police business big dataset make clear that our improved approach performs well and is more suitable for a big data environment. Guiduo Duan, Tianxi Huang, Jürgen Kurths |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2019 | Median Filtering Detection of Small-Size Image Using AlexCaps-Network
Guiduo Duan, Jiayu Miao, Tianxi Huang |
IWDW | 1 |
| 2019 | A Novel Task Allocation Algorithm in Mobile Crowdsensing with Spatial Privacy PreservationabstractThe Internet of Things (IoT) has attracted the interests of both academia and industry and enables various real-world applications. The acquirement of large amounts of sensing data is a fundamental issue in IoT. An efficient way is obtaining sufficient data by the mobile crowdsensing. It is a promising paradigm which leverages the sensing capacity of portable mobile devices. The crowdsensing platform is the key entity who allocates tasks to participants in a mobile crowdsensing system. The strategy of task allocating is crucial for the crowdsensing platform, since it affects the data requester’s confidence, the participant’s confidence, and its own benefit. Traditional allocating algorithms regard the privacy preservation, which may lose the confidence of participants. In this paper, we propose a novel three-step algorithm which allocates tasks to participants with privacy consideration. It maximizes the benefit of the crowdsensing platform and meanwhile preserves the privacy of participants. Evaluation results on both benefit and privacy aspects show the effectiveness of our proposed algorithm. Wenyi Tang, Xu Zheng 0001, Guangchun Luo, Guiduo Duan |
Wirel. Commun. Mob. Comput. | 5 |
| 2015 | Rotation Invariant Texture Retrieval Considering the Scale Dependence of Gabor WaveletabstractObtaining robust and efficient rotation-invariant texture features in content-based image retrieval field is a challenging work. We propose three efficient rotation-invariant methods for texture image retrieval using copula model based in the domains of Gabor wavelet (GW) and circularly symmetric GW (CSGW). The proposed copula models use copula function to capture the scale dependence of GW/CSGW for improving the retrieval performance. It is well known that the Kullback-Leibler distance (KLD) is the commonly used similarity measurement between probability models. However, it is difficult to deduce the closed-form of KLD between two copula models due to the complexity of the copula model. We also put forward a kind of retrieval scheme using the KLDs of marginal distributions and the KLD of copula function to calculate the KLD of copula model. The proposed texture retrieval method has low computational complexity and high retrieval precision. The experimental results on VisTex and Brodatz data sets show that the proposed retrieval method is more effective compared with the state-of-the-art methods. Chaorong Li, Guiduo Duan, Fujin Zhong |
IEEE Trans. Image Process. | 2 |