EDBT 2026 Demo / reviewers in the wild / expert
Dengdi Sun
dblp:01/9924
· DBLP profile ↗
42ranked-venue papers
15as first author
38since 2021 · last 2027
0000-0002-0164-7944ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Neural network-assisted evolutionary search for large-scale sparse multiobjective optimization
Zhuanlian Ding, Junzhe Liu, Dengdi Sun, Xingyi Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Neurological disorder detection based on an adaptive self-supervised multi-scale spatiotemporal interaction network
Changxu Dong, Zongyun Gu, Donghua Li, Xinyuan Xi, Bin Luo 0001, Dengdi Sun |
Expert Syst. Appl. | 6 |
| 2026 | A Unified Hypergraph-Mamba Framework for Adaptive Electroencephalogram Modeling in Multi-view Seizure PredictionabstractSeizure prediction from Electroencephalogram (EEG) signals is a critical task for proactive intervention in epilepsy management. Existing models often struggle to capture high-order inter-channel dependencies dynamically and adapt to the spectral variations preceding seizure onset, especially in cross-patient scenarios. To address these issues, a novel Unified Hypergraph-Mamba (UHM) framework, which for the first time integrates hypergraph-based spatial modeling with Mamba-based adaptive spectral modeling. Specifically, a hypergraph attention mechanism is designed to capture high-order spatial interactions among EEG channels, enabling dynamic representation of inter-channel dependencies. Concurrently, an adaptive spectral modeling module based on the Mamba architecture selectively emphasizes frequency components most indicative of preictal states. Together, these components form a unified architecture capable of jointly modeling spatiotemporal EEG dynamics. Extensive experiments conducted on both patient-specific and cross-patient settings demonstrate that our model consistently outperforms state-of-the-art baselines, achieving superior sensitivity and AUC. Dengdi Sun, Changxu Dong, Zongyun Gu |
Int. J. Neural Syst. | 1 |
| 2026 | Towards higher quality and fewer hallucinations: A multi-agent collaboration framework for LLMs
Shuanghong Shen, Dengdi Sun, Zixuan Qin, Yu Su 0002, Linbo Zhu, Junyu Lu 0003, Zhenya Huang, Shijin Wang 0001 |
Inf. Process. Manag. | 2 |
| 2026 | A manifold embedding-based evolutionary algorithm for many-objective optimization with irregular Pareto front shapes
Zhuanlian Ding, Xihong Jiang, Xingyi Zhang 0001, Dengdi Sun |
Inf. Sci. | 5 |
| 2026 | Structure and progress aware diffusion for medical image segmentation
Siyuan Song, Guyue Hu 0001, Chenglong Li 0002, Dengdi Sun, Zhe Jin 0001, Jin Tang 0001 |
Pattern Recognit. | 4 |
| 2026 | Multiview Graph Contrastive Learning Based on Learnable Graph AugmentationabstractGraph contrastive learning has recently gained prominence as a key technique in the domain of graph representation learning. Most graph contrastive learning methods generate two graph views by augmenting the input graph and maximizing the consistency of their representations. However, due to a lack of prior knowledge of graph data, existing graph augmentation methods often generate low-quality views, which can destroy the core structural information of the graph and thus affect the model’s learning. The fundamental limitation lies in the semantic-agnostic nature of such augmentations, which fail to adapt to the underlying data distribution and the model’s evolving training state. In addition, existing contrastive learning methods typically employ a single contrastive strategy and rely on numerous similarity calculations, which makes it challenging to fully capture the diverse features in a graph. To address these limitations, we propose a multiview graph contrastive learning method based on learnable graph augmentation (MGLGA). Specifically, this method abandons the traditional random graph augmentation, constructs views through dynamic feature learning, and generates feature views using a graph feature learner and a postprocessing module, thereby focusing on the essential structural features of the graph while minimizing interference with the view generation. This learnable paradigm ensures that the augmented views maintain high semantic fidelity and provide adaptive, curriculum-like learning signals throughout the training process, thereby theoretically promoting better generalization. Moreover, we design a three-branch structure to capture both fine-grained and coarse-grained knowledge of the graph through node-level and graph-level contrasts. We further develop a structure-aware group discrimination loss through contrastive objective optimization, enhancing the model’s capacity for capturing graph structural patterns. Comprehensive evaluations demonstrate state-of-the-art performance across multiple downstream benchmarks. Dengdi Sun, Weilong Gong, Bin Luo 0001, Zhuanlian Ding |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2026 | Visual-Label Alignment and Attribute-Aware Prompt for Multi-Label Image Recognition with Partial LabelsabstractThe problem of Multi-Label Image Recognition with Partial Labels (MLIR-PL) is a significant challenge in computer vision, primarily due to the scarcity and high cost of complete annotations. Recent advances have leveraged large-scale vision-language models, such as CLIP, to establish rich correspondences between images and their labels, thereby improving the MLIR-PL performance. However, the existing CLIP-based methods have not fully exploited fine-grained local image features to mitigate interference from semantically irrelevant regions. Moreover, many studies have oversimplified the use of prompt contexts, limiting their ability to comprehensively capture the multi-dimensional attributes of categories. To address these limitations, this article proposes a novel MLIR-PL model with Visual–Label Alignment and Attribute-Aware Prompt (VA \({}^{3}\) P), which sufficiently harnesses the capabilities of large-scale pre-trained vision-language models. In the model, we design a Visual–Label Alignment module to establish a mapping between local image features and category text representations, conspicuously reducing the interference from irrelevant regions. Additionally, our Attribute-Aware Prompt module offers diverse contextual information, providing a more comprehensive representation of the category’s attributes. Extensive experimental results on the COCO 2014 and VOC 2007 datasets, compared with multiple state-of-the-art methods, demonstrate that our model achieves the best performance comprehensively, verifying the advantages of the proposed model in the MLIR-PL task. Dengdi Sun, Hongxing Xie, Zhendong Cai, Leilei Ma 0002, Bin Luo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Texture and Geometry Optimization for 3D Reconstruction
Yanping Fu, Hongjing Zhang, Shaojie Zhang 0002, Dengdi Sun, Haifeng Zhao 0001 |
CGI (1) | 4 |
| 2025 | Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt TuningabstractModeling label correlations has always played a pivotal role in multi-label image classification (MLC), attracting significant attention from researchers. However, recent studies have overemphasized co-occurrence relationships among labels, which can lead to overfitting risk on this overemphasis, resulting in suboptimal models. To tackle this problem, we advocate for balancing correlative and discriminative relationships among labels to mitigate the risk of overfitting and enhance model performance. To this end, we propose the Multi-Label Visual Prompt Tuning framework, a novel and parameter-efficient method that groups classes into multiple class subsets according to label co-occurrence and mutual exclusivity relationships, and then models them respectively to balance the two relationships. In this work, since each group contains multiple classes, multiple prompt tokens are adopted within Vision Transformer (ViT) to capture the correlation or discriminative label relationship within each group, and effectively learn correlation or discriminative representations for class subsets. On the other hand, each group contains multiple group-aware visual representations that may correspond to multiple classes, and the mixture of experts (MoE) model can cleverly assign them from the group-aware to the label-aware, adaptively obtaining label-aware representation, which is more conducive to classification. Experiments on multiple benchmark datasets show that our proposed approach achieves competitive results and outperforms SOTA methods on multiple pre-trained models. Leilei Ma 0002, Ming-Kun Xie, Lei Wang 0095, Dengdi Sun, Haifeng Zhao 0001 |
CVPR | 5 |
| 2025 | Efficient RGBT Tracking via Heterogeneous Hierarchical Knowledge DistillationabstractThe increasing demand for real-time RGBT (RGB and Thermal) tracking in applications such as video surveillance, autonomous driving, and robotic navigation underscores the need for lightweight and efficient tracking frameworks. However, existing approaches require dual-stream architectures for repeated feature extraction, incurring high costs, while complex interaction strategies further reduce efficiency, limiting the real-time performance of RGBT trackers. To address this, we propose a novel Heterogeneous Hierarchical Knowledge Distillation framework (H2KD) to enable a single-stream RGBT tracker that maintains high efficiency while delivering performance comparable to existing dual-stream trackers. In particular, H2KD takes an existing dual-stream network tracker as the teacher and builds a simple single-stream network tracker as the student by concatenating the inputs. To inherit powerful representation of the heterogeneous teacher network, we expand the channel dimensions of single-stream networks to align with the fusion features of teacher network, and employ a hierarchical distillation strategy between their backbone networks. Moreover, H2KD also introduces architecture-independent prediction-level distillation between their prediction score maps to inherit the tracking capability of teachers more directly. Extensive experiments on three major RGBT tracking benchmarks and multiple dual-stream RGBT trackers demonstrate the effectiveness and generalization of the proposed method, which achieves competitive accuracy while achieving 120.2 FPS inference speed. Dengdi Sun, Chenglong Li 0002, Andong Lu |
ICME | 1 |
| 2025 | Towards Space and Semantics: Object-Purified Representation Learning for Multi-Label Image ClassificationabstractMulti-label image classification requires simultaneously recognizing multiple objects with complex interdependencies. While existing attention-based methods are prominent, their performance is hampered by two forms of representation entanglement: 1) Spatial entanglement, where contextual interference from backgrounds and co-occurring objects confuses specific object representations; 2) Semantic entanglement, where models overfit label co-occurrence priors, thereby impairing a genuine semantic understanding of the image. To address these challenges, we propose an Object-Purified Representation Learning framework. Concretely, for spatial entanglement, we propose the Spatial-wise Representation Purification Module that employs Spatial-Purified Attention to eliminate object-irrelevant feature activations for contextual interference reduction, combined with Spatial-Aware Supervision to enhance object perception capability. For semantic entanglement, we develop the Semantic-wise Association Purification Module that synergistically integrates our proposed average message with the original co-occurrence-based message. This design effectively models co-occurrence relationships while preventing their overemphasis. Furthermore, we design the Bidirectional Representation Refinement Module to efficiently enhance representations, further boosting classification performance. Extensive experiments on multiple benchmark datasets with different configurations demonstrate that our proposed method achieves state-of-the-art performance. Haifeng Zhao 0001, Leilei Ma 0002, Lei Wang 0095, Dengdi Sun |
ACM Multimedia | 6 |
| 2025 | Segment Anything Model Meets Semi-supervised Medical Image Segmentation: A Novel PerspectiveabstractThe scarcity of annotated medical imaging data has driven significant progress in semi-supervised learning to alleviate reliance on expensive expert labeling. While foundational vision models such as the Segment Anything Model (SAM) exhibit robust generalization in generic segmentation tasks, their direct application to medical images often results in suboptimal performance. To address this challenge, in this work, we propose a novel fully SAM-based semi-supervised medical image segmentation framework and develop the corresponding knowledge distillation-based learning strategy. Specifically, we first employ an efficient SAM variant as the backbone network of the semi‑supervised framework and update the default prompt embedding of SAM to unleash its full potential. Then, we utilize an original SAM, which is rich in prior knowledge, as the teacher to optimize our efficient student SAM backbone through hierarchical knowledge distillation and a dynamic loss weighting strategy. Extensive experiments on various medical datasets demonstrate that our method outperforms state-of-the-art semi-supervised segmentation approaches. Especially, our model requires less than 10% of the parameter size of the original SAM, enabling substantially lower deployment and storage overhead in real-world clinical settings. Haifeng Zhao 0001, Leilei Ma 0002, Dengdi Sun |
NeurIPS | 4 |
| 2025 | Fully Automated SAM for Single-source Domain Generalization in Medical Image Segmentation
Huanli Zhuo, Leilei Ma 0002, Haifeng Zhao 0001, Dengdi Sun, Yanping Fu |
SMC | 5 |
| 2025 | Semantic knowledge transfer for semi-supervised medical image segmentation
Haifeng Zhao 0001, Leilei Ma 0002, Dengdi Sun |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Multi-view brain network classification based on Adaptive Graph Isomorphic Information Bottleneck Mamba
Changxu Dong, Dengdi Sun, Zhenda Yu, Bin Luo 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Hybrid graph-based radiology report generation
Dengdi Sun, Chaofan Mu, Xuyang Fan, Jin Tang 0001, Zegeng Li, Zhuanlian Ding |
Expert Syst. Appl. | 1 |
| 2025 | DAMR: Multi-scale graph contrastive learning with dynamic adjustment and mutual rectification
Dengdi Sun, Mingwei Cao, Zhifu Tao, Zhuanlian Ding |
Knowl. Based Syst. | 1 |
| 2025 | Dual-level semantic alignment for video moment retrieval and highlight detection
Haifeng Zhao 0001, Wenhai Qin, Leilei Ma 0002, Dengdi Sun |
Multim. Syst. | 5 |
| 2025 | Self-supervised spatial-temporal contrastive network for EEG-based brain network classification
Changxu Dong, Dengdi Sun, Bin Luo 0001 |
Neural Networks | 2 |
| 2025 | Challenge-aware U-net for breast lesion segmentation in ultrasound images
Dengdi Sun, Changxu Dong, Bo Jiang 0002, Yayang Duan, Zhengzheng Tu, Chaoxue Zhang |
Pattern Recognit. | 1 |
| 2025 | TIRAGNN: Temporal and Implicit Relation-Aware Graph Neural Networks for Social RecommendationabstractSocial recommendation systems predict user preferences by using social relationships to address data sparsity and cold-start problems. Since social relations and user–item interactions can naturally be modeled as graph structures, graph neural networks (GNNs) have achieved significant success in social recommendation. However, most existing models fail to incorporate temporal information when modeling user–item interactions and rely solely on explicit social relationships to capture user influence, resulting in suboptimal performance. To address these problems, this article presents a temporal and implicit relation-aware graph neural network for social recommendation (TIRAGNN). Specifically, we model user and item representations using rating and temporal information from user–item interactions, integrating their relational influence within both the social graph and the constructed auxiliary graphs. Additionally, attention mechanisms are employed to model interaction sequences and aggregate relational influences, thereby enhancing the learning of user and item representations. Experimental results on two real-world datasets verify the superiority of TIRAGNN over state-of-the-art approaches. Chenxu Wang 0001, Dengdi Sun, Tao Qin 0002 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Single image shadow removal using 2D signed distance field
Yanping Fu, Dengdi Sun, Shaojie Zhang 0002, Haifeng Zhao 0001 |
Vis. Comput. | 3 |
| 2024 | Text-Region Matching for Multi-Label Image Recognition with Missing LabelsabstractRecently, large-scale visual language pre-trained (VLP) models have demonstrated impressive performance across various downstream tasks. Motivated by these advancements, pioneering efforts have emerged in multi-label image recognition with missing labels, leveraging VLP prompt-tuning technology. However, they usually cannot match text and vision features well, due to complicated semantics gaps and missing labels in a multi-label image. To tackle this challenge, we propose Text-Region Matching for optimizing Multi-Label prompt tuning, namely TRM-ML, a novel method for enhancing meaningful cross-modal matching. Compared to existing methods, we advocate exploring the information of category-aware regions rather than the entire image or pixels, which contributes to bridging the semantic gap between textual and visual representations in a one-to-one matching manner. Concurrently, we further introduce multimodal contrastive learning to narrow the semantic gap between textual and visual modalities and establish intra-class and inter-class relationships. Additionally, to deal with missing labels, we propose a multimodal category prototype that leverages intra- and inter-category semantic relationships to estimate unknown labels, facilitating pseudo-label generation. Extensive experiments on the MS-COCO, PASCAL VOC, Visual Genome, NUS-WIDE, and CUB-200-211 benchmark datasets demonstrate that our proposed framework outperforms the state-of-the-art methods by a significant margin. Our code is available here. Leilei Ma 0002, Hongxing Xie, Lei Wang 0095, Yanping Fu, Dengdi Sun, Haifeng Zhao 0001 |
ACM Multimedia | 5 |
| 2024 | BCS-NeRF: Bundle Cross-Sensing Neural Radiance Fields
Mingwei Cao, Fengna Wang, Dengdi Sun, Haifeng Zhao 0001 |
MMAsia | 3 |
| 2024 | Domain Adaptive Lung Nodule Detection in X-Ray ImageabstractMedical images from different healthcare centers exhibit varied data distributions, posing significant challenges for adapting lung nodule detection due to the domain shift between training and application phases. Traditional unsupervised domain adaptive detection methods often struggle with this shift, leading to suboptimal outcomes. To overcome these challenges, we introduce a novel domain adaptive approach for lung nodule detection that leverages mean teacher self-training and contrastive learning. First, we propose a hierarchical contrastive learning strategy to refine nodule representations and enhance the distinction between nodules and background. Second, we introduce a nodule-level domain-invariant feature learning (NDL) module to capture domain-invariant features through adversarial learning across different domains. Additionally, we propose a new annotated dataset of X-ray images to aid in advancing lung nodule detection research. Extensive experiments conducted on multiple X-ray datasets demonstrate the efficacy of our approach in mitigating domain shift impacts. Haifeng Zhao 0001, Lixiang Jiang, Leilei Ma 0002, Dengdi Sun, Yanping Fu |
SMC | 4 |
| 2024 | Hybrid attention mechanism of feature fusion for medical image segmentationabstractAbstract Traditional convolution neural networks (CNN) have achieved good performance in multi‐organ segmentation of medical images. Due to the lack of ability to model long‐range dependencies and correlations between image pixels, CNN usually ignores the information of channel dimension. To further improve the performance of multi‐organ segmentation, a hybrid attention mechanism model is proposed. First, a CNN was used to extract multi‐scale feature maps and fed into the Channel Attention Enhancement Module (CAEM) to selectively pay attention to target organs in medical images, and the Transformer encoded tokenized image patches from CNN feature maps as the input sequence to model long‐range dependencies. Second, the decoder upsampled the output from Transformer and fused with the CAEM features in multi‐scale through skip connections. Finally, we introduced a Refinement Module (RM) after the decoder to improve feature correlations of the same organ and the feature discriminability between different organs. The model outperformed on dice coefficient (%) and hd95 on both the synapse multi‐organ segmentation and cardiac diagnosis challenge datasets. The hybrid attention mechanisms exhibited high efficiency and high segmentation accuracy in medical images. Shanshan Tong, Zhentao Zuo, Zuxiang Liu, Dengdi Sun, Tiangang Zhou |
IET Image Process. | 4 |
| 2024 | Spatial-Temporal Dynamic Hypergraph Information Bottleneck for Brain Network ClassificationabstractRecently, Graph Neural Networks (GNNs) have gained widespread application in automatic brain network classification tasks, owing to their ability to directly capture crucial information in non-Euclidean structures. However, two primary challenges persist in this domain. First, within the realm of clinical neuro-medicine, signals from cerebral regions are inevitably contaminated with noise stemming from physiological or external factors. The construction of brain networks heavily relies on set thresholds and feature information within brain regions, making it susceptible to the incorporation of such noises into the brain topology. Additionally, the static nature of the artificially constructed brain network's adjacent structure restricts real-time changes in brain topology. Second, mainstream GNN-based approaches tend to focus solely on capturing information interactions of nearest neighbor nodes, overlooking high-order topology features. In response to these challenges, we propose an adaptive unsupervised Spatial-Temporal Dynamic Hypergraph Information Bottleneck (ST-DHIB) framework for dynamically optimizing brain networks. Specifically, adopting an information theory perspective, Graph Information Bottleneck (GIB) is employed for purifying graph structure, and dynamically updating the processed input brain signals. From a graph theory standpoint, we utilize the designed Hypergraph Neural Network (HGNN) and Bi-LSTM to capture higher-order spatial-temporal context associations among brain channels. Comprehensive patient-specific and cross-patient experiments have been conducted on two available datasets. The results demonstrate the advancement and generalization of the proposed framework. Changxu Dong, Dengdi Sun |
Int. J. Neural Syst. | 2 |
| 2024 | UAV-Ground Visual Tracking: A Unified Dataset and Collaborative Learning ApproachabstractVisual tracking from the ground view and the UAV view has received increasing attention due to its wide range of practical applications. These two tasks have strong complementary benefits in the description of the target object, such as detailed appearance in the ground view and global motion information in the UAV view, and their combination has the potential to allow the tracking system to be more robust. However, no work has studied this problem in-depth, and it is challenging to accurately combine the ground view information and the UAV view information. To fill the gap and address the challenge, we propose a new computer vision task called UAV-Ground visual tracking. Considering the lack of relevant data and methods, we first propose a unified video dataset called UGVT, which includes 210 pairs of UAV and ground high-resolution video sequences with a total of more than 204K frames, which can be used as a comprehensive evaluation platform for relevant tracking methods. Secondly, based on the newly constructed dataset, we propose a co-learning method called MvCL to fuse the information of ground and UAV views. It first associates the same tracking target in the two views based on cross-attention operation and then fuses the complementary information of the two views. In particular, as a plug-and-play module based on Transformer structure, this method can be flexibly embedded into different tracking frameworks. Extensive experiments are conducted on the newly created dataset. The results demonstrate the effectiveness of the proposed method in improving the robustness of the tracking system compared with 10 state-of-the-art tracking methods and also indicate the prospect and significance of potential UAV-Ground visual tracking research. The dataset is available at:https://github.com/mmic-lcl/Datasets-and-benchmark-code/. Dengdi Sun, Leilei Cheng, Chenglong Li 0002, Yun Xiao 0003, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Transformer RGBT Tracking With Spatio-Temporal Multimodal TokensabstractMany RGBT tracking researches primarily focus on modal fusion design, while overlooking the effective handling of target appearance changes. While some approaches have introduced historical frames or fuse and replace initial templates to incorporate temporal information, they have the risk of disrupting the original target appearance and accumulating errors over time. To alleviate these limitations, we propose a novel Transformer RGBT tracking approach, which mixes spatio-temporal multimodal tokens from the static multimodal templates and multimodal search regions in Transformer to handle target appearance changes, for robust RGBT tracking. We introduce independent dynamic template tokens to interact with the search region, embedding temporal information to address appearance changes, while also retaining the involvement of the initial static template tokens in the joint feature extraction process to ensure the preservation of the original reliable target appearance information that prevent deviations from the target appearance caused by traditional temporal updates. We also use attention mechanisms to enhance the target features of multimodal template tokens by incorporating supplementary modal cues, and make the multimodal search region tokens interact with multimodal dynamic template tokens via attention mechanisms, which facilitates the conveyance of multimodal-enhanced target change information. Our module is inserted into the transformer backbone network and inherits joint feature extraction, search-template matching, and cross-modal interaction. Extensive experiments on three RGBT benchmark datasets show that the proposed approach maintains competitive performance compared to other state-of-the-art tracking algorithms while running at 39.1 FPS. The project-related materials are available at:https://github.com/yinghaidada/STMT. Dengdi Sun, Yajie Pan, Andong Lu, Chenglong Li 0002, Bin Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Efficient Sparse Large-Scale Multiobjective Optimization Based on Cross-Scale Knowledge FusionabstractDue to the curse of dimensionality and the unknown sparsity of search spaces, evolutionary algorithms face immense challenges in approximating optimal solutions for widely studied sparse large-scale multiobjective optimization problems (SLMOPs). Most bilevel encoding scheme (BLES)-based algorithms primarily focus on exploring sparsity in the binary layer, neglecting the real layer. Moreover, the interactions between two layers may be disregarded in these algorithms, thus the latent gap between the two encoding scales could lead to evolutionary ambiguity and performance limitations. To tackle the above issues, this article proposes a novel BLES-based collaborative algorithm using cross-scale knowledge fusion for SLMOPs. The algorithm integrates dual grouping and dual dimension reduction techniques via two subpopulations in a coevolutionary manner. Additionally, the interaction strategy is designed for each technique, leveraging the binary layer to guide the real layer, thus facilitating sufficient cross-scale cooperation. Extensive experiments on benchmark SLMOPs and four real-world applications validate the proposed algorithm’s strong competitiveness in solving SLMOPs compared to state-of-the-art algorithms. Zhuanlian Ding, Dengdi Sun, Xingyi Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2023 | Semantic-Aware Dual Contrastive Learning for Multi-Label Image ClassificationabstractExtracting image semantics effectively and assigning corresponding labels to multiple objects or attributes for natural images is challenging due to the complex scene contents and confusing label dependencies. Recent works have focused on modeling label relationships with graph and understanding object regions using class activation maps (CAM). However, these methods ignore the complex intra- and inter-category relationships among specific semantic features, and CAM is prone to generate noisy information. To this end, we propose a novel semantic-aware dual contrastive learning framework that incorporates sample-to-sample contrastive learning (SSCL) as well as prototype-to-sample contrastive learning (PSCL). Specifically, we leverage semantic-aware representation learning to extract category-related local discriminative features and construct category prototypes. Then based on SSCL, label-level visual representations of the same category are aggregated together, and features belonging to distinct categories are separated. Meanwhile, we construct a novel PSCL module to narrow the distance between positive samples and category prototypes and push negative samples away from the corresponding category prototypes. Finally, the discriminative label-level features related to the image content are accurately captured by the joint training of the above three parts. Experiments on five challenging large-scale public datasets demonstrate that our proposed method is effective and outperforms the state-of-the-art methods. Code and supplementary materials are released on https://github.com/yu-gi-oh-leilei/SADCL. Leilei Ma 0002, Dengdi Sun, Lei Wang 0095, Haifeng Zhao 0001, Bin Luo 0001 |
ECAI | 2 |
| 2023 | Large-scale multimodal multiobjective evolutionary optimization based on hybrid hierarchical clustering
Zhuanlian Ding, Lve Cao, Dengdi Sun, Xingyi Zhang 0001, Zhifu Tao |
Knowl. Based Syst. | 4 |
| 2023 | EGARNet: adjacent residual lightweight super-resolution network based on extended group-enhanced convolution
Longfeng Shen, Fenglan Qin, Hongying Zhu, Dengdi Sun, Hai Min |
Multim. Syst. | 4 |
| 2022 | SILP-autoencoder for face de-occlusion
Dengdi Sun, Wandong Xie, Zhuanlian Ding, Jin Tang 0001 |
Neurocomputing | 1 |
| 2022 | Multitask Multigranularity Aggregation With Global-Guided Attention for Video Person Re-IdentificationabstractThe goal of video-based person re-identification (Re-ID) is to identify the same person across multiple non-overlapping cameras. The key to accomplishing this challenging task is to sufficiently exploit both spatial and temporal cues in video sequences. However, most current methods are incapable of accurately locating semantic regions or efficiently filtering discriminative spatio-temporal features; so it is difficult to handle issues such as spatial misalignment and occlusion. Thus, we propose a novel feature aggregation framework, multi-task and multi-granularity aggregation with global-guided attention (MMA-GGA), which aims to adaptively generate more representative spatio-temporal aggregation features. Specifically, we develop a multi-task multi-granularity aggregation (MMA) module to extract features at different locations and scales to identify key semantic-aware regions that are robust to spatial misalignment. Then, to determine the importance of the multi-granular semantic information, we propose a global-guided attention (GGA) mechanism to learn weights based on the global features of the video sequence, allowing our framework to identify stable local features while ignoring occlusions. Therefore, the MMA-GGA framework can efficiently and effectively capture more robust and representative features. Extensive experiments on four benchmark datasets demonstrate that our MMA-GGA framework outperforms current state-of-the-art methods. In particular, our method achieves a rank-1 accuracy of 91.0% on the MARS dataset, the most widely used database, significantly outperforming existing methods. Dengdi Sun, Jin Tang 0001, Zhuanlian Ding |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | LasHeR: A Large-Scale High-Diversity Benchmark for RGBT TrackingabstractRGBT tracking receives a surge of interest in the computer vision community, but this research field lacks a large-scale and high-diversity benchmark dataset, which is essential for both the training of deep RGBT trackers and the comprehensive evaluation of RGBT tracking methods. To this end, we present a La rge- s cale H igh-diversity [Formula: see text]nchmark for short-term R GBT tracking (LasHeR) in this work. LasHeR consists of 1224 visible and thermal infrared video pairs with more than 730K frame pairs in total. Each frame pair is spatially aligned and manually annotated with a bounding box, making the dataset well and densely annotated. LasHeR is highly diverse capturing from a broad range of object categories, camera viewpoints, scene complexities and environmental factors across seasons, weathers, day and night. We conduct a comprehensive performance evaluation of 12 RGBT tracking algorithms on the LasHeR dataset and present detailed analysis. In addition, we release the unaligned version of LasHeR to attract the research interest for alignment-free RGBT tracking, which is a more practical task in real-world applications. The datasets and evaluation protocols are available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code. Chenglong Li 0002, Wanlin Xue, Yaqing Jia, Zhichen Qu, Bin Luo 0001, Jin Tang 0001, Dengdi Sun |
IEEE Trans. Image Process. | 7 |
| 2021 | Dual-decoder graph autoencoder for unsupervised graph representation learning
Dengdi Sun, Dashuang Li, Zhuanlian Ding, Xingyi Zhang 0001, Jin Tang 0001 |
Knowl. Based Syst. | 1 |
| 2020 | Information Enhanced Graph Convolutional Networks for Skeleton-based Action RecognitionabstractSkeleton-based action recognition has recently attracted much attention in computer vision. The latest methods are mostly based on graph convolutional networks (GCNs), which construct the human body as spatial-temporal Skeleton graphs, and has achieved excellent performance. However, previous studies only capture the local and rough information based on the physical dependencies among joints, which may miss implicit joint correlations. In this work, we propose a novel action recognition model, namely Information Enhanced Graph Convolutional Networks (IE-GCN). To improve the accuracy and robustness of recognition, this model capture higher-order dependency in the skeleton-based graph by expanding the joint neighbors, and combine second stage skeleton features (the lengths and directions of bones) to enhance the discriminative information simultaneously. In addition, an training strategy is designed to solve the framework. Extensive experiments on two large-scale public datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate the superior performance of the proposed algorithms over the state-of-the-art methods. Dengdi Sun, Fanchen Zeng, Bin Luo 0001, Jin Tang 0001, Zhuanlian Ding |
IJCNN | 1 |
| 2018 | Low-rank subspace learning based network community detection
Zhuanlian Ding, Xingyi Zhang 0001, Dengdi Sun, Bin Luo 0001 |
Knowl. Based Syst. | 3 |
| 2017 | Protein functional annotation refinement based on graph regularized ℓ1-norm PCA
Dengdi Sun, Huadong Liang, Meiling Ge, Zhuanlian Ding, Wan-Ting Cai, Bin Luo 0001 |
Pattern Recognit. Lett. | 1 |
| 2011 | Angular DecompositionabstractDimensionality reduction plays a vital role in pattern recognition. However, for normalized vector data, existing methods do not utilize the fact that the data is normalized. In this paper, we propose to employ an Angular Decomposition of the normalized vector data which corresponds to embedding them on a unit surface. On graph data for similarity/ kernel matrices with constant diagonal elements, we propose the Angular Decomposition of the similarity matrices which corresponds to embedding objects on a unit sphere. In these angular embeddings, the Euclidean distance is equivalent to the cosine similarity. Thus data structures best described in the cosine similarity and data structures best captured by the Euclidean distance can both be effectively detected in our angular embedding. We provide the theoretical analysis, derive the computational algorithm, and evaluate the angular embedding on several datasets. Experiments on data clustering demonstrate that our method can provide a more discriminative subspace. Dengdi Sun, Chris Ding, Bin Luo 0001, Jin Tang 0001 |
IJCAI | 1 |