Donglin Di

dblp:242/4517 · DBLP profile ↗
← Back
56ranked-venue papers
7as first author
52since 2021 · last 2026
0000-0002-2270-3378ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 3 first-author · 26 since 2021Artificial intelligence and machine learning · 23 · 4 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 16 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DiTalker: A unified DiT-based framework for high-quality and style-controllable portrait animation
Yongjia Ma, Lei Fan 0007, Donglin Di, Tonghua Su
Comput. Vis. Image Underst.4
2026 OSTE: Omni-Scene Text Editing with Latent Decoupling
Tonghua Su, Fuxiang Yang, Lei Fan 0007, Donglin Di, Zhongjie Wang 0003, Xiangqian Wu 0002
Comput. Vis. Image Underst.4
2026 M 3 Surv : Fusing Multi-slide and Multi-omics for Memory-augmented robust Survival prediction
abstract
Multimodal survival prediction is crucial for personalized oncology. However, existing methods typically integrate only Formalin-Fixed Paraffin-Embedded (FFPE) slides with a single omics type, such as genomics, overlooking Fresh Frozen (FF) slides that better preserve molecular information, as well as richer multi-omics data like proteomics and transcriptomics. More critically, the complete absence of certain modalities due to clinical constraints ( e.g. , time or cost) severely limits the applicability of conventional fusion models that rely on inter-modality correlations. To address these gaps, we propose M 3 Surv, a framework designed to integrate multi-pathology slides (both FF and FFPE) with multi-omics profiles. For multi-slide fusion, we design a divide-and-conquer hypergraph learning approach to capture both intra-slide higher-order cellular structures and inter-slide relationships, yielding a unified pathology representation. To enrich the biological context, we integrate multi-omics data and employ interactive cross-attention to fuse the pathological and omics modalities. To tackle the missing modality, we introduce a prototype-based memory bank. During training, this memory bank learns and stores representative pathology-omics feature prototypes. At inference, even if a modality is entirely missing, the model can query the bank with available features and robustly impute information from the most similar prototype. Extensive experiments on five TCGA cancer datasets and an in-house dataset demonstrate that M 3 Surv outperforms state-of-the-art methods, achieving an average 2.2% improvement in C-Index. The framework also shows strong stability across various missing modality scenarios, highlighting its clinical potential in real-world, data-incomplete scenarios.
Mingcheng Qu, Donglin Di, Yue Gao 0002, Yang Song 0001, Lei Fan 0007
Medical Image Anal.3
2026 STAG: Biologically guided spatial transcriptomics prediction via hypergraph learning
abstract
Spatial transcriptomics (ST) enables spatially resolved gene expression profiling within intact tissue sections. However, its widespread adoption is constrained by the high cost and low throughput of current sequencing-based protocols. This has motivated growing interest in computationally predicting gene expression directly from routinely acquired histology images. Existing methods are largely restricted to isolated 2D tissue slices and fail to capture richer spatial relationships or structured dependencies among spot-level gene expression profiles. In this paper, we propose STAG, a dual-branch framework for gene-aware expression prediction and spatial context modeling. A Query branch predicts ST expression for an individual target spot, while a Neighbor branch acts as an auxiliary branch to model structured relationships among multiple spots. By leveraging hypergraph learning, the Neighbor branch captures higher-order spatial and molecular dependencies, enabling unified modeling of both intra-slice and inter-slice relationships. This design supports standard 2D settings (a single slice) and naturally extends to 3D scenarios when adjacent tissue sections are available. Moreover, STAG leverages gene semantic information as biological guidance by encoding gene names with a foundation model, enabling coordinated gene-aware interactions beyond independent gene prediction. STAG achieves an average gain of 5.16% in PCC@250 across six datasets. Under highly variable gene selection, STAG maintains the lowest RMSE and highest PCC@50 across three datasets. The effectiveness of the learned representations is further demonstrated in pseudo-3D prediction and downstream cancer classification tasks. Code is available at https://github.com/MCPathology/STAG.
Mingcheng Qu, Yuchuan Zhao, Donglin Di, Xiu Su, Hongyan Xu 0002, Yang Song 0001, Lei Fan 0007
Medical Image Anal.4
2026 Learning priority-aware controllable poster layout generation
Fuxiang Yang, Wendi Hou, Lei Fan 0007, Tonghua Su, Lingxiao He, Chengzhou Li, Meng Wang 0001, Qianlong Xie, Donglin Di, Xun Yang 0001
Pattern Recognit.10
2026 Noise-aware cross attention for image manipulation localization
abstract
• A Gated Noise Extractor that dynamically captures noise features from multiple strategies. • Dual-granularity contrastive learning for more discriminative noise extraction. • Noise-domain guided fusion module t • reduce interference from irrelevant in- formation. • An efficient model with low parameter count and computational complexity. Modern image manipulation techniques have achieved visual realism that often deceives the human eye and semantic-based detectors. However, manipulation operations typically disturb the intrinsic statistical properties of images. Unlike high-level semantic content, which remains visually consistent, such disturbances manifest as anomalies in noise characteristics, including inconsistencies in sensor pattern noise, distinct high-frequency residuals, and unnatural frequency-domain artifacts introduced by resampling or synthesis. These subtle forensic cues provide more reliable evidence for manipulation localization but are often suppressed by standard RGB-domain feature extractors. Existing IML methods often rely on a single noise feature extraction strategy or treat all tampering techniques uniformly, leading to two major limitations, incomplete noise characterization and insufficient tampering-type awareness . We propose a Noise-aware Contrastive localization Network (NC-Net), which introduces two key modules. Firstly, a Gated Noise Extractor that captures mixed noise-domain patterns using a gated network combining features derived from BayarConv and Discrete Wavelet Transform (DWT) operations. This extractor is further enhanced by a dual-granularity contrastive learning strategy, which models distributional discrepancies both within images (between manipulated and authentic regions) and across images (among different manipulation types). Secondly, a Multi-Scale Fusion Module that adaptively integrates noise-domain and RGB-domain semantic features via a cross-domain attention mechanism and a top-down feature pyramid. A lightweight decoder then produces the final localization map with high precision. NC-Net enables end-to-end joint optimization of the noise extraction and RGB branches, achieving state-of-the-art performance with competitive computational overhead. Extensive experiments demonstrate its superiority over existing methods. Source code is available at https://github.com/HIT-liar/NC-Net .
Hongshi Zhang, Tonghua Su, Fuxiang Yang, Donglin Di, Yang Song 0001, Lei Fan 0007
Pattern Recognit.5
2026 Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
abstract
Creating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotemporal inconsistencies and high computational resource demands. We propose Global-Local Collaborative Diffusion (GLC-Diffusion), a tuning-free method for long video generation. It models the long video denoising process by establishing denoising trajectories through Global-Local Collaborative Denoising (GLCD) to ensure overall content consistency and temporal coherence between frames. Additionally, we introduce a Noise Reinitialization strategy which combines local noise shuffling with frequency fusion to improve global content consistency and visual diversity. Further, we propose a Video Motion Consistency Refinement (VMCR) module that computes the gradient of pixel-wise and frequency-wise losses to enhance visual consistency and temporal smoothness. Extensive experiments, including quantitative and qualitative evaluations on videos of varying lengths (e.g., 3× and 6× longer), demonstrate that our method effectively integrates with existing video diffusion models, producing coherent, high-fidelity long videos superior to previous approaches.
Yongjia Ma, Junlin Chen, Donglin Di, Qi Xie 0009, Lei Fan 0007, Wei Chen 0089, Na Zhao 0004, Xun Yang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2025 MV-VTON: Multi-View Virtual Try-On with Diffusion Models
abstract
The goal of image-based virtual try-on is to generate an image of the target person naturally wearing the given clothing. However, existing methods solely focus on the frontal try-on using the frontal clothing. When the views of the clothing and person are significantly inconsistent, particularly when the person's view is non-frontal, the results are unsatisfactory. To address this challenge, we introduce Multi-View Virtual Try-ON (MV-VTON), which aims to reconstruct the dressing results from multiple views using the given clothes. Given that single-view clothes provide insufficient information for MV-VTON, we instead employ two images, i.e., the frontal and back views of the clothing, to encompass the complete view as much as possible. Moreover, we adopt diffusion models that have demonstrated superior abilities to perform our MV-VTON. In particular, we propose a view-adaptive selection method where hard-selection and soft-selection are applied to the global and local clothing feature extraction, respectively. This ensures that the clothing features are roughly fit to the person's view. Subsequently, we suggest joint attention blocks to align and fuse clothing features with person features. Additionally, we collect a MV-VTON dataset MVG, in which each person has multiple photos with diverse views and poses. Experiments show that the proposed method not only achieves state-of-the-art results on MV-VTON task using our MVG dataset, but also has superiority on frontal-view virtual try-on task using VITON-HD and DressCode datasets.
Zhilu Zhang 0001, Donglin Di, Shiliang Zhang, Wangmeng Zuo
AAAI3
2025 GRPose: Learning Graph Relations for Human Image Generation with Pose Priors
abstract
Recent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this paper, we propose a framework that delves into the graph relations of pose priors to provide control information for human image generation. The main idea is to establish a graph topological structure between the pose priors and latent representation of diffusion models to capture the intrinsic associations between different pose parts. A Progressive Graph Integrator (PGI) is designed to learn the spatial relationships of the pose priors with the graph structure, adopting a hierarchical strategy within an Adapter to gradually propagate information across different pose parts. Besides, a pose perception loss is introduced based on a pretrained pose estimation network to minimize the pose differences. Extensive qualitative and quantitative experiments conducted on the Human-Art and LAION-Human datasets clearly demonstrate that our model can achieve significant performance improvement over the latest benchmark models.
Xiangchen Yin, Donglin Di, Lei Fan 0007, Hao Li 0030, Wei Chen 0089, Gouxiao Fei, Yang Song 0001, Xiao Sun 0003, Xun Yang 0001
AAAI2
2025 Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation Learning
abstract
Universal, transferable whole-slide image (WSI) representations are central to computational pathology. Incorporating multiple markers (e.g., immunohistochemistry, IHC) alongside H&E enriches H&E-based features with diverse, biologically meaningful information. However, progress is limited by the scarcity of well-aligned multi-stain datasets. Inter-stain Misalignment shifts corresponding tissue across slides, hindering consistent patch-level features and degrading slide-level embeddings. To address this, we curated a slide-level aligned, five-stain dataset (H&E, HER2, KI67, ER, PGR) to enable paired H&E-IHC learning and robust cross-stain representation. Leveraging this dataset, we propose Cross-Stain Contrastive Learning (CSCL), a two-stage pretraining framework: a lightweight adapter trained with patch-wise contrastive alignment to improve the compatibility of H&E features with corresponding IHC-derived contextual cues; and slide-level representation learning with Multiple Instance Learning (MIL), which uses a cross-stain attention fusion module to integrate stain-specific patch features and a crossstain global alignment module to enforce consistency among slide-level embeddings across different stains. Experiments on cancer subtype classification, IHC biomarker status classification, and survival prediction, show consistent gains by yielding high-quality, transferable H&E slide-level representations. The code and data are available at: https://github.com/lily-zyz/CSCL.
Yizhi Zhang, Lei Fan 0007, Zhulin Tao, Donglin Di, Yang Song 0001, Sidong Liu, Cong Cong 0001
BIBM4
2025 MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects
abstract
We present MANTA, a visual-text anomaly detection dataset for tiny objects. The visual component comprises over 137.3K images across 38 object categories spanning five typical domains, of which 8.6K images are labeled as anomalous with pixel-level annotations. Each image is captured from five distinct viewpoints to ensure comprehensive object coverage. The text component consists of two subsets: Declarative Knowledge, including 875 words that describe common anomalies across various domains and specific categories, with detailed explanations for ⟨what, why, how⟩, including causes and visual characteristics; and Constructivist Learning, providing 2K multiple-choice questions with varying levels of difficulty, each paired with images and corresponded answer explanations. We also propose a baseline for visual-text tasks and conduct extensive benchmarking experiments to evaluate advanced methods across different settings, highlighting the challenges and efficacy of our dataset.
Lei Fan 0007, Dongdong Fan, Zhiguang Hu, Yiwen Ding, Donglin Di, Kai Yi, Maurice Pagnucco, Yang Song 0001
CVPR5
2025 MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation
abstract
The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frameworks for modeling single basic emotional expressions, which restricts the generation of complex emotions such as compound emotions; b) the lack of comprehensive datasets rich in human emotional expressions, which limits the potential of models. To address these challenges, we propose the following innovations: 1) the Mixture of Emotion Experts (MoEE) model, which decouples six fundamental emotions to enable the precise synthesis of both singular and compound emotional states; 2) the DH-FaceEmoVid-150 dataset, specifically curated to include six prevalent human emotional expressions as well as four types of compound emotions, thereby expanding the training potential of emotion-driven models. Furthermore, to enhance the flexibility of emotion control, we propose an emotion-to-latents module that leverages multimodal inputs, aligning diverse control signals—such as audio, text, and labels—to ensure more varied control inputs as well as the ability to control emotions using audio alone. Through extensive quantitative and qualitative evaluations, we demonstrate that the MoEE framework, in conjunction with the DH-FaceEmoVid-150 dataset, excels in generating complex emotional expressions and nuanced facial details, setting a new benchmark in the field. These datasets will be publicly released.
Huaize Liu, Wenzhang Sun, Donglin Di, Shibo Sun, Changqing Zou, Hujun Bao
CVPR3
2025 DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
Donglin Di, Wenzhang Sun, Yongjia Ma, Hao Li 0030, Wei Chen 0089, Lei Fan 0007, Tonghua Su, Xun Yang 0001
ICCV1
2025 Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly Detection
Lei Fan 0007, Donglin Di, Anyang Su, Tianyou Song, Maurice Pagnucco, Yang Song 0001
ICCV3
2025 QR-LoRA: Efficient and Disentangled Fine-Tuning via QR Decomposition for Customized Generation
abstract
Existing text-to-image models often rely on parameter fine-tuning techniques such as Low-Rank Adaptation (LoRA) to customize visual attributes. However, when combining multiple LoRA models for content-style fusion tasks, unstructured modifications of weight matrices often lead to undesired feature entanglement between content and style attributes. We propose QR-LoRA, a novel fine-tuning framework leveraging QR decomposition for structured parameter updates that effectively separate visual attributes. Our key insight is that the orthogonal Q matrix naturally minimizes interference between different visual features, while the upper triangular R matrix efficiently encodes attribute-specific transformations. Our approach fixes both Q and R matrices while only training an additional task-specific $ΔR$ matrix. This structured design reduces trainable parameters to half of conventional LoRA methods and supports effective merging of multiple adaptations without cross-contamination due to the strong disentanglement properties between $ΔR$ matrices. Experiments demonstrate that QR-LoRA achieves superior disentanglement in content-style fusion tasks, establishing a new paradigm for parameter-efficient, disentangled fine-tuning in generative models. The project page is available at: https://luna-ai-lab.github.io/QR-LoRA/.
Yongjia Ma, Donglin Di, Jianxun Cui, Hao Li 0030, Wei Chen 0089, Xun Yang 0001, Wangmeng Zuo
ICCV3
2025 EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
abstract
Diffusion models are highly regarded for their controllability and the diversity of images they generate. However, class-conditional generation methods based on diffusion models often focus on more common categories. In large-scale fine-grained image generation, issues of semantic information entanglement and insufficient detail in the generated images still persist. This paper attempts to introduce a concept of a "tiered embedder" in fine-grained image generation, which integrates semantic information from both super and child classes, allowing the diffusion model to better incorporate semantic information and address the issue of semantic entanglement. To address the issue of insufficient detail in fine-grained images, we introduce the concept of super-resolution during the perceptual information generation stage, enhancing the detailed features of fine-grained images through enhancement and degradation models. Furthermore, we propose an efficient ProAttention mechanism that can be effectively implemented in the diffusion model. We evaluate our method through extensive experiments on public benchmarks, demonstrating that our approach outperforms other state-of-the-art fine-tuning methods in terms of performance.
Donglin Di, Tonghua Su, Lei Fan 0007
ICME2
2025 Global-Local Aware Scene Text Editing
abstract
Scene Text Editing (STE) involves replacing text in a scene image with new target text while preserving both the original text style and background texture. Existing methods suffer from two major challenges: inconsistency and length-insensitivity. They often fail to maintain coherence between the edited local patch and the surrounding area, and they struggle to handle significant differences in text length before and after editing. To tackle these challenges, we propose an end-to-end framework called Global-Local Aware Scene Text Editing (GLASTE), which simultaneously incorporates high-level global contextual information along with delicate local features. Specifically, we design a global-local combination structure, joint global and local losses, and enhance text image features to ensure consistency in text style within local patches while maintaining harmony between local and global areas. Additionally, we express the text style as a vector independent of the image size, which can be transferred to target text images of various sizes. We use an affine fusion to fill target text images into the editing patch while maintaining their aspect ratio unchanged. Extensive experiments on real-world datasets validate that our GLASTE model outperforms previous methods in both quantitative metrics and qualitative results and effectively mitigates the two challenges.
Fuxiang Yang, Tonghua Su, Donglin Di, Xiangqian Wu 0002, Zhongjie Wang 0003, Lei Fan 0007
ICME3
2025 Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance
abstract
Multimodal pathology-genomic analysis has become increasingly prominent in cancer survival prediction. However, existing studies mainly utilize multi-instance learning to aggregate patch-level features, neglecting the information loss of contextual and hierarchical details within pathology images. Furthermore, the disparity in data granularity and dimensionality between pathology and genomics leads to a significant modality imbalance. The high spatial resolution inherent in pathology data renders it a dominant role while overshadowing genomics in multimodal integration. In this paper, we propose a multimodal survival prediction framework that incorporates hypergraph learning to effectively capture both contextual and hierarchical details from pathology images. Moreover, it employs a modality rebalance mechanism and an interactive alignment fusion strategy to dynamically reweight the contributions of the two modalities, thereby mitigating the pathology-genomics imbalance. Quantitative and qualitative experiments are conducted on five TCGA datasets, demonstrating that our model outperforms advanced methods by over 3.4% in C-Index performance. Code: https://github.com/MCPathology/MRePath.
Mingcheng Qu, Donglin Di, Tonghua Su, Yue Gao 0002, Yang Song 0001, Lei Fan 0007
IJCAI3
2025 Spatially Gene Expression Prediction Using Dual-Scale Contrastive Learning
Mingcheng Qu, Yuncong Wu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007
MICCAI (15)3
2025 Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning
Mingcheng Qu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007
MICCAI (10)3
2025 Hypergraph Tversky-Aware Domain Incremental Learning for Brain Tumor Segmentation with Missing Modalities
Junze Wang, Lei Fan 0007, Weipeng Jing 0001, Donglin Di, Yang Song 0001, Sidong Liu, Cong Cong 0001
MICCAI (11)4
2025 LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
abstract
While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfactual reasoning remains underexplored. To tackle these challenges, we introduce LLaPa, a vision-language model framework designed for multimodal procedural planning. LLaPa generates executable action sequences from textual task descriptions and visual environmental images using vision-language models (VLMs). Furthermore, we enhance LLaPa with two auxiliary modules to improve procedural planning. The first module, the Task-Environment Reranker (TER), leverages task-oriented segmentation to create a task-sensitive feature space, aligning textual descriptions with visual environments and emphasizing critical regions for procedural execution. The second module, the Counterfactual Activities Retriever (CAR), identifies and emphasizes potential counterfactual conditions, enhancing the model's reasoning capability in counterfactual scenarios. Extensive experiments on ActPlan-1K and ALFRED benchmarks demonstrate that LLaPa generates higher-quality plans with superior LCS and correctness, outperforming advanced models. The code and models are available https://github.com/sunshibo1234/LLaPa.
Shibo Sun, Xue Li 0011, Donglin Di, Lanshun Nie, Weinan Zhang 0003, Dechen Zhan, Yang Song 0001, Lei Fan 0007
ACM Multimedia3
2025 PRISM: A Benchmark for Unveiling Cross-modal Knowledge Inconsistency in Large Vision-Language Models
abstract
Recent advances in Large Vision-Language Models (LVLMs) have unearthed boosted performance of multi-modal understanding. In this paper, however, we for the first time uncover a critically under-explored challenge persisting in this trend, that LVLMs unfortunately exhibit cross-modal knowledge inconsistencies. Cross-modal knowledge inconsistency refers to the tendency of providing semantically inconsistent responses to contexts that are semantically equivalent but expressed in different modalities. In real-world applications, users can rely on either text or image to express their ideas. Inconsistent responses across modalities can confuse users, challenging the reliabilities of LVLMs in practice. Therefore, we argue that evaluating performance on either multi-modal or text-only task is insufficient; and waiving the mentioned cross-modal knowledge inconsistency is crucial. The paper proposes PRISM, the first-ever benchmark for measuring the inconsistency, and the corresponding evaluation metric Know-Inc. PRISM covers commonsense, encyclopedia, and mathematics knowledge, with manually-screened samples of semantic alignment. From the evaluation results of up to 27 LVLMs with diverse structures, we conclude that: 1) LVLMs show a preference for textual input, 2) there is a correlation between inconsistency and accuracy, and 3) the inconsistency is more prominent in encyclopedia knowledge. These findings can shed light on further optimization and development of LVLMs.
Weinan Zhang 0003, Donglin Di, Wei Chen 0089, Ting Liu 0001
ACM Multimedia5
2025 SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment
abstract
While Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle with industrial anomaly detection and reasoning, particularly in delivering interpretable explanations and generalizing to unseen categories. This limitation stems from the inherently domain-specific nature of anomaly detection, which hinders the applicability of existing VLMs in industrial scenarios that require precise, structured, and context-aware analysis. To address these challenges, we propose SAGE, a VLM-based framework that enhances anomaly reasoning through Self-Guided Fact Enhancement (SFE) and Entropy-aware Direct Preference Optimization (E-DPO). SFE integrates domain-specific knowledge into visual reasoning via fact extraction and fusion, while E-DPO aligns model outputs with expert preferences using entropy-aware optimization. Additionally, we introduce AD-PL, a preference-optimized dataset tailored for industrial anomaly reasoning, consisting of 28,415 question-answering instances with expert-ranked responses. To evaluate anomaly reasoning models, we develop Multiscale Logical Evaluation (MLE), a quantitative framework analyzing model logic and consistency. SAGE demonstrates superior performance on industrial anomaly datasets under zero-shot and one-shot settings. The code, model, and dataset are available at https://github.com/amoreZgx1n/SAGE.
Guoxin Zang, Xue Li 0011, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song 0001, Lei Fan 0007
ACM Multimedia3
2025 GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
abstract
Articulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object representations, we propose GAOT, a three-phase framework that generates articulated objects from text prompts, leveraging diffusion models and hypergraph learning in a three-step process.
Lei Fan 0007, Donglin Di, Shaohui Liu
MMAsia3
2025 UniCP: A Unified Caching and Pruning Framework for Efficient Video Generation
abstract
Diffusion Transformers (DiT) excel in video generation but encounter significant computational challenges due to the quadratic complexity of attention. Notably, attention differences between adjacent diffusion steps follow a U-shaped pattern. Current methods leverage this property by caching attention blocks; however, they still struggle with sudden error spikes and large discrepancies. To address these issues, we propose UniCP—a unified caching and pruning framework for efficient video generation. UniCP optimizes both temporal and spatial dimensions through: Error-Aware Dynamic Cache Window (EDCW): Dynamically adjusts cache window sizes for different blocks at various timesteps to adapt to abrupt error changes. PCA-based Slicing (PCAS) and Dynamic Weight Shift (DWS): PCAS prunes redundant attention components, while DWS integrates caching and pruning by enabling dynamic switching between pruned and cached outputs. By adjusting cache windows and pruning redundant components, UniCP enhances computational efficiency and maintains video detail fidelity. Experimental results show that UniCP outperforms existing methods, delivering superior performance and efficiency.
Wenzhang Sun, Qirui Hou, Donglin Di, Yongjia Ma, Jianxun Cui
MMAsia3
2025 MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
abstract
Achieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained facial dynamics or maintain temporal identity coherence. To address these limitations, we propose MoCA, a novel Video Diffusion Model built on a Diffusion Transformer (DiT) backbone, incorporating a Mixture of Cross-Attention mechanism inspired by the Mixture-of-Experts paradigm. Our framework improves inter-frame identity consistency by embedding MoCA layers into each DiT block, where Hierarchical Temporal Pooling captures identity features over varying timescales, and Temporal-Aware Cross-Attention Experts dynamically model spatiotemporal relationships. We further incorporate a Latent Video Perceptual Loss to enhance identity coherence and fine-grained details across video frames. To train this model, we collect CelebIPVid, a dataset of 10,000 high-resolution videos from 1,000 diverse individuals, promoting cross-ethnic generalization. Extensive experiments on CelebIPVid show that MoCA outperforms existing T2V methods by over 5% across facial similarity.
Qi Xie 0009, Yongjia Ma, Donglin Di, Xuehao Gao, Xun Yang 0001
MMAsia3
2025 EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character Animation
abstract
Consistent pose‐driven character animation has achieved remarkable progress in single‐character scenarios. However, extending these advances to multi‐character settings is non‐trivial, especially when position swap is involved. Beyond mere scaling, the core challenge lies in enforcing correct Identity Correspondence (IC) between characters in reference and generated frames. To address this, we introduce EverybodyDance, a systematic solution targeting IC correctness in multi-character animation. EverybodyDance is built around the **Identity Matching Graph (IMG)**, which models characters in the generated and reference frames as two node sets in a weighted complete bipartite graph. Edge weights, computed via our proposed Mask–Query Attention (MQA), quantify the affinity between each pair of characters. Our key insight is to formalize IC correctness as a graph structural metric and to optimize it during training. We also propose a series of targeted strategies tailored for multi-character animation, including identity-embedded guidance, a multi-scale matching strategy, and pre-classified sampling, which work synergistically. Finally, to evaluate IC performance, we curate the **Identity Correspondence Evaluation** benchmark, dedicated to multi‐character IC correctness. Extensive experiments demonstrate that EverybodyDance substantially outperforms state‐of‐the‐art baselines in both IC and visual fidelity.
Haotian Ling, Zequn Chen, Qiuying Chen, Donglin Di, Yongjia Ma, Hao Li 0030, Zhulin Tao, Xun Yang 0001
NeurIPS4
2025 Hyper-3DG: Text-to-3D Gaussian Generation via Hypergraph
Donglin Di, Chaofan Luo, Zhou Xue, Wei Chen 0089, Xun Yang 0001, Yue Gao 0002
Int. J. Comput. Vis.1
2025 Multi-modal hypergraph contrastive learning for medical image segmentation
Weipeng Jing 0001, Junze Wang, Donglin Di, Yang Song 0001, Lei Fan 0007
Pattern Recognit.3
2025 Learning Frequency-Domain Fusion for Multimodal Remote Sensing Semantic Segmentation
Guangsheng Chen, Fangyu Sun, Weipeng Jing 0001, Weitao Zou, Donglin Di, Yang Song 0001, Lei Fan 0007
IEEE Trans. Geosci. Remote. Sens.5
2025 Hypergraph BiFormer for Semantic Segmentation of High-Resolution Remote Sensing Images
abstract
While transformers are powerful neural network architectures for feature learning, current Transformer-based approaches for semantic segmentation of high-resolution remote sensing images (HRRSIs) struggle with the extraction of local semantic features. To address this issue, we incorporate a hypergraph into the Transformer. Hypergraph-based methods are proficient at discovering high-order correlations within limited-scale data, extracting pertinent representations to enhance the Transformer’s learning capabilities. We also propose dual pooling and feature aggregation modules (FAMs), inspired by the adaptive pooling’s potent local modeling capabilities, to additionally extract fine-grained features from HRRSIs. In particular, we conceive a hypergraph BiFormer (HGBT) based on these three proposed modules along with a BiFormer backbone. HGBT has the potential to learn general latent features as well as generate high-order representations of HRRSIs by modeling correlations of multiscale features and local topology within an entirely nonlinear space, leading to the aggregation of features in a compact and localized manner, enhancing the model’s ability to capture detailed variations within small areas. We validate our approach through extensive experiments on ISPRS Vaihingen and Potsdam datasets, where HGBT attains mean intersection over union (mIoU) of 83.71% and 87.88%, respectively. Both quantitative and qualitative assessments underscore the dominance of HGBT. Our code will be accessible at:https://github.com/ZhangIceNight/HGBFormer.
Weipeng Jing 0001, Donglin Di, Chao Li 0066, Mahmoud Emam, Ajmal Mian
IEEE Trans. Geosci. Remote. Sens.3
2025 GrainBrain: Multiview Identification and Stratification of Defective Grain Kernels
abstract
Grain appearance inspection is crucial for evaluating grain quality and determining seed stratification. Typically, trained inspectors manually examine each grain kernel to identify and remove defective ones, which is time-consuming and error-prone. In this article, we present GrainBrain, a robotic vision-based system comprising a hardware prototype (A100) and a deep learning model (GrainAD). A100 is equipped with five cameras to capture high-quality, multiview images of each kernel. The identification of defective kernels is treated as an unsupervised anomaly detection task. GrainAD trains a classifier to distinguish between healthy and pseudoanomaly samples generated at both image and feature levels, and a supervised contrastive learning loss is employed to obtain compact feature representations of healthy kernels. In addition, we release a large-scale dataset containing over 100K annotated images of four types of cereal grains. Extensive experiments were conducted to verify the superiority of our system, achieving an average AUROC of 94.4/90.4% at the image/pixel level. Our system excelled in both efficiency and consistency, as demonstrated by experiments comparing human experts to the system.
Lei Fan 0007, Dongdong Fan, Yiwen Ding, Donglin Di, Maurice Pagnucco, Yang Song 0001
IEEE Trans. Ind. Informatics5
2025 TrAME: Trajectory-Anchored Multi-View Editing for Text-Guided 3D Gaussian Manipulation
abstract
Despite significant strides in the field of 3D scene editing, current methods encounter substantial challenge, particularly in preserving 3D consistency during the multi-view editing process. To tackle this challenge, we propose a progressive 3D editing strategy that ensures multi-view consistency via a Trajectory-Anchored Scheme (TAS) with a dual-branch editing mechanism. Specifically, TAS facilitates a tightly coupled iterative process between 2D view editing and 3D updating, preventing error accumulation yielded from the text-to-image process. Additionally, we explore the connection between optimization-based methods and reconstruction-based methods, offering a unified perspective for selecting superior design choices, supporting the rationale behind the designed TAS. We further present a tuning-free View-Consistent Attention Control (VCAC) module that leverages cross-view semantic and geometric reference from the source branch to yield aligned views from the target branch during the editing of 2D views. To validate the effectiveness of our method, we analyze 2D examples to demonstrate the improved consistency with the VCAC module. Extensive quantitative and qualitative results in text-guided 3D scene editing clearly indicate that our method can achieve superior editing quality compared with state-of-the-art 3D scene editing methods. Our project site is athttps://fkcptlst.github.io/TrAME/
Chaofan Luo, Donglin Di, Xun Yang 0001, Yongjia Ma, Zhou Xue, Wei Chen 0089, Xiaofei Gou, Yebin Liu
IEEE Trans. Multim.2
2025 Mode Hypergraph Neural Network
abstract
The hypergraph neural network (HGNN) is an emerging powerful tool for modeling and learning complex, high-order correlations among entities upon hypergraph structures. While existing HGNN-based approaches excel in modeling high-order correlations among data using hyperedges, they often have difficulties in distinguishing diverse semantics (e.g., bioactivities between drug and target in biological networks) of different correlations, making it challenging to learn accurate final representations. The underlying reason is that the specific semantic information of each hyperedge cannot be captured and distinguished during the modeling and learning process. To address this, we propose a mode HGNN ( $\textsf {MHGNN}$ ) framework that extends the vanilla hypergraph structure by endowing hyperedges with mode information for encapsulating their semantics and then performs mode-aware high-order message passing upon mode hypergraph for achieving comprehensive node representations. Extensive evaluations on four real-world datasets under two representative tasks have demonstrated the outstanding performance of $\textsf {MHGNN}$ against the state of the arts.
Shuyi Ji, Yifan Feng 0001, Donglin Di, Shihui Ying, Yue Gao 0002
IEEE Trans. Neural Networks Learn. Syst.3
2025 SRF: SpectrumRecombineFormer for Hyperspectral Image Classification
abstract
Hyperspectral imaging is a valuable technique for accurately classifying materials because of the abundance of spectral information and high resolution it provides. However, the characteristics of Hyperspectral Imaging, such as high-dimensional features and information redundancy, pose significant challenges to data processing. Traditional dimensionality reduction methods often have information loss, high computational complexity, and easy to ignore the strong correlation between HSI bands when dealing with the HSI data. Although other methods can achieve satisfactory classification performance, they do not consider the dimensionality reduction of HSI, and they focus on the model performance, which limits further improvement in classification performance. This article proposes a transformer-based framework called “SpectrumRecombineFormer” (SRF), which is composed of two key modules, namely “Spatial–Spectral Recombination” (SSRC) and “Cross-Layer Fusion” (CF). The SSRC is capable of utilizing both adjacent and non-adjacent spectrums to generate the spatial-sequential perceptive representations, which alleviate the effect of the strong correlation between HSI bands. The CF can avoid the loss of information during the feed-forward procedure among layers. Extensive experiments on five existing datasets (widely adopted Indian Pines, Houston2013, Pavia University, Salinas, and KSC) demonstrate the capability of our proposed method to address the above-mentioned challenges. Both quantitative and qualitative experimental ablation studies, including visualization results, reveal that the proposed SRF method can successfully and efficiently classify HSIs and surpass the other state-of-the-art methods. For access to the source code, please visit https://github.com/kangpeilun/SRF-HSI-Classification-master .
Weipeng Jing 0001, Peilun Kang, Donglin Di, Juntao Gu, Mahmoud Emam, Linda F. Mohaisen, Xun Yang 0001, Chao Li 0066
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Boundary-Guided Learning for Gene Expression Prediction in Spatial Transcriptomics
abstract
Spatial transcriptomics (ST) has emerged as an advanced technology that provides spatial context to gene expression. Recently, deep learning-based methods have shown the capability to predict gene expression from WSI data using ST data. Existing approaches typically extract features from images and the neighboring regions using pretrained models, and then develop methods to fuse this information to generate the final output. However, these methods often fail to account for the cellular structure similarity, cellular density and the interactions within the microenvironment.In this paper, we propose a framework named BG-TRIPLEX, which leverages boundary information extracted from pathological images as guiding features to enhance gene expression prediction from WSIs. Specifically, our model consists of three branches: the spot, in-context and global branches. In the spot and in-context branches, boundary information, including edge and nuclei characteristics, is extracted using pretrained models. These boundary features guide the learning of cellular morphology and the characteristics of microenvironment through Multi-Head Cross-Attention. Finally, these features are integrated with global features to predict the final output.Extensive experiments were conducted on three public ST datasets. The results demonstrate that our BG-TRIPLEX consistently outperforms existing methods in terms of Pearson Correlation Coefficient (PCC). This method highlights the crucial role of boundary features in understanding the complex interactions between WSI and gene expression, offering a promising direction for future research. Codes are available at: https://github.com/WcloudC0416/BG-TRIPLEX
Mingcheng Qu, Yuncong Wu, Donglin Di, Anyang Su, Tonghua Su, Yang Song 0001, Lei Fan 0007
BIBM3
2024 Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video Understanding
abstract
Understanding of video creativity and content often varies among individuals, with differences in focal points and cognitive levels across different ages, experiences, and genders. There is currently a lack of research in this area, and most existing benchmarks suffer from several drawbacks: 1) a limited number of modalities and answers with restrictive length; 2) the content and scenarios within the videos are excessively monotonous, transmitting allegories and emotions that are overly simplistic. To bridge the gap to real-world applications, we introduce a large-scale Video Subjective Multi-modal Evaluation dataset, namely Video-SME. Specifically, we collected real changes in Electroencephalographic (EEG) and eye-tracking regions from different demographics while they viewed identical video content. Utilizing this multi-modal dataset, we developed tasks and protocols to analyze and evaluate the extent of cognitive understanding of video content among different users. Along with the dataset, we designed a Hypergraph Multi-modal Large Language Model (HMLLM) to explore the associations among different demographics, video elements, EEG and eye-tracking indicators. HMLLM could bridge semantic gaps across rich modalities and integrate information beyond different modalities to perform logical reasoning. Extensive experimental evaluations on Video-SME and other additional video-based generative performance benchmarks demonstrate the effectiveness of our method. The code and dataset are available at https://github.com/mininglamp-MLLM/HMLLM
Anyang Su, Donglin Di, Tianyu Fu 0001, Da An, Meng Ma 0001, Kun Yan 0008, Ping Wang 0003
ACM Multimedia4
2024 Divide-Aggregate Heterogeneous Hypergraph for large-scale user intention detection
Mingcheng Qu, Xianyang Song, Donglin Di, Tonghua Su
Knowl. Based Syst.3
2023 Self-Supervised Cross-Language Scene Text Editing
abstract
We propose and formulate the task of cross-language scene text editing, modifying the text content of a scene image into new text in another language, while preserving the scene text style and background texture. The key challenges of this task lie in the difficulty in distinguishing text and background, great distribution differences among languages, and the lack of fine-labeled real-world data. To tackle these problems, we propose a novel network named Cross-LAnguage Scene Text Editing (CLASTE), which is capable of separating the foreground text and background, as well as further decomposing the content and style of the foreground text. Our model can be trained in a self-supervised training manner on the unlabeled and multi-language data in real-world scenarios, where the source images serve as both input and ground truth. Experimental results on the Chinese-English cross-language dataset show that our proposed model can generate realistic text images, specifically, modifying English to Chinese and vice versa. Furthermore, our method is universal and can be extended to other languages such as Arabic, Korean, Japanese, Hindi, Bengali, and so on.
Fuxiang Yang, Tonghua Su, Donglin Di, Zhongjie Wang 0003
ACM Multimedia4
2023 A sampling method based on forecasting and combinatorial optimization for high performance A/B testing
Tiancheng Zhang 0001, Shengjia Cui, Hengyu Liu 0001, Zhibin Ren, Donglin Di, Po Zhang, Ge Yu 0001
Frontiers Comput. Sci.6
2023 Dual attentional transformer for video visual relation prediction
Mingcheng Qu, Ganlin Deng, Donglin Di, Jianxun Cui, Tonghua Su
Neurocomputing3
2023 StfMLP: Spatiotemporal Fusion Multilayer Perceptron for Remote-Sensing Images
abstract
Remote-sensing (RS) images with high spatial and temporal resolutions play a significant role in monitoring periodic landscape changes for earth observation science. To enrich RS images, spatiotemporal fusion (STF) is considered a promising approach. The key challenge in the current STF-based methods is the requirement for large-scale data. In this work, we propose a deep-learning-based method called spatiotemporal fusion multilayer perceptron (StfMLP) to tackle this challenge. First, our method focuses on the given data in the manner of transductive learning. Second, we propose a designed multilayer perceptron (MLP) model to capture the time dependency and consistency among the input images. Consequently, StfMLP is capable of simultaneously achieving more accurate fusion and requiring a small-scale of data. We conduct extensive experiments on two widely adopted public datasets, namely Coleambally irrigation area (CIA) and the lower Gwydir catchment (LGC). The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods effectively. Code, trained model, and cropped images are available online ( https://github.com/luhailaing-max/StfMLP-master ).
Guangsheng Chen, Hailiang Lu 0004, Donglin Di, Mahmoud Emam, Weipeng Jing 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Generating Hypergraph-Based High-Order Representations of Whole-Slide Histopathological Images for Survival Prediction
abstract
Patient survival prediction based on gigapixel whole-slide histopathological images (WSIs) has become increasingly prevalent in recent years. A key challenge of this task is achieving an informative survival-specific global representation from those WSIs with highly complicated data correlation. This article proposes a multi-hypergraph based learning framework, called "HGSurvNet," to tackle this challenge. HGSurvNet achieves an effective high-order global representation of WSIs via multilateral correlation modeling in multiple spaces and a general hypergraph convolution network. It has the ability to alleviate over-fitting issues caused by the lack of training data by using a new convolution structure called hypergraph max-mask convolution. Extensive validation experiments were conducted on three widely-used carcinoma datasets: Lung Squamous Cell Carcinoma (LUSC), Glioblastoma Multiforme (GBM), and National Lung Screening Trial (NLST). Quantitative analysis demonstrated that the proposed method consistently outperforms state-of-the-art methods, coupled with the Bayesian Concordance Readjust loss. We also demonstrate the individual effectiveness of each module of the proposed framework and its application potential for pathology diagnosis and reporting empowered by its interpretability potential.
Donglin Di, Changqing Zou, Yifan Feng 0001, Rongrong Ji, Qionghai Dai, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Building Dialogue Understanding Models for Low-resource Language Indonesian from Scratch
abstract
Using off-the-shelf resources from resource-rich languages to transfer knowledge to low-resource languages has received a lot of attention. The requirements of enabling the model to achieve the reliable performance, including the scale of required annotated data and the effective framework, are not well guided. To address the first question, we empirically investigate the cost-effectiveness of several methods for training intent classification and slot-filling models from scratch in Indonesia (ID) using English data. Confronting the second challenge, we propose a Bi-Confidence-Frequency Cross-Lingual transfer framework (BiCF), which consists of “BiCF Mixing”, “Latent Space Refinement” and “Joint Decoder”, respectively, to overcome the lack of low-resource language dialogue data. BiCF Mixing based on the word-level alignment strategy generates code-mixed data by utilizing the importance-frequency and translating-confidence. Moreover, Latent Space Refinement trains a new dialogue understanding model using code-mixed data and word embedding models. Joint Decoder based on Bidirectional LSTM (BiLSTM) and Conditional Random Field (CRF) is used to obtain experimental results of intent classification and slot-filling. We also release a large-scale fine-labeled Indonesia dialogue dataset (ID-WOZ 1 ) and ID-BERT for experiments. BiCF achieves 93.56% and 85.17% (F1 score) on intent classification and slot filling, respectively. Extensive experiments demonstrate that our framework performs reliably and cost-efficiently on different scales of manually annotated Indonesian data.
Donglin Di, Xianyang Song, Weinan Zhang 0003, Yue Zhang 0004, Fanglin Wang
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2022 GrainSpace: A Large-scale Dataset for Fine-grained and Domain-adaptive Recognition of Cereal Grains
abstract
Cereal grains are a vital part of human diets and are important commodities for people's livelihood and international trade. Grain Appearance Inspection (GAI) serves as one of the crucial steps for the determination of grain quality and grain stratification for proper circulation, storage and food processing, etc. GAI is routinely performed manually by qualified inspectors with the aid of some hand tools. Automated GAI has the benefit of greatly assisting inspectors with their jobs but has been limited due to the lack of datasets and clear definitions of the tasks. In this paper we formulate GAI as three ubiquitous computer vision tasks: fine-grained recognition, domain adaptation and out-of-distribution recognition. We present a large-scale and publicly available cereal grains dataset called GrainSpace. Specifically, we construct three types of device prototypes for data acquisition, and a total of 5.25 million images determined by professional inspectors. The grain samples including wheat, maize and rice are collected from five countries and more than 30 regions. We also develop a comprehensive benchmark based on semi-supervised learning and self-supervised learning techniques. To the best of our knowledge, GrainSpace is the first publicly released dataset for cereal grain inspection, https://github.com/hellodfan/GrainSpace.
Lei Fan 0007, Yiwen Ding, Dongdong Fan, Donglin Di, Maurice Pagnucco, Yang Song 0001
CVPR4
2022 Binary Neural Network for Multispectral Image Classification
abstract
Compared with traditional images, multispectral images (MSIs) contain more spectral bands and higher data dimensions. The existing MSI classification model has high computational complexity and consumes a lot of computing resources. In this letter, we propose a lightweight multispectral classification method named CABNN based on binary neural networks (BNNs) to effectively have a trade-off between model performance and computational cost. First, we modify and binarize the MobileNetV1 network and add almost computation-free shortcuts to enhance the expressive capability. Secondly, since the BNN is sensitive to the distribution of activation functions, we introduce RPReLU with learnable coefficients to automatically adjust activation distribution at almost no extra cost. Lastly, considering that MSIs have multiple channels, we utilize an efficient channel attention (ECA) module to assign different weights to each channel to concentrate on crucial features and suppress insignificant features. We conduct experiments on four public MSI datasets, including NaSC-TG2, EuroSAT, GID Fine land-cover classification, and UC Merced Land Use. Extensive experiments demonstrate that the proposed CABNN has higher efficiency and better comprehensive performance than the state-of-the-art methods across the board.
Weipeng Jing 0001, Xu Zhang 0016, Jian Wang 0061, Donglin Di, Guangsheng Chen, Houbing Song
IEEE Geosci. Remote. Sens. Lett.4
2022 Context-Aware Attentional Graph U-Net for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) registers hundreds of spectral bands, whose intraclass variability and interclass similarity are resourceful information to be mined. Intraclass variability reflects the nonuniform and redundancy of the spatial and semantic features extracted from HSI. Interclass similarity represents the inherent relationship between adjacent features and snapshots. Existing models extract the superficial correlation representation for HSI to tackle the classification task but fail to embed the interclass and intraclass correlations due to these models’ intrinsic bottlenecks. Confronting the challenges of capturing interrelation for complex data in practice, we propose a Context-Aware Attentional Graph U-Net (CAGU) to improve these two modes of representation, which is more flexible in feature enhancement. In this method, attentional Graph U-Net is capable of extracting the intraclass embeddings within a non-Euclidean space by combining similar distributing feature vertices. The gated recurrent unit (GRU) is another critical component of our model to capture the context-aware dynamic interclass embeddings. Extensive experiments demonstrate that our model can efficiently outperform state-of-the-art methods across-the-board on five wide-adopted public data sets, namely, Pavia University, Indian Pines, Salinas Scene-show, Houston 2013, and Houston 2018, on par with the same scale of model parameters.
Moule Lin, Weipeng Jing 0001, Donglin Di, Guangsheng Chen, Houbing Song
IEEE Geosci. Remote. Sens. Lett.3
2022 Multi-Scale U-Shape MLP for Hyperspectral Image Classification
abstract
Hyperspectral images (HSIs) have significant applications in various domains, since they register numerous semantic and spatial information in the spectral band with spatial variability of spectral signatures. Two critical challenges in identifying pixels of the HSI are, respectively, representing the correlated information among the local and global, as well as the abundant parameters of the model. To tackle this challenge, we propose a multi-scale U-shape multi-layer perceptron (MUMLP) a model consisting of the designed multi-scale channel (MSC) block and the U-shape multi-layer perceptron (UMLP) structure. MSC transforms the channel dimension and mixes spectral band feature to embed the deep-level representation adequately. UMLP is designed by the encoder–decoder structure with multi-layer perceptron layers, which is capable of compressing the large-scale parameters. Extensive experiments are conducted to demonstrate that our model can outperform state-of-the-art methods across the board on three wide-adopted public datasets, namely Pavia University (PaviaU), Houston 2013, and Houston 2018.
Moule Lin, Weipeng Jing 0001, Donglin Di, Guangsheng Chen, Houbing Song
IEEE Geosci. Remote. Sens. Lett.3
2022 Big-Hypergraph Factorization Neural Network for Survival Prediction From Whole Slide Image
abstract
Survival prediction for patients based on histopa- thological whole-slide images (WSIs) has attracted increasing attention in recent years. Due to the massive pixel data in a single WSI, fully exploiting cell-level structural information (e.g., stromal/tumor microenvironment) from the gigapixel WSI is challenging. Most of the current studies resolve the problem by sampling limited image patches to construct a graph-based model (e.g., hypergraph). However, the sampling scale is a critical bottleneck since it is a fundamental obstacle of broadening samples for transductive learning. To overcome the limitation of the sampling scale for constructing a big hypergraph model, we propose a factorization neural network that embeds the correlation among large-scale vertices and hyperedges into two low-dimensional latent semantic spaces separately, empowering the dense sampling. Thanks to the compressed low-dimensional correlation embedding, the hypergraph convolutional layers generate the high-order global representation for each WSI. To minimize the effect of the uncertainty data as well as to achieve the metric-driven learning, we also propose a multi-level ranking supervision to enable the network learning by a queue of patients on the global horizon. Extensive experiments are conducted on three public carcinoma datasets (i.e., LUSC, GBM, and NLST), and the quantitative results demonstrate the proposed method outperforms state-of-the-art methods across-the-board.
Donglin Di, Jun Zhang 0018, Fuqiang Lei, Qi Tian 0001, Yue Gao 0002
IEEE Trans. Image Process.1
2021 Hypergraph learning for identification of COVID-19 with CT imaging
Donglin Di, Feng Shi 0001, Fuhua Yan, Liming Xia, Zhanhao Mo, Zhongxiang Ding, Bin Song 0002, Shengrui Li, Ying Wei 0009, Ying Shao, Miaofei Han, Yaozong Gao, He Sui, Yue Gao 0002, Dinggang Shen
Medical Image Anal.1
2021 geoGAT: Graph Model Based on Attention Mechanism for Geographic Text Classification
abstract
In the area of geographic information processing, there are few researches on geographic text classification. However, the application of this task in Chinese is relatively rare. In our work, we intend to implement a method to extract text containing geographical entities from a large number of network texts. The geographic information in these texts is of great practical significance to transportation, urban and rural planning, disaster relief, and other fields. We use the method of graph convolutional neural network with attention mechanism to achieve this function. Graph attention networks (GAT) is an improvement of graph convolutional neural networks (GCN). Compared with GCN, the advantage of GAT is that the attention mechanism is proposed to weight the sum of the characteristics of adjacent vertices. In addition, We construct a Chinese dataset containing geographical classification from multiple datasets of Chinese text classification. The Macro-F Score of the geoGAT we used reached 95% on the new Chinese dataset.
Weipeng Jing 0001, Xianyang Song, Donglin Di, Houbing Song
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2020 Ranking-Based Survival Prediction on Histopathological Whole-Slide Images
Donglin Di, Shengrui Li, Jun Zhang 0018, Yue Gao 0002
MICCAI (5)1
2019 A Neural Network Approach to Verb Phrase Ellipsis Resolution
abstract
Verb Phrase Ellipsis (VPE) is a linguistic phenomenon, where some verb phrases as syntactic constituents are omitted and typically referred by an auxiliary verb. It is ubiquitous in both formal and informal text, such as news articles and dialogues. Previous work on VPE resolution mainly focused on manually constructing features extracted from auxiliary verbs, syntactic trees, etc. However, the optimization of feature representation, the effectiveness of continuous features and the automatic composition of features are not well addressed. In this paper, we explore the advantages of neural models on VPE resolution in both pipeline and end-to-end processes, comparing the differences between statistical and neural models. Two neural models, namely multi-layer perception and the Transformer, are employed for the subtasks of VPE detection and resolution. Experimental results show that the neural models outperform the state-of-the-art baselines in both subtasks and the end-to-end results.
Weinan Zhang 0003, Yue Zhang 0004, Yuanxing Liu 0001, Donglin Di, Ting Liu 0001
AAAI4
2019 Annotating Objects and Relations in User-Generated Videos
abstract
Understanding the objects and relations between them is indispensable to fine-grained video content analysis, which is widely studied in recent research works in multimedia and computer vision. However, existing works are limited to evaluating with either small datasets or indirect metrics, such as the performance over images. The underlying reason is that the construction of a large-scale video dataset with dense annotation is tricky and costly. In this paper, we address several main issues in annotating objects and relations in user-generated videos, and propose an annotation pipeline that can be executed at a modest cost. As a result, we present a new dataset, named VidOR, consisting of 10k videos (84 hours) together with dense annotations that localize 80 categories of objects and 50 categories of predicates in each video. We have made the training and validation set public and extendable for more tasks to facilitate future research on video object and relation recognition.
Xindi Shang, Donglin Di, Junbin Xiao, Xun Yang 0001, Tat-Seng Chua
ICMR2
2019 Relation Understanding in Videos: A Grand Challenge Overview
abstract
ACM Multimedia 2019 Video Relation Understanding Challenge is the first grand challenge aiming at pushing video content analysis at the relational and structural level. This year, the challenge asks the participants to explore and develop innovative algorithms to detect object entities and their relations based on a large-scale user-generated video dataset. The tasks will advance the foundation of future visual systems that are able to perform complex inferences. This paper presents an overview of the grand challenge, including background, detailed descriptions of the three proposed tasks, the corresponding datasets for training, validation and testing, and the evaluation process.
Xindi Shang, Junbin Xiao, Donglin Di, Tat-Seng Chua
ACM Multimedia3