Di Wu 0057

dblp:52/328-57 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0001-6589-7136ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Dynamics as Feedback: An Adaptive Entropy Flow Dynamics Framework for Long-tailed Human Action Recognition
abstract
Deep human action recognition models trained on real-world data are often challenged by long-tailed distributions, where performance on rare classes is severely degraded. Current solutions typically apply static or heuristic interventions that are disconnected from the model's evolving internal state. To overcome this limitation, we reconceptualize long-tailed human action recognition as a closed-loop, self-regulating system, inspired by ecological theory. We further introduce an Adaptive Ecological Entropy Dynamics (AEED) framework, which is built upon three synergistic components. First, AEED perceives the learning state through entropy flow, providing a robust and directional signal of learning progress. Second, this signal drives an adaptation mechanism, which dynamically adjusts class-specific loss weights to allocate more learning resources to underperforming classes. Finally, AEED facilitates intelligent knowledge transfer via Confidence-Guided Symbiosis (CS-Mix). Extensive experiments demonstrate that AEED achieves state-of-the-art performance on challenging skeleton-based action recognition benchmarks, including NTU-60-LT and Kinetics-400-LT.
Zhe Zhao 0008, Liheng Yu, Di Wu 0057, Pengkun Wang 0001
AAAI4
2026 Rethinking Crystal Symmetry Prediction: A Decoupled Perspective
abstract
Efficiently and accurately determining the symmetry is a crucial step in the structural analysis of crystalline materials. Existing methods usually mindlessly apply deep learning models while ignoring the underlying chemical rules. More importantly, experiments show that they face a serious sub-property confusion SPC problem. To address the above challenges, from a decoupled perspective, we introduce the XRDecoupler framework, a problem-solving arsenal specifically designed to tackle the SPC problem. Imitating the thinking process of chemists, we innovatively incorporate multidimensional crystal symmetry information as superclass guidance to ensure that the model's prediction process aligns with chemical intuition. We further design a hierarchical PXRD pattern learning model and a multi-objective optimization approach to achieve high-quality representation and balanced optimization. Comprehensive evaluations on three mainstream databases (e.g., CCDC, CoREMOF, and InorganicData) demonstrate that XRDecoupler excels in performance, interpretability, and generalization.
Liheng Yu, Zhe Zhao 0008, Xucong Wang, Di Wu 0057, Pengkun Wang 0001
AAAI4
2026 Memory-Efficient Intrinsic Gating Adaptation for Enhanced On-Device Epilepsy Diagnosis
abstract
Recently, advances in neuroscience and the rise of artificial intelligence have significantly enhanced the capabilities of epilepsy diagnosis. While EEG-based diagnosis offer a promising avenue for detecting and predicting seizure activity, practical implementation in real-world scenarios remains hindered by the heterogeneity of epilepsy and the variability of patient-specific biomarkers over time. Conventional deep learning models, trained on historical EEG, often fail to adapt to such biomarker variations, leading to degraded performance. Moreover, the computational and memory constraints of edge devices further exacerbate the challenge of on-device learning. To address these challenges, we introduce a novel framework, Memory-Efficient Intrinsic Gating Adaptation (MEIGA), designed to enhance real-world epilepsy diagnosis on resource-constrained edge devices. Our approach pre-trains a model using historical EEG data and employs lightweight adapter networks for efficient on-device tuning across new sessions, addressing session-to-session variability. By leveraging Direct Feedback Alignment (DFA), MEIGA reduces memory usage and computational overhead while maintaining high classification accuracy. Extensive experiments on the CHB-MIT epilepsy dataset demonstrate that MEIGA outperforms the pretrained-only Vision Transformer baseline, raising seizure prediction accuracy from 47.88% to 86.77% with only 3,908 tunable parameters (5.05% of the backbone). For seizure detection, MEIGA improves accuracy from 85.06% to 96.29% by adapting 2,008 parameters (17.40% of the base architecture). Further experiments on the AES dataset demonstrate that MEIGA consistently delivers strong performance across subjects and scales effectively to larger networks.
Shanjin Li, Di Wu 0057, Shiqi Zhao 0001, Jie Yang 0033, Mohamad Sawan
IEEE J. Biomed. Health Informatics2
2026 Neuro-BERT: Rethinking Masked Autoencoding for Self-Supervised Neurological Pretraining
abstract
Deep learning associated with neurological signals is poised to drive major advancements in diverse fields such as medical diagnostics, neurorehabilitation, and brain-computer interfaces. The challenge in harnessing the full potential of these signals lies in the dependency on extensive, high-quality annotated data, which is often scarce and expensive to acquire, requiring specialized infrastructure and domain expertise. To address the appetite for data in deep learning, we present Neuro-BERT, a self-supervised pre-training framework of neurological signals based on masked autoencoding in the Fourier domain. The intuition behind our approach is simple: frequency and phase distribution of neurological signals can reveal intricate neurological activities. We propose a novel pre-training task dubbed Fourier Inversion Prediction (FIP), which randomly masks out a portion of the input signal and then predicts the missing information using the Fourier inversion theorem. Pre-trained models can be potentially used for various downstream tasks such as sleep stage classification and gesture recognition. Unlike contrastive-based methods, which strongly rely on carefully hand-crafted augmentations and siamese structure, our approach works reasonably well with a simple transformer encoder with no augmentation requirements. By evaluating our method on several benchmark datasets, we show that Neuro-BERT improves downstream neurological-related tasks by a large margin.
Di Wu 0057, Siyuan Li 0002, Jie Yang 0033, Mohamad Sawan
IEEE J. Biomed. Health Informatics1
2025 Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings
abstract
Recent advancements in brain-computer interfaces (BCIs) and deep learning have made decoding lexical tones from intracranial recordings possible, providing the potential to restore the communication ability of speech-impaired tonal language speakers. However, data heterogeneity induced by both physiological and instrumental factors poses a significant challenge for unified invasive brain tone decoding. Particularly, the existing heterogeneous decoding paradigm (training subject-specific models with individual data) suffers from the intrinsic limitation that fails to learn generalized neural representations and leverages data across subjects. To this end, we introduce Homogeneity-Heterogeneity Disentangled Learning for Neural Representations (H2DiLR), a framework that disentangles and learns the homogeneity and heterogeneity from intracranial recordings of multiple subjects. To verify the effectiveness of H2DiLR, we collected stereoelectroencephalography (sEEG) from multiple participants reading Mandarin materials containing 407 syllables (covering nearly all Mandarin characters). Extensive experiments demonstrate that H2DiLR, as a unified decoding paradigm, outperforms the naive heterogeneous decoding paradigm by a large margin. We also empirically show that H2DiLR indeed captures homogeneity and heterogeneity during neural representation learning.
Di Wu 0057, Siyuan Li 0002, Jie Yang 0033, Mohamad Sawan
ICLR1
2025 IIB-DDI: Invariant Information Bottleneck Theory for Out-of-Distribution Drug-Drug Interaction Prediction
abstract
Rapid and accurate identification of drug-drug interactions (DDIs) among multiple medications is crucial for various medical treatments, and clinical therapies. Currently, the increasing significance of molecular substructure interactions in DDI prediction has become a consensus. However, due to the uneven distribution of substructures in molecules and most methods do not consider the differences in substructure distributions, models may rely on spurious substructure relationships, thereby weakening their out-of-distribution (OOD) generalization ability. To address this challenge, we propose an OOD-DDI framework called IIB-DDI. Specifically, information bottleneck theory is initially utilized to extract core subgraphs of drug pairs. Then, considering the diversity and unknown nature of environments, we introduce the vector quantization to design an environment codebook, where the potential environments in the dataset are clustered into a specified number of categories. Subsequently, we position the extracted core subgraphs under various latent environmental factors to attain invariant core substructures (rationales). Additionally, the learned environmental distribution also could be acted as noise injection for optimizing mutual information, achieving a smoother and more stable training curve, thereby leading to lower loss. Extensive experiments conducted on real-world DDI datasets demonstrate the superiority of our model over state-of-the-art baselines.
Xuqiang Li, Di Wu 0057, Wenjie Du 0003, Yang Wang 0015
IEEE Trans. Comput. Biol. Bioinform.4
2025 GenURL: A General Framework for Unsupervised Representation Learning
abstract
Unsupervised representation learning (URL) that learns compact embeddings of high-dimensional data without supervision has achieved remarkable progress recently. However, the development of URLs for different requirements is independent, which limits the generalization of the algorithms, especially prohibitive as the number of tasks grows. For example, dimension reduction (DR) methods, t-SNE and UMAP, optimize pairwise data relationships by preserving the global geometric structure, while self-supervised learning, SimCLR and BYOL, focuses on mining the local statistics of instances under specific augmentations. To address this dilemma, we summarize and propose a unified similarity-based URL framework, GenURL, which can adapt to various URL tasks smoothly. In this article, we regard URL tasks as different implicit constraints on the data geometric structure that help to seek optimal low-dimensional representations that boil down to data structural modeling (DSM) and low-dimensional transformation (LDT). Specifically, DSM provides a structure-based submodule to describe the global structures, and LDT learns compact low-dimensional embeddings with given pretext tasks. Moreover, an objective function, general Kullback-Leibler (GKL) divergence, is proposed to connect DSM and LDT naturally. Comprehensive experiments demonstrate that GenURL achieves consistent state-of-the-art performance in self-supervised visual learning, unsupervised knowledge distillation (KD), graph embeddings (GEs), and DR.
Siyuan Li 0002, Zicheng Liu 0006, Zelin Zang, Di Wu 0057, Zhiyuan Chen 0008, Stan Z. Li
IEEE Trans. Neural Networks Learn. Syst.4
2024 XRDMamba: Large-scale Crystal Material Space Group Identification with Selective State Space Model
abstract
In material science, the properties of crystalline materials largely depend on their structures, and space group is a key descriptor of crystal structure. With the rapid advancement of deep learning, the traditional artificial structure analysis method based on X-ray diffraction (XRD) has become cumbersome and is being gradually supplanted by neural networks. However, existing models are too simplistic and lack a comprehensive understanding of material structure. Our approach XRDMamba integrates chemical knowledge and presents a fresh crystal planes perspective on XRD data. We also introduce a knowledge-driven model for space group identification tasks. We have thoroughly analyzed our approach through numerous experiments, observing its SOTA performance and excellent generalization capabilities. The code is available in ~https://github.com/baigeiguai/XRDMamba.
Liheng Yu, Pengkun Wang 0001, Zhe Zhao 0008, Zhongchao Yi, Sun Nan, Di Wu 0057, Yang Wang 0015
CIKM6
2024 MogaNet: Multi-order Gated Aggregation Network
abstract
By contextualizing the kernel as global as possible, Modern ConvNets have shown great potential in computer vision tasks. However, recent progress on \textit{multi-order game-theoretic interaction} within deep neural networks (DNNs) reveals the representation bottleneck of modern ConvNets, where the expressive interactions have not been effectively encoded with the increased kernel size. To tackle this challenge, we propose a new family of modern ConvNets, dubbed MogaNet, for discriminative visual representation learning in pure ConvNet-based models with favorable complexity-performance trade-offs. MogaNet encapsulates conceptually simple yet effective convolutions and gated aggregation into a compact module, where discriminative features are efficiently gathered and contextualized adaptively. MogaNet exhibits great scalability, impressive efficiency of parameters, and competitive performance compared to state-of-the-art ViTs and ConvNets on ImageNet and various downstream vision benchmarks, including COCO object detection, ADE20K semantic segmentation, 2D\&3D human pose estimation, and video prediction. Notably, MogaNet hits 80.0\% and 87.8\% accuracy with 5.2M and 181M parameters on ImageNet-1K, outperforming ParC-Net and ConvNeXt-L, while saving 59\% FLOPs and 17M parameters, respectively. The source code is available at https://github.com/Westlake-AI/MogaNet.
Siyuan Li 0002, Zedong Wang, Zicheng Liu 0006, Cheng Tan 0012, Di Wu 0057, Zhiyuan Chen 0008, Jiangbin Zheng 0002, Stan Z. Li
ICLR6
2024 Exploring Effective Stimulus Encoding via Vision System Modeling for Visual Prostheses
abstract
Visual prostheses are potential devices to restore vision for blind people, which highly depends on the quality of stimulation patterns of the implanted electrode array. However, existing processing frameworks prioritize the generation of stimulation while disregarding the potential impact of restoration effects and fail to assess the quality of the generated stimulation properly. In this paper, we propose for the first time an end-to-end visual prosthesis framework (StimuSEE) that generates stimulation patterns with proper quality verification using V1 neuron spike patterns as supervision. StimuSEE consists of a retinal network to predict the stimulation pattern, a phosphene model, and a primary vision system network (PVS-net) to simulate the signal processing from the retina to the visual cortex and predict the firing rate of V1 neurons. Experimental results show that the predicted stimulation shares similar patterns to the original scenes, whose different stimulus amplitudes contribute to a similar firing rate with normal cells. Numerically, the predicted firing rate and the recorded response of normal neurons achieve a Pearson correlation coefficient of 0.78.
Chuanqing Wang, Di Wu 0057, Chaoming Fang, Jie Yang 0033, Mohamad Sawan
ICLR2
2024 VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling
abstract
Similar to natural language models, pre-trained genome language models are proposed to capture the underlying intricacies within genomes with unsupervised sequence modeling. They have become essential tools for researchers and practitioners in biology. However, the hand-crafted tokenization policies used in these models may not encode the most discriminative patterns from the limited vocabulary of genomic data. In this paper, we introduce VQDNA, a general-purpose framework that renovates genome tokenization from the perspective of genome vocabulary learning. By leveraging vector-quantized codebook as learnable vocabulary, VQDNA can adaptively tokenize genomes into pattern-aware embeddings in an end-to-end manner. To further push its limits, we propose Hierarchical Residual Quantization (HRQ), where varying scales of codebooks are designed in a hierarchy to enrich the genome vocabulary in a coarse-to-fine manner. Extensive experiments on 32 genome datasets demonstrate VQDNA’s superiority and favorable parameter efficiency compared to existing genome language models. Notably, empirical analysis of SARS-CoV-2 mutations reveals the fine-grained pattern awareness and biological significance of learned HRQ vocabulary, highlighting its untapped potential for broader applications in genomics.
Siyuan Li 0002, Zedong Wang, Zicheng Liu 0006, Di Wu 0057, Cheng Tan 0012, Jiangbin Zheng 0002, Yufei Huang 0002, Stan Z. Li
ICML4
2024 MMGNN: A Molecular Merged Graph Neural Network for Explainable Solvation Free Energy Prediction
Wenjie Du 0003, Di Wu 0057, Jun Xia 0001, Ziyuan Zhao, Junfeng Fang, Yang Wang 0015
IJCAI3
2023 Improving efficiency in rationale discovery for Out-of-Distribution molecular representations
abstract
Graph Neural Networks (GNNs), as the dominant approach in Molecular Representation Learning (MRL), have exhibited remarkable efficacy in diverse tasks, including molecular property prediction and drug discovery. Considering the dynamic and diverse nature of test molecules in real-world contexts, the previous independently and identically distributed (i.i.d.) assumption for training and test molecules is not align with the requirements of practical applications in chemistry. Rationalization has been proposed to enhance the Out-of-Distribution (OOD) generalization of GNN, but a critical concern remains unexplored: low efficiency in rationale discovery. We consider this bottleneck is due to large search space and overly flexible modeling. Here, we introduce a framework called Molecule Stratifier and Invariant Rational Improver (MSIRI) to overcome the challenge. Specifically, MSIRI adopt vector quantization to obtain invariant rationales and the search space is narrowed by substructure and two assumptions based on subgraph matching. Experimental results on ten datasets demonstrate that our approach both achieves significant improvements across various GNN backbones and outperforms other six Out-of-Distribution method, clearly demonstrating its effectiveness in addressing the OOD challenges in MRL.
Jiahe Li 0011, Wenjie Du 0003, Di Wu 0057, Yang Wang 0015
BIBM5
2023 Architecture-Agnostic Masked Image Modeling - From ViT back to CNN
abstract
Masked image modeling, an emerging self-supervised pre-training method, has shown impressive success across numerous downstream vision tasks with Vision transformers. Its underlying idea is simple: a portion of the input image is masked out and then reconstructed via a pre-text task. However, the working principle behind MIM is not well explained, and previous studies insist that MIM primarily works for the Transformer family but is incompatible with CNNs. In this work, we observe that MIM essentially teaches the model to learn better middle-order interactions among patches for more generalized feature extraction. We then propose an Architecture-Agnostic Masked Image Modeling framework (A$^2$MIM), which is compatible with both Transformers and CNNs in a unified way. Extensive experiments on popular benchmarks show that A$^2$MIM learns better representations without explicit design and endows the backbone model with the stronger capability to transfer to various downstream tasks.
Siyuan Li 0002, Di Wu 0057, Fang Wu 0002, Zelin Zang, Stan Z. Li
ICML2
2023 Fusing 2D and 3D molecular graphs as unambiguous molecular descriptors for conformational and chiral stereoisomers
abstract
The rapid progress of machine learning (ML) in predicting molecular properties enables high-precision predictions being routinely achieved. However, many ML models, such as conventional molecular graph, cannot differentiate stereoisomers of certain types, particularly conformational and chiral ones that share the same bonding connectivity but differ in spatial arrangement. Here, we designed a hybrid molecular graph network, Chemical Feature Fusion Network (CFFN), to address the issue by integrating planar and stereo information of molecules in an interweaved fashion. The three-dimensional (3D, i.e., stereo) modality guarantees precision and completeness by providing unabridged information, while the two-dimensional (2D, i.e., planar) modality brings in chemical intuitions as prior knowledge for guidance. The zipper-like arrangement of 2D and 3D information processing promotes cooperativity between them, and their synergy is the key to our model's success. Experiments on various molecules or conformational datasets including a special newly created chiral molecule dataset comprised of various configurations and conformations demonstrate the superior performance of CFFN. The advantage of CFFN is even more significant in datasets made of small samples. Ablation experiments confirm that fusing 2D and 3D molecular graphs as unambiguous molecular descriptors can not only effectively distinguish molecules and their conformations, but also achieve more accurate and robust prediction of quantum chemical properties.
Wenjie Du 0003, Xiaoting Yang, Di Wu 0057, Fenfen Ma, Baicheng Zhang, Chaochao Bao, Yaoyuan Huo, Yang Wang 0015
Briefings Bioinform.3
2022 Exploring Localization for Self-supervised Fine-grained Contrastive Learning
Di Wu 0057, Siyuan Li 0002, Zelin Zang, Stan Z. Li
BMVC1
2022 AutoMix: Unveiling the Power of Mixup for Stronger Classifiers
Zicheng Liu 0006, Siyuan Li 0002, Di Wu 0057, Zhiyuan Chen 0008, Lirong Wu, Stan Z. Li
ECCV (24)3
2022 DLME: Deep Local-Flatness Manifold Embedding
Zelin Zang, Siyuan Li 0002, Di Wu 0057, Kai Wang 0036, Baigui Sun, Hao Li 0030, Stan Z. Li
ECCV (21)3
2022 Towards Task-aware Signal Compression for Efficient Continuous Health Monitoring
abstract
High-precision multi-channel bio-signals are the basis of reliable and accurate wearable and implantable continuous health monitoring systems. However, the limitations of transmission bandwidth and computation resources of these systems pose heavy constraints on either the communication or direct processing of the large volume of physiological signals. Although signal compression can be adopted to compress the signals, most existing compression methods are computationally expensive and completely overlook the actual monitoring task purpose, which causes the discard of task-relevant information. Moreover, a complex reconstruction process is needed for further signal analysis at the cost of a heavy computational burden for downstream devices. We propose in this paper a novel flexible health monitoring framework where the signal is compressed with a low computation and hardware cost in-sensor compression matrix, trained in a task-aware fashion to preserve task-relevant information. The resulting compressed signals can be transmitted with significantly lower bandwidth, analyzed directly without a dedicated reconstruction process, or reconstructed with high fidelity. We demonstrate the effectiveness of our proposed framework by showcasing a seizure monitoring system. Prediction accuracy, sensitivity, false prediction rate, and signal reconstruction quality are reported under different compression ratios. Extensive experiments show that the proposed framework is accurate, with an average seizure prediction accuracy of 91.44%.
Di Wu 0057, Jie Yang 0033, Mohamad Sawan
ISCAS1
2022 Deep manifold embedding of attributed graphs
Zelin Zang, Siyuan Li 0002, Di Wu 0057, Jianzhu Guo, Yongjie Xu 0001, Stan Z. Li
Neurocomputing3