EDBT 2026 Demo / reviewers in the wild / expert
Enzhi Zhang
dblp:232/2921
· DBLP profile ↗
17ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0002-6421-0192ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improved Competitive Swarm Optimizer with Linear Population Reduction for Large-scale OptimizationabstractCompetitive swarm optimizer (CSO) is an efficient and effective swarm intelligence approach, especially for large-scale optimization. This paper presents an enhanced version of CSO termed improved CSO with linear population reduction (L-ICSO). The novel triple-individuals competitive mechanism is introduced to strengthen the optimization performance of L-ICSO, and the linear population reduction mechanism from L-SHADE is integrated into L-ICSO to highlight the explorative search in the initial phase of optimization and emphasize the exploitative behavior in the late phase. We conduct comprehensive numerical experiments in 100-dimensional CEC2017 benchmark functions. Ten state-of-the-art optimizers such as L-SHADE, jSO, L-SHADE-cnEpSin, and the original CSO are employed as competitor algorithms. The Mann–Whitney U and Holm multiple comparison tests are used to measure the statistical significance between L-ICSO and competitor algorithms. The experimental results and statistical analysis confirm the efficiency and effectiveness of our proposed L-ICSO in addressing large-scale optimization problems. The source code of L-ICSO can be found at https://github.com/RuiZhong961230/L-ICSO. Rui Zhong 0004, Jun Yu 0012, Xingbang Du, Enzhi Zhang, Abdelazim G. Hussien |
CEC | 4 |
| 2025 | NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule Generationabstract3D molecule generation is crucial for drug discovery and material design. While prior efforts focus on 3D diffusion models for their benefits in modeling continuous 3D conformers, they overlook the advantages of 1D SELFIES-based Language Models (LMs), which can generate 100\% valid molecules and leverage the billion-scale 1D molecule datasets. To combine these advantages for 3D molecule generation, we propose a foundation model -- NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule Generation. NExT-Mol uses an extensively pretrained molecule LM for 1D molecule generation, and subsequently predicts the generated molecule's 3D conformers with a 3D diffusion model. We enhance NExT-Mol's performance by scaling up the LM's model size, refining the diffusion neural architecture, and applying 1D to 3D transfer learning. Notably, our 1D molecule LM significantly outperforms baselines in distributional similarity while ensuring validity, and our 3D diffusion model achieves leading performances in conformer prediction. Given these improvements in 1D and 3D modeling, NExT-Mol achieves a 26\% relative improvement in 3D FCD for de novo 3D generation on GEOM-DRUGS, and a 13\% average relative gain for conditional 3D generation on QM9-2014. Our codes and pretrained checkpoints are available at https://github.com/acharkq/NExT-Mol. Zhiyuan Liu 0001, Yanchen Luo, Enzhi Zhang, Sihang Li 0002, Junfeng Fang, Yaorui Shi, Xiang Wang 0010, Kenji Kawaguchi, Tat-Seng Chua |
ICLR | 4 |
| 2025 | SHF: Symmetrical Hierarchical Forest with Pretrained Vision Transformer Encoder for High-Resolution Medical SegmentationabstractThis paper presents a novel approach to addressing the long-sequence problem in high-resolution medical images for Vision Transformers (ViTs). Using smaller patches as tokens can enhance ViT performance, but quadratically increases computation and memory requirements. Therefore, the common practice for applying ViTs to high-resolution images is either to: (a) employ complex sub-quadratic attention schemes or (b) use large to medium-sized patches and rely on additional mechanisms within the model to capture the spatial hierarchy of details. We propose Symmetrical Hierarchical Forest (SHF), a lightweight approach that adaptively patches the input image to increase token information density and encode hierarchical spatial structures into the input embedding. We then apply a reverse depatching scheme to the output embeddings of the transformer encoder, eliminating the need for convolution-based decoders. Unlike previous methods that modify attention mechanisms \wahib{or use a complex hierarchy of interacting models}, SHF can be retrofitted to any ViT model to allow it to learn the hierarchical structure of details in high-resolution images without requiring architectural changes. Experimental results demonstrate significant gains in computational efficiency and performance: on the PAIP WSI dataset, we achieved a 3$\sim$32$\times$ speedup or a 2.95\% to 7.03\% increase in accuracy (measured by Dice score) at a $64K^2$ resolution with the same computational budget, compared to state-of-the-art production models. On the 3D medical datasets BTCV and KiTS, training was 6$\times$ faster, with accuracy gains of 6.93\% and 5.9\%, respectively, compared to models without SHF. Enzhi Zhang, Peng Chen 0035, Rui Zhong 0004, Du Wu, Jun Igarashi, Isaac Lyngaas, Xiao Wang 0004, Masaharu Munetomo, Mohamed Wahib |
NeurIPS | 1 |
| 2025 | CWTLNet: Ultra-Short-Term Cryptocurrency Forecasting with Wavelet-Enhanced Deep ArchitectureabstractIn this study, we propose a novel deep neural model, CWTLNet, specifically designed to address the unique characteristics of cryptocurrency trading, including significant short-term volatility, which operates around the clock, and is prone to sudden price jumps and drops. The architecture consists of two branches. The CWT branch employs a continuous wavelet transform (CWT) to extract time-frequency feature maps from the input time series. These feature maps are then processed by a stack of CWT Blocks, which are composed of residual structures with two Conv1D layers, allowing the model to capture local, ultra-short-term patterns in the data. The linear branch first decomposes the time series into its trend and residual components. Each component is modeled separately using linear models and subsequently recombined. The linear branch is particularly effective at capturing the periodic patterns embedded in the time series. Finally, the outputs of the two branches are fused through a gated mechanism with a temperature coefficient, enabling adaptive weighting of the branch outputs. Experimental results demonstrate that CWTLNet outperforms commonly used models — including LSTM, Transformer, and DLinear — on minute-level trading data for seven different cryptocurrencies. For example, The performance of CWTLNet on the Bitcoin dataset, in terms of Mean Squared Error (MSE) and Mean Absolute Error (MAE), shows improvements of at least 2.7% and 1.7%, respectively, compared to the aforementioned three models. Xingbang Du, Enzhi Zhang, Rui Zhong 0004, Masaharu Munetomo |
SMC | 3 |
| 2025 | SAM2v-BTR: Accelerating SAM 2 Training for 3D Medical Image Segmentation Through Bootstrap and Memory AnnealingabstractVision Transformers (ViTs) have revolutionized image classification, but their application in dense prediction tasks such as segmentation, particularly for 3D medical imaging, faces scalability challenges. To address these issues, there has been growing interest in utilizing foundation models like the Segment Anything Model (SAM) for medical image analysis. The recent introduction of SAM 2 has generated optimism regarding enhancements for tasks such as 3D MRI segmentation. However, SAM 2’s memory bank mechanism, designed for object tracking and segmentation, introduces computational complexity and memory overhead, limiting training efficiency. In this paper, we address these limitations by proposing a bootstrap and memory annealing approach. The bootstrap mechanism replaces the memory bank with ground truth data, significantly enhancing training speed without compromising performance, achieving a 6x improvement on 3D medical image benchmarks such as KiTS and ACDC. To counteract potential overfitting caused by the bootstrap approach, we introduce memory annealing, which adaptively adjusts the selection of ground truth frames based on validation loss, resulting in a 5.46% improvement on the KiTS dataset. Our approach accelerates model training while maintaining or improving segmentation performance, offering a robust solution for efficient 3D medical image segmentation. Enzhi Zhang, Junichiro Iwasawa, Keita Oda, Yuta Tokuoka |
SMC | 1 |
| 2025 | Vision transformer-based meta loss landscape exploration with actor-critic method
Enzhi Zhang, Rui Zhong 0004, Xingbang Du, Mohamed Wahib, Masaharu Munetomo |
J. Supercomput. | 1 |
| 2025 | Competitive differential evolution with knowledge inheritance for single-objective human-powered aircraft design
Rui Zhong 0004, Enzhi Zhang, Masaharu Munetomo |
J. Supercomput. | 3 |
| 2024 | ProtT3: Protein-to-Text Generation for Text-based Protein UnderstandingabstractZhiyuan Liu, An Zhang, Hao Fei, Enzhi Zhang, Xiang Wang, Kenji Kawaguchi, Tat-Seng Chua. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhiyuan Liu 0001, An Zhang 0003, Hao Fei 0001, Enzhi Zhang, Xiang Wang 0010, Kenji Kawaguchi, Tat-Seng Chua |
ACL (1) | 4 |
| 2024 | Unnecessary Budget Reduction in Federated Active LearningabstractFederated active learning has been proposed as a potential solution to limited labeling budget in federated learning. The global model contains balanced knowledge across clients, while the local model provides a better fit to local knowledge. As the number of locally labeled samples increases, the local model can be iterated locally, resulting in a richer local knowledge. Therefore, using the predictive consistency of the global model and the iterative local models to filter pseudo-labels can improve the correctness of pseudo-labels. Based on this, we propose unnecessary budget reduction(UBR). UBR provides the oracle with inconsistent candidate samples to request true labels, while consistent candidate samples are given pseudo-labels. In order to maintain the model performance, candidate samples with pseudo labels are added to the model training. A progressive training strategy is employed to avoid pseudo-labeling noise affecting model training, in which pseudo-labeled samples are added to model training only when the pseudo-labeling correctness rate is sufficiently high. We conduct experiments combining UBR with several federated active learning methods, which demonstrates UBR effectively reduces the unnecessary query budget while maintaining model performance. Enzhi Zhang, Liu Yang 0010 |
ICTAI | 1 |
| 2024 | Validation Loss Landscape Exploration with Deep Q-LearningabstractOverfitting is a well-documented and studied issue in supervised learning. Human experts have been designing methods to reduce over-fitting by observing the validation knowledge, e.g., learning rate schedules, dropout, and adversarial training. We propose a validation-loss landscape exploration/exploitation method called VKI (Validation Knowledge Inheritance). We reformulate the traditional gradient optimization problem as a reinforcement learning task to explore the validation-loss landscape. In particular, by treating the gradient descent process as a Markov Decision Process (MDP), where the validation losses are treated as the costs, we train a Q-network as a controller to learn the future rewards and later use it to decrease the validation loss of workers. We conduct the experiments in two reduced gradient action spaces on the MNIST, CIFAR-10, and CIFAR-100 datasets with naive dense neural networks and ResNet-56. Our results show that VKI could rediscover hyperparameters schedule rules and improve the models’ training and generalization by only exploring/exploiting the validation loss landscape, e.g., increasing the learning rate accelerated training and penalty factor when overfitting. In comparison to other metaHPO (Hyper Parameter Optimization) methods, we empirically show that VKI can leverage its weights-loss regression of Q-Net to enable the transfer to the target dataset without heavy retraining but with light fine-tuning. Enzhi Zhang, Rui Zhong 0004, Masaharu Munetomo, Mohamed Wahib |
IJCNN | 1 |
| 2024 | On Softmax Direct Preference Optimization for RecommendationabstractRecommender systems aim to predict personalized rankings based on user preference data. With the rise of Language Models (LMs), LM-based recommenders have been widely explored due to their extensive world knowledge and powerful reasoning abilities. Most of the LM-based recommenders convert historical interactions into language prompts, pairing with a positive item as the target response and fine-tuning LM with a language modeling loss. However, the current objective fails to fully leverage preference data and is not optimized for personalized ranking tasks, which hinders the performance of LM-based recommenders. Inspired by the current advancement of Direct Preference Optimization (DPO) in human preference alignment and the success of softmax loss in recommendations, we propose Softmax-DPO (\textbf{S-DPO}) to instill ranking information into the LM to help LM-based recommenders distinguish preferred items from negatives, rather than solely focusing on positives. Specifically, we incorporate multiple negatives in user preference data and devise an alternative version of DPO loss tailored for LM-based recommenders, which is extended from the traditional full-ranking Plackett-Luce (PL) model to partial rankings and connected to softmax sampling strategies. Theoretically, we bridge S-DPO with the softmax loss over negative sampling and find that it has an inherent benefit of mining hard negatives, which assures its exceptional capabilities in recommendation tasks. Empirically, extensive experiments conducted on three real-world datasets demonstrate the superiority of S-DPO to effectively model user preference and further boost recommendation performance while providing better rewards for preferred items. Our codes are available at https://github.com/chenyuxin1999/S-DPO. Junfei Tan, An Zhang 0003, Zhengyi Yang 0007, Leheng Sheng, Enzhi Zhang, Xiang Wang 0010, Tat-Seng Chua |
NeurIPS | 6 |
| 2024 | Adaptive Patching for High-resolution Image Segmentation with TransformersabstractAttention-based models are proliferating in the space of image analytics, including segmentation. The standard method of feeding images to transformer encoders is to divide the images into patches and then feed the patches to the model as a linear sequence of tokens. For high-resolution images, e.g. microscopic pathology images, the quadratic compute and memory cost prohibits the use of an attention-based model, if we are to use smaller patch sizes that are favorable in segmentation. The solution is to either use custom complex multi-resolution models or approximate attention schemes. We take inspiration from Adapative Mesh Refinement (AMR) methods in HPC by adaptively patching the images, as a pre-processing step, based on the image details to reduce the number of patches being fed to the model, by orders of magnitude. This method has a negligible overhead, and works seamlessly with any attention-based model, i.e. it is a pre-processing step that can be adopted by any attention-based model without friction. We demonstrate superior segmentation quality over SoTA segmentation models for real-world pathology datasets while gaining a geomean speedup of $6.9 \times$ for resolutions up to $64 K^{2}$, on up to 2,048 GPUs. Enzhi Zhang, Isaac Lyngaas, Peng Chen 0035, Xiao Wang 0004, Jun Igarashi, Yuankai Huo, Masaharu Munetomo, Mohamed Wahib |
SC | 1 |
| 2024 | Meta generative image and text data augmentation optimization
Enzhi Zhang, Bochen Dong, Mohamed Wahib, Rui Zhong 0004, Masaharu Munetomo |
J. Supercomput. | 1 |
| 2024 | Evolutionary multi-mode slime mold optimization: a hyper-heuristic algorithm inspired by slime mold foraging behaviors
Rui Zhong 0004, Enzhi Zhang, Masaharu Munetomo |
J. Supercomput. | 2 |
| 2023 | Adjacent Intensity Matrix with Linkage Identification for Large-Scale Optimization in Noisy EnvironmentsabstractThis paper proposes a novel decomposition method: Adjacent Intensity Matrix with Linkage Identification (AIM-LI) for large-scale optimization problems (LSOPs) in noisy environments. The most advanced differential grouping (DG)-based methods such as DG2, Recursive DG, and dual DG are environmentally sensitive, and the presence of noise may misguide these methods to identify the originally separable decision variables as non-separable. Although the noise may affect the absolute intensity of interactions, we can detect the intensity between separable and non-separable decision variables with potential relative differences. This difference can help us classify the separability. Based on this hypothesis, we define the interaction intensity of pairwise decision variables and save them to AIM. A significant intensity (SI) is selected from AIM to transform the AIM into the dependency structure matrix (DSM). To verify the feasibility of AIM-LI, we first run the pre-experiments and mathematically analyze the possibility of interaction identification in noisy environments. Furthermore, we implement a decomposition experiment on CEC2013 with different strengths of noise. Theoretical analysis and experimental results show that our proposal has great potential to detect and classify the separability of problems in noisy environments. Rui Zhong 0004, Binan Tu, Enzhi Zhang, Masaharu Munetomo |
CEC | 3 |
| 2023 | Rethinking Tokenizer and Decoder in Masked Graph Modeling for MoleculesabstractMasked graph modeling excels in the self-supervised representation learning of molecular graphs. Scrutinizing previous studies, we can reveal a common scheme consisting of three key components: (1) graph tokenizer, which breaks a molecular graph into smaller fragments (\ie subgraphs) and converts them into tokens; (2) graph masking, which corrupts the graph with masks; (3) graph autoencoder, which first applies an encoder on the masked graph to generate the representations, and then employs a decoder on the representations to recover the tokens of the original graph. However, the previous MGM studies focus extensively on graph masking and encoder, while there is limited understanding of tokenizer and decoder. To bridge the gap, we first summarize popular molecule tokenizers at the granularity of node, edge, motif, and Graph Neural Networks (GNNs), and then examine their roles as the MGM's reconstruction targets. Further, we explore the potential of adopting an expressive decoder in MGM. Our results show that a subgraph-level tokenizer and a sufficiently expressive decoder with remask decoding have a \yuan{large impact on the encoder's representation learning}. Finally, we propose a novel MGM method SimSGT, featuring a Simple GNN-based Tokenizer (SGT) and an effective decoding strategy. We empirically validate that our method outperforms the existing molecule self-supervised learning methods. Our codes and checkpoints are available at https://github.com/syr-cn/SimSGT. Zhiyuan Liu 0001, Yaorui Shi, An Zhang 0003, Enzhi Zhang, Kenji Kawaguchi, Xiang Wang 0010, Tat-Seng Chua |
NeurIPS | 4 |
| 2023 | Training Knowledge Inheritance Through Deep Q-NetabstractWhen training neural networks, the weights of the model are updated at each optimization step, and the older weights are discarded. In this paper, we propose a method called, Training Knowledge Inheritance (TKI), to use the knowledge about the progression of weight and loss data in reducing overfitting and improving the generalization in the later stages of training. We reformulate the traditional gradient optimization problem as a reinforcement learning task. In particular, by treating the trainable weight space as an environment, the learning rate as action, and the validation accuracies as the rewards, we train a Q-network (controller) to learn the discounted future validation accuracy and guide the later training of another network (worker). We conduct the experiments on the MNIST, CIFAR-10, and CIFAR-100 datasets with naive dense neural networks and ResNet-56. Our results show that TKI could rediscover learning rate schedule rules similar to previous works, including increasing, decaying, and cyclical repeating. Enzhi Zhang, Ruqin Wang, Mohamed Wahib, Rui Zhong 0004, Masaharu Munetomo |
SMC | 1 |