VLDB 2026 Research / reviewers in the wild / expert
Zhengdong Luo
dblp:250/5839
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0005-8745-8528ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Progressive Self-Learning for Domain Adaptation on Symbolic Regression of Integer SequencesabstractSymbolic Regression of Integer Sequences (SRIS) aims to discover precise mathematical formulas from integer sequences. The neural machine translation-based method of SRIS trains the model using randomly generated data, and directly utilizes the trained model for inference on target sequences. However, the method often fails to effectively generalize to the target sequence, since the randomly generated data can not adequately cover the distributions of target data, i.e., there are distribution differences between them. In this work, we propose a progressive self-learning (PSL) method to explicitly capture sequence-formula distributions of the target domain. Specifically, a source domain dataset is generated by incorporating initial terms of the target domain to reduce the sequence distribution gap between the source domain and the target domain. Meanwhile, a self-learning loop strategy is adopted to improve the ability of the model to capture the sequence-formula distribution of the target domain. In this strategy, a neural machine translation model is used to learn the mappings from sequences to formulas in an end-to-end fashion. Then, this model is employed to explore candidate formulas of the target sequence using beam search. After verifying these candidate formula correctness, some of them are retained as training data for the next learning. Experimental results on OEIS datasets demonstrate that the proposed method surpasses current state-of-the-art methods in accuracy, and also discovers new formulas. Kaiming Sun, Zhengdong Luo, Lingfeng Wang 0002 |
AAAI | 3 |
| 2025 | FedSe: Group-Based Sequential Training Strategies for Mitigating Label Skew in Federated LearningabstractFederated Learning (FL) has emerged as a promising approach for distributed machine learning, enabling clients to collaboratively train models without sharing their data. However, existing FL methods continue to face challenges when dealing with non-IID data, particularly under conditions of extreme label skew. This divergence among client models can lead to a significant degradation in the accuracy of the global model. To address this critical issue, we introduce the FedSe approach, which incorporates two novel components: homogeneous grouping and sequential training. First, clients are grouped based on the distribution of data labels to ensure that each group contains a balanced representation of all labels. Second, task models are trained sequentially within these groups while training occurs in parallel across groups. The final step involves aggregating the models through averaging. Extensive experimental results on several datasets with label skew demonstrate the effectiveness of FedSe, and its results surpass current state-of-the-art methods. Ketu Qiao, Baoquan Wang, Zhengdong Luo |
ICASSP | 4 |
| 2025 | GraphVCM: Virtual Center Mixing with Distance-Aware Regulation for Class Imbalanced Node ClassificationabstractClass imbalance is a prevalent issue in real-world graph-structure data, such as social and citation networks, posing significant challenges for Graph Neural Networks (GNNs). Existing solutions often focus on balancing class distributions via oversampling techniques, which may lead to overfitting and blurred decision boundaries between majority and minority classes. To address this issue, we propose GraphVCM, a novel virtual center mixing with distance-aware regulation approach for class imbalanced node classification. GraphVCM first leverages spectral clustering to generate virtual center nodes, representing the features of majority class samples. These virtual centers are then mixed with minority class nodes to synthesize new nodes for enhancing minority class representation. Moreover, we introduce a distance-aware regulation module to optimize the inter-class decision boundaries by simultaneously maximizing inter-class distances and minimizing intra-class distances. Extensive experiments on multiple public datasets demonstrate that GraphVCM consistently outperforms state-of-the-art methods and effectively alleviates the negative impact of class imbalance. Yixiao Ren, Yunfei Han, Zhengdong Luo, Yupeng Ma |
ICASSP | 4 |
| 2025 | Calibrating feature representations for few-shot image recognition via vicinal mixup
Wuyuan Ye, Zhengdong Luo, Mengcheng Chen |
Multim. Syst. | 2 |
| 2024 | TabCGOK: Intra-Class Groups Retrieval and Inter-Class Ordinal Knowledge Augmented Network for Ordinal Tabular Data PredictionabstractOrdinal tabular data, with advantages of structured knowledge representation in tabular data and the characteristic of inter-class ranks, has drawn increasing attention. However, existing retrieval-based tabular deep learning methods designed primarily for classical tabular data pay less attention to ordinal tabular data. Ordinal knowledge of ordinal tabular data provides a more explicit objective for tabular ordinal classification by considering both classification and regression properties. Furthermore, these approaches overlook the significance of intra-class group features which can balance the retrieved probability of various sample size groups and capture shared knowledge among multiple samples within same group. In this work, we propose the Intra-Class Groups Retrieval and Inter-Class Ordinal Knowledge Augmented Network (TabCGOK) model for ordinal tabular data prediction, equipped with Intra-Class Groups Retrieval (CG) module and Inter-Class Ordinal Knowledge Augmented (OK) module. The CG module provides intra-class group features candidate set for subsequent retrieval operation. It divides each class into several groups, then extracts the representation of each group as intra-class group features. And the intra-class group features candidate set consists of all intra-class group features from each class. The OK module is designed to capture inter-class ordinal knowledge. It estimates the ordinal distances by calculating inter-class feature distances, which could correspond to the inter-class non-isometric nature of ordinal knowledge, and then aggregates the previous ordinal distances to clarify the containment relationship of ordinal knowledge. OK module utilizes the attention mechanism for fusing the captured ordinal knowledge to retrieved intra-class group features. Finally, TabCGOK integrates fused intra-class group features with sample level features for ordinal tabular data prediction. Extensive experiments on several ordinal tabular datasets demonstrate the effectiveness of our method. The source code is available at https://github.com/luozhengdong/TabCGOK. Zhengdong Luo, Abibulla Atawulla, Fengyi Yang, Yixiao Ren, Yunfei Han |
ECAI | 1 |
| 2020 | ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention NetworkabstractFood recognition has received more and more attention in the multimedia community for its various real-world applications, such as diet management and self-service restaurants. A large-scale ontology of food images is urgently needed for developing advanced large-scale food recognition algorithms, as well as for providing the benchmark dataset for such algorithms. To encourage further progress in food recognition, we introduce the dataset ISIA Food-500 with 500 categories from the list in the Wikipedia and 399,726 images, a more comprehensive food dataset that surpasses existing popular benchmark datasets by category coverage and data volume. Furthermore, we propose a stacked global-local attention network, which consists of two sub-networks for food recognition. One sub-network first utilizes hybrid spatial-channel attention to extract more discriminative features, and then aggregates these multi-scale discriminative features from multiple layers into global-level representation (e.g., texture and shape information about food). The other one generates attentional regions (e.g., ingredient relevant regions) from different regions via cascaded spatial transformers, and further aggregates these multi-scale regional features from different layers into local-level representation. These two types of features are finally fused as comprehensive representation for food recognition. Extensive experiments on ISIA Food-500 and other two popular benchmark datasets demonstrate the effectiveness of our proposed method, and thus can be considered as one strong baseline. The dataset, code and models can be found at http://123.57.42.89/FoodComputing-Dataset/ISIA-Food500.html. Weiqing Min, Linhu Liu, Zhengdong Luo, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2020 | Multi-Scale Multi-View Deep Feature Aggregation for Food RecognitionabstractRecently, food recognition has received more and more attention in image processing and computer vision for its great potential applications in human health. Most of the existing methods directly extracted deep visual features via convolutional neural networks (CNNs) for food recognition. Such methods ignore the characteristics of food images and are, thus, hard to achieve optimal recognition performance. In contrast to general object recognition, food images typically do not exhibit distinctive spatial arrangement and common semantic patterns. In this paper, we propose a multi-scale multi-view feature aggregation (MSMVFA) scheme for food recognition. MSMVFA can aggregate high-level semantic features, mid-level attribute features, and deep visual features into a unified representation. These three types of features describe the food image from different granularity. Therefore, the aggregated features can capture the semantics of food images with the greatest probability. For that solution, we utilize additional ingredient knowledge to obtain mid-level attribute representation via ingredient-supervised CNNs. High-level semantic features and deep visual features are extracted from class-supervised CNNs. Considering food images do not exhibit distinctive spatial layout in many cases, MSMVFA fuses multi-scale CNN activations for each type of features to make aggregated features more discriminative and invariable to geometrical deformation. Finally, the aggregated features are more robust, comprehensive, and discriminative via two-level fusion, namely multi-scale fusion for each type of features and multi-view aggregation for different types of features. In addition, MSMVFA is general and different deep networks can be easily applied into this scheme. Extensive experiments and evaluations demonstrate that our method achieves state-of-the-art recognition performance on three popular large-scale food benchmark datasets in Top-1 recognition accuracy. Furthermore, we expect this paper will further the agenda of food recognition in the community of image processing and computer vision. Shuqiang Jiang, Weiqing Min, Linhu Liu, Zhengdong Luo |
IEEE Trans. Image Process. | 4 |
| 2019 | Ingredient-Guided Cascaded Multi-Attention Network for Food RecognitionabstractRecently, food recognition is gaining more attention in the multimedia community due to its various applications, e.g., multimodal foodlog and personalized healthcare. Most of existing methods directly extract visual features of the whole image using popular deep networks for food recognition without considering its own characteristics. Compared with other types of object images, food images generally do not exhibit distinctive spatial arrangement and common semantic patterns, and thus are very hard to capture discriminative information. In this work, we achieve food recognition by developing an Ingredient-Guided Cascaded Multi-Attention Network (IG-CMAN), which is capable of sequentially localizing multiple informative image regions with multi-scale from category-level to ingredient-level guidance in a coarse-to-fine manner. At the first level, IG-CMAN generates the initial attentional region from the category-supervised network with Spatial Transformer (ST). Taking this localized attentional region as the reference, IG-CMAN combined ST with LSTM to sequentially discover diverse attentional regions with fine-grained scales from ingredient-guided sub-network in the following levels. Furthermore, we introduce a new dataset ISIA Food-200 with 200 food categories from the list in the Wikipedia, about 200,000 food images and 319 ingredients. We conducted extensive experiment on two popular food datasets and newly proposed ISIA Food-200, and verified the effectiveness of our method. Qualitative results along with visualization further show that IG-CMAN can introduce the explainability for localized regions, and is able to learn relevant regions for ingredients. Weiqing Min, Linhu Liu, Zhengdong Luo, Shuqiang Jiang |
ACM Multimedia | 3 |