EDBT 2026 Demo / reviewers in the wild / expert
Shalayiding Sirejiding
dblp:353/2588
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0003-1255-4994ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discretized Gaussian Representation for Tomographic Reconstruction
Shaokai Wu, Yapan Guo, Suizhi Huang, Shalayiding Sirejiding, Qichen He, Jing Tong, Yanbiao Ji, Yue Ding 0001, Hongtao Lu 0001 |
ICCV | 7 |
| 2025 | CLIP-MT: Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense PredictionabstractRecent advancements in visual multi-task learning (MTL) have sparked significant interest. However, existing dense prediction MTL methods predominantly rely on single-modality image data, limiting their performance due to the absence of complementary knowledge from other modalities. Additionally, different dense tasks exhibit heterogeneous preferences during information decoding, posing a critical challenge in effectively allocating multi-scale encoded features. To address these limitations, we propose CLIP-MT, a Multi-Modal Knowledge-Driven Adaptive Scale Feature Allocation for Multi-Task Dense Prediction. Specifically, to enrich task-shared image features with multi-modal knowledge, we introduce a novel CLIP-Guided Global Feature Enhancer (CGGF), which leverages aligned text-image information to augment object-level representations through a dual-path feature fusion architecture. Furthermore, to tackle the task-specific scale preference problem, we design an Adaptive Scale Selection Gate (ASSG), a learnable gating mechanism that dynamically selects high- or low-scale features based on task-specific demands. Finally, we integrate multi-modal and multi-scale information through a Task-Aware Feature Fusion Module (TAFF). Extensive experiments on the NYUDv2 and PASCAL-Context datasets demonstrate that CLIP-MT achieves state-of-the-art performance, outperforming existing methods across multiple dense prediction tasks. Shalayiding Sirejiding, Yue Ding 0001, Xinyi Hou, Shaokai Wu, Qichen He, Hongtao Lu 0001 |
ACM Multimedia | 1 |
| 2024 | Fedhca2: Towards Hetero-Client Federated Multi-Task LearningabstractFederated Learning (FL) enables joint training across distributed clients using their local data privately. Federated Multi-Task Learning (FMTL) builds on FL to handle multiple tasks, assuming model congruity that identical model architecture is deployed in each client. To relax this assumption and thus extend real-world applicability, we introduce a novel problem setting, Hetero-Client Fed-erated Multi-Task Learning (HC-FMTL), to accommodate diverse task setups. The main challenge of HC-FMTL is the model incongruity issue that invalidates conventional aggregation methods. It also escalates the difficulties in model aggregation to deal with data and task heterogeneity inherent in FMTL. To address these challenges, we pro-pose the$FedHCA^{2}$framework, which allows for federated training of personalized models by modeling relationships among heterogeneous clients. Drawing on our theoretical insights into the difference between multi-task and federated optimization, we propose the Hyper Conflict-Averse Aggregation scheme to mitigate conflicts during encoder updates. Additionally, inspired by task interaction in MTL, the Hyper Cross Attention Aggregation scheme uses layer-wise cross attention to enhance decoder interactions while alleviating model incongruity. Moreover, we employ learnable Hyper Aggregation Weights for each client to customize personalized parameter updates. Extensive experiments demon-strate the superior performance of$FedHCA^{2}$in various HC-FMTL scenarios compared to representative methods. Code is available at https://github.com/innovator-zero/FedHCA2. Suizhi Huang, Yuwen Yang, Shalayiding Sirejiding, Yue Ding 0001, Hongtao Lu 0001 |
CVPR | 4 |
| 2024 | YOLO-Med : Multi-Task Interaction Network for Biomedical ImagesabstractObject detection and semantic segmentation are pivotal components in biomedical image analysis. Current single-task networks exhibit promising outcomes in both detection and segmentation tasks. Multi-task networks have gained prominence due to their capability to simultaneously tackle segmentation and detection tasks, while also accelerating the segmentation inference. Nevertheless, recent multi-task networks confront distinct limitations such as the difficulty in striking a balance between accuracy and inference speed. Additionally, they often overlook the integration of cross-scale features, which is especially important for biomedical image analysis. In this study, we propose an efficient end-to-end multi-task network capable of concurrently performing object detection and semantic segmentation called YOLO-Med. Our model employs a backbone and a neck for multi-scale feature extraction, complemented by the inclusion of two task-specific decoders. A cross-scale task-interaction module is employed in order to facilitate information fusion between various tasks. Our model exhibits promising results in balancing accuracy and speed when evaluated on the Kvasir-seg dataset and a private biomedical image dataset. Suizhi Huang, Shalayiding Sirejiding, Yue Ding 0001, Leheng Liu, Hongtao Lu 0001 |
ICASSP | 2 |
| 2024 | Task Indicating Transformer for Task-Conditional Dense PredictionsabstractThe task-conditional model is a distinctive stream for efficient multi-task learning. Existing works encounter a critical limitation in learning task-agnostic and task-specific representations, primarily due to shortcomings in global context modeling arising from CNN-based architectures, as well as a deficiency in multi-scale feature interaction within the decoder. In this paper, we introduce a novel task-conditional framework called Task Indicating Transformer (TIT) to tackle this challenge. Our approach designs a Mix Task Adapter module within the transformer block, which incorporates a Task Indicating Matrix through matrix decomposition, thereby enhancing long-range dependency modeling and parameter-efficient feature adaptation by capturing intra- and inter-task features. Moreover, we propose a Task Gate Decoder module that harnesses a Task Indicating Vector and gating mechanism to facilitate adaptive multi-scale feature refinement guided by task embeddings. Experiments on two public multi-task dense prediction benchmarks, NYUD-v2 and PASCAL-Context, demonstrate that our approach surpasses state-of-the-art task-conditional methods. Shalayiding Sirejiding, Bayram Bayramli, Suizhi Huang, Yue Ding 0001, Hongtao Lu 0001 |
ICASSP | 2 |
| 2024 | BARTENDER: A simple baseline model for task-level heterogeneous federated learningabstractThis study presents the Task-level Heterogeneous Federated Learning (TH-FL), a novel paradigm that fuses the principles of Federated Learning (FL) and Multi-Task Learning (MTL). In the TH-FL scenario, each client can learn an indefinite number of tasks, which may vary in type and originate from distinct domains. We introduce a unique baseline model, BARTENDER, that integrates a Conditional Prompt (CP) module. This module encodes task-specific and domain-specific information, enabling the model to generate tailored outputs based on the encoding inputs. This innovative strategy not only minimizes the communication costs associated with FL but also enhances model generalization across a variety of task types. Through extensive experiments, we establish that the BARTENDER model surpasses traditional multi-decoder architecture models across diverse scenarios. We also explore the influence of the parameter decoupling strategy on model training and outline the assumptions necessary for achieving a $O\left( {1/\sqrt T } \right)$ convergence speed in the TH-FL scenario. Yuwen Yang, Suizhi Huang, Shalayiding Sirejiding, Chang Liu 0078, Muyang Yi, Zhaozhi Xie, Yue Ding 0001, Hongtao Lu 0001 |
ICME | 4 |
| 2024 | Federated Multi-Task Learning on Non-IID Data Silos: An Experimental StudyabstractThe innovative Federated Multi-Task Learning (FMTL) approach consolidates the benefits of Federated Learning (FL) and Multi-Task Learning (MTL), enabling collaborative model training on multi-task learning datasets. However, a comprehensive evaluation method, integrating the unique features of both FL and MTL, is currently absent in the field. This paper fills this void by introducing a novel framework, FMTL-Bench, for systematic evaluation of the FMTL paradigm. This benchmark covers various aspects at the data, model, and optimization algorithm levels, and comprises seven sets of comparative experiments, encapsulating a wide array of non-independent and identically distributed (Non-IID) data partitioning scenarios. We propose a systematic process for comparing baselines of diverse indicators and conduct a case study on communication expenditure, time, and energy consumption. Through our exhaustive experiments, we aim to provide valuable insights into the strengths and limitations of existing baseline methods, contributing to the ongoing discourse on optimal FMTL application in practical scenarios. The source code can be found at https://github.com/youngfish42/FMTL-Benchmark. Yuwen Yang, Suizhi Huang, Shalayiding Sirejiding, Hongtao Lu 0001, Yue Ding 0001 |
ICMR | 4 |
| 2024 | Task-Interaction-Free Multi-Task Learning with Efficient Hierarchical Feature RepresentationabstractTraditional multi-task learning often relies on explicit task interaction mechanisms to enhance multi-task performance. However, these approaches encounter challenges such as negative transfer when jointly learning multiple weakly correlated tasks. Additionally, these methods handle encoded features at a large scale, which escalates computational complexity to ensure dense prediction task performance. In this study, we introduce a Task-Interaction-Free Network (TIF) for multi-task learning, which diverges from explicitly designed task interaction mechanisms. Firstly, we present a Scale Attentive-Feature Fusion Module (SAFF) to enhance each scale in the shared encoder to have rich task-agnostic encoded features. Subsequently, our proposed task and scale-specific decoders efficiently decode the enhanced features shared across tasks without necessitating task-interaction modules. Concretely, we utilize a Self-Feature Distillation Module (SFD) to explore task-specific features at lower scales and the Low-To-High Scale Feature Diffusion Module (LTHD) to diffuse global pixel relationships from low-level to high-level scales. Experiments on publicly available multi-task learning datasets validate that our TIF attains state-of-the-art performance. Shalayiding Sirejiding, Bayram Bayramli, Yuwen Yang, Tamam Alsarhan, Hongtao Lu 0001, Yue Ding 0001 |
ACM Multimedia | 1 |
| 2024 | TFUT: Task fusion upward transformer model for multi-task learning on dense prediction
Zewei Xin, Shalayiding Sirejiding, Yue Ding 0001, Tamam Alsarhan, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2024 | Adaptive Task-Wise Message Passing for Multi-Task Learning: A Spatial Interaction PerspectiveabstractRecent advancements have facilitated the simultaneous processing of multiple dense prediction tasks, utilizing diverse correlations between these tasks. However, many of these advances predominantly focus on a singular or fixed task interaction, leading to negative transfer effects. In this paper, we introduce an end-to-end model called the Adaptive Task-Wise Message Passing Network (ATMPNet) for multi-task learning. Our proposed model focuses on excavating comprehensive spatial messages among tasks in an adaptive manner. To achieve this, ATMPNet incorporates the Adaptive Spatial Message Interaction (ASMI) module, which models various local spatial message interactions and global interactions among tasks. ASMI explores potential spatial relationships by generating a task-specific message pool for each target task. Furthermore, we propose an Adaptive Task Message Passing (ATMP) module, a novel method for aggregating messages. The ATMP module generates refined global-local messages from each message pool and adaptively transfers them to the corresponding target tasks through a well-designed message passing scheme. We conduct extensive experiments on the NYUD-v2 and PASCAL-Context datasets to evaluate the effectiveness of ATMPNet. The results demonstrate the state-of-the-art performance of our proposed model in handling multi-task learning scenarios. Code will be publicly available in here. Shalayiding Sirejiding, Bayram Bayramli, Suizhi Huang, Hongtao Lu 0001, Yue Ding 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Prompt Guided Transformer for Multi-Task Dense PredictionabstractTask-conditional architecture offers advantage in parameter efficiency but falls short in performance compared to state-of-the-art multi-decoder methods. How to trade off performance and model parameters is an important and difficult problem. In this paper, we introduce a simple and lightweight task-conditional model called Prompt Guided Transformer (PGT) to optimize this challenge. Our approach designs a Prompt-conditioned Transformer block, which incorporates task-specific prompts in the self-attention mechanism to achieve global dependency modeling and parameter-efficient feature adaptation across multiple tasks. This block is integrated into both the shared encoder and decoder, enhancing the capture of intra- and inter-task features. Moreover, we design a lightweight decoder to further reduce parameter usage, which accounts for only 2.7% of the total model parameters. Extensive experiments on two multi-task dense prediction benchmarks, PASCAL-Context and NYUD-v2, demonstrate that our approach achieves state-of-the-art results among task-conditional methods while using fewer parameters, and maintains a significant balance between performance and parameter size. Shalayiding Sirejiding, Yue Ding 0001, Hongtao Lu 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Scale-Aware Task Message Transferring for Multi-Task LearningabstractExploring cross-task interaction has been the mainstream in recent multi-task learning for dense predictions. However, existing works that focus on excavating cross-task contextual information are briefly based on hierarchical all-scale features, ignoring the complementary information from task relationships in multi-scale features. Meanwhile, recent research advances rarely pay attention to efficient information transformation between tasks. In this paper, we propose a novel end-to-end multi-task learning network termed as Scale-Aware Task Message Transferring Network (SATMTN) to explore the task relationships in multi-scale features to enrich the contextual cross-task information in the all-scale features. Specifically, we design Multi-Scale Message Passing Decoder (MSMP) to model interaction information between tasks in the multi-scale features, which is characterized by stacked bi-directional fully connected graph neural networks. Further, we devise another All-Scale Message Passing decoder (ASMP) to extract rich contextual information between tasks in the all-scale features, which is based on adaptive message transformation in the graph. Extensive experiments are implemented to show that our method surpasses the current state-of-the-art works on public multi-task benchmarks. Shalayiding Sirejiding, Hongtao Lu 0001, Yue Ding 0001 |
ICME | 1 |