EDBT 2026 Demo / reviewers in the wild / expert
Yao Zhang 0010
dblp:57/3892-10
· DBLP profile ↗
22ranked-venue papers
4as first author
19since 2021 · last 2026
0000-0002-8759-4811ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SlimNet: High-Quality and Efficient Object Removal by Eliciting Latent Capabilities of Diffusion ModelsabstractObject removal aims to eliminate undesired objects from images while plausibly restoring the underlying content with high visual fidelity. Existing diffusion-based methods often achieve strong results by relying on large-scale model finetuning, auxiliary control networks, or powerful diffusion backbones, which incur substantial computational overhead and limit practical efficiency. In this work, we propose SlimNet, a high-quality and efficient object removal framework that elicits the latent object removal capabilities of pretrained diffusion models. SlimNet injects lightweight adapters into a frozen diffusion backbone to selectively modulate intermediate representations, avoiding heavy architectural modifications. To further improve visual fidelity, we design a composite perceptual loss that enforces object removal accuracy, background preservation, and smooth boundary transitions, together with a semantic-aware data processing pipeline for automatic mask generation. Extensive experiments demonstrate that SlimNet achieves competitive removal quality compared to state-of-the-art methods, with favorable reductions in inference time and VRAM consumption, making it a practical solution for resource-constrained multimedia applications. Yao Zhang 0010, Zhongchao Shi, Jianping Fan 0007, Guihua Zeng |
ICMR | 3 |
| 2025 | DoGA: Enhancing Grounded Object Detection via Grouped Pre-Training with AttributesabstractRecent advances in vision-language pre-training have significantly enhanced the model capabilities on grounded object detection. However, these studies often pre-train with coarse-grained text prompts, such as plain category names and brief grounded phrases. This limitation curtails the model's capacity for fine-grained linguistic comprehension and leads to a significant decline in performance when faced with detailed descriptions or contextual information. To tackle these problems, we develop DoGA: Detect objects with Grouped Attributes, which employs commonly apparent attributes to bridge different granular semantics and uses specific attributes to identify the object discrepancy. Our DoGA incorporates three principle components: 1) Generation of attribute-based prompts, consisting of linguistic definitions enriched with common-sense visible attributes and hard negative notations deriving from the image-specific attribute features; 2) Paralleled entity fusion and optimization, designed to manage long attribute-based descriptions and negative concepts efficiently; and 3) Prompt-wise grouped training to accommodate model to perform many-to-many assignments, facilitating simultaneous training and inferring with multiple attribute-based synonyms. Extensive experiments demonstrate that training with synonymous attribute-based prompts allows DoGA to generalize multi-granular prompts and surpass previous state-of-the-art approaches, yielding 50.2 on the COCO and 38.0 on the LVIS benchmarks under the zero-short setting. We will make our code publicly available upon acceptance. Yang Liu 0250, Feng Hou, Yunjie Peng, Gangjian Zhang, Yao Zhang 0010, Peng Wang 0095, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
AAAI | 5 |
| 2025 | Hypertasking: From Information Web to Computing Utility
Zi-Shu Yu, Feng-Zhi Li, Yao Zhang 0010 |
J. Comput. Sci. Technol. | 4 |
| 2025 | Gradient-aware domain-invariant learning for domain generalization
Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
Multim. Syst. | 2 |
| 2025 | RECISTSurv: Hybrid Multi-Task Transformer for Hepatocellular Carcinoma Response and Survival EvaluationabstractTransarterial Chemoembolization (TACE) is a widely applied alternative treatment for patients with hepatocellular carcinoma who are not eligible for liver resection or transplantation. However, the clinical outcomes after TACE are highly heterogeneous. There remains an urgent need for effective and efficient strategies to accurately assess tumor response and predict long-term outcomes using longitudinal and multi-center datasets. To address this challenge, we here introduce RECISTSurv, a novel response-driven Transformer model that integrates multi-task learning with a response-driven co-attention mechanism to simultaneously perform liver and tumor segmentation, predict tumor response to TACE, and estimate overall survival based on longitudinal Computed Tomography (CT) imaging. The proposed Response-driven Co-attention layer models the interactions between pre-TACE and post-TACE features guided by the treatment response embedding. This design enables the model to capture complex relationships between imaging features, treatment response, and survival outcomes, thereby enhancing both prediction accuracy and interpretability. In a multi-center validation study, RECISTSurv-predicted prognosis has demonstrated superior precision than state-of-the-art methods with C-indexes ranging from 0.595 to 0.780. Furthermore, when integrated with multi-modal data, RECISTSurvhas emerged as an independent prognostic factor in all three validation cohorts, with hazard ratio (HR) ranging from 1.693 to 20.7 (P = 0.001-0.042). Our results highlight the potential of RECISTSurvas a powerful tool for personalized treatment planning and outcome prediction in hepatocellular carcinoma patients undergoing TACE. The experimental code is made publicly available at https://github.com/rushier/RECISTSurv. Rushi Jiao, Qiuping Liu, Yao Zhang 0010, Bangzheng Pu, Bingsen Xue, Kailan Yang, Xisheng Liu, Jinrong Qu, Cheng Jin 0005, Ya Zhang 0002, Yanfeng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | DomainVerse: A Benchmark Towards Real-World Distribution Shifts for Training-Free Adaptive Domain GeneralizationabstractTraditional cross-domain tasks, including unsupervised domain adaptation (UDA), domain generalization (DG) and test-time adaptation (TTA), rely heavily on the training model by source domain data whether for specific or arbitrary target domains. With the recent advance of vision-language models (VLMs), recognized as natural source models that can be transferred to various downstream tasks without any parameter training, we propose a novel cross-domain task directly combining the strengths of both UDA and DG, named Training-Free Adaptive Domain Generalization (TF-ADG). However, current cross-domain datasets have many limitations, such as unrealistic domains, unclear domain definitions, and the inability to fine-grained domain decomposition, which hinder the real-world application of current cross-domain models due to the lack of accurate and fair evaluation of fine-grained realistic domains. These insights motivate us to establish a novel realistic benchmark for TF-ADG. Benefiting from the introduced hierarchical definition of domain shifts, our proposed dataset DomainVerse addresses these issues by providing about 0.5 million images from 390 realistic, hierarchical, and balanced domains, allowing for decomposition across multiple domains within each image. With the help of the constructed DomainVerse and VLMs, we further propose two algorithms called Domain CLIP and Domain++ CLIP for training-free adaptive domain generalization. Extensive and comprehensive experiments demonstrate the significance of the dataset and the effectiveness of the proposed methods. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002, Yong Rui |
IEEE Trans. Multim. | 4 |
| 2024 | A Survey of Visual TransformersabstractTransformer, an attention-based encoder-decoder model, has already revolutionized the field of natural language processing (NLP). Inspired by such significant achievements, some pioneering works have recently been done on employing Transformer-liked architectures in the computer vision (CV) field, which have demonstrated their effectiveness on three fundamental CV tasks (classification, detection, and segmentation) as well as multiple sensory data stream (images, point clouds, and vision-language data). Because of their competitive modeling capabilities, the visual Transformers have achieved impressive performance improvements over multiple benchmarks as compared with modern convolution neural networks (CNNs). In this survey, we have reviewed over 100 of different visual Transformers comprehensively according to three fundamental CV tasks and different data stream types, where taxonomy is proposed to organize the representative methods according to their motivations, structures, and application scenarios. Because of their differences on training settings and dedicated vision tasks, we have also evaluated and compared all these existing visual Transformers under different configurations. Furthermore, we have revealed a series of essential but unexploited aspects that may empower such visual Transformers to stand out from numerous architectures, e.g., slack high-level semantic embeddings to bridge the gap between the visual Transformers and the sequential ones. Finally, two promising research directions are suggested for future investment. We will continue to update the latest articles and their released source codes at https://github.com/liuyang-ict/awesome-visual-transformers. Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | SAP-DETR: Bridging the Gap Between Salient Points and Queries-Based Transformer Detector for Fast Model ConvergencyabstractRecently, the dominant DETR-based approaches apply central-concept spatial prior to accelerating Transformer detector convergency. These methods gradually refine the reference points to the center of target objects and imbue object queries with the updated central reference information for spatially conditional attention. However, centralizing reference points may severely deteriorate queries' saliency and confuse detectors due to the indiscriminative spatial prior. To bridge the gap between the reference points of salient queries and Transformer detectors, we propose SAlient Point-based DETR (SAP-DETR) by treating object detection as a transformation from salient points to instance objects. Concretely, we explicitly initialize a query-specific reference point for each object query, gradually aggregate them into an instance object, and then predict the distance from each side of the bounding box to these points. By rapidly attending to query-specific reference regions and the conditional box edges, SAP-DETR can effectively bridge the gap between the salient point and the query-based Transformer detector with a significant convergency speed. Experimentally, SAP-DETR achieves 1.4× convergency speed with competitive performance and stably promotes the SoTA approaches by ∼1.0 AP. Based on ResNet-DC-101, SAP-DETR achieves 46.9 AP. The code will be released at https://github.com/liuyang-ict/SAP-DETR. Yang Liu 0250, Yao Zhang 0010, Yixin Wang 0003, Yang Zhang 0002, Jiang Tian, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
CVPR | 2 |
| 2023 | Learning How to Learn Domain-Invariant Parameters for Domain GeneralizationabstractDue to domain shift, deep neural networks (DNNs) usually fail to generalize well on unknown test data in practice. Domain generalization (DG) aims to overcome this issue by capturing domain-invariant representations from source domains. Motivated by the insight that only partial parameters of DNNs are optimized to extract domain-invariant representations, we expect a general model that is capable of well perceiving and emphatically updating such domain-invariant parameters. In this paper, we propose two modules of Domain Decoupling and Combination (DDC) and Domain-invariance-guided Backpropagation (DIGB), which can encourage such general model to focus on the parameters that have a unified optimization direction between pairs of contrastive samples. Our extensive experiments on two benchmarks have demonstrated that our proposed method has achieved state-of-the-art performance with strong generalization capability. Feng Hou, Yao Zhang 0010, Yang Liu 0250, Yang Zhang 0002, Zhongchao Shi, Jianping Fan 0007, Zhiqiang He 0002 |
ICASSP | 2 |
| 2023 | Balanced masking strategy for multi-label image classification
Yao Zhang 0010, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui |
Neurocomputing | 2 |
| 2023 | Graph Attention Transformer Network for Multi-label Image ClassificationabstractMulti-label classification aims to recognize multiple objects or attributes from images. The key to solving this issue relies on effectively characterizing the inter-label correlations or dependencies, which bring the prevailing graph neural network. However, current methods often use the co-occurrence probability of labels based on the training set as the adjacency matrix to model this correlation, which is greatly limited by the dataset and affects the model’s generalization ability. This article proposes a Graph Attention Transformer Network, a general framework for multi-label image classification by mining rich and effective label correlation. First, we use the cosine similarity value of the pre-trained label word embedding as the initial correlation matrix, which can represent richer semantic information than the co-occurrence one. Subsequently, we propose the graph attention transformer layer to transfer this adjacency matrix to adapt to the current domain. Our extensive experiments have demonstrated that our proposed methods can achieve highly competitive performance on three datasets. Shikai Chen, Yao Zhang 0010, Zhongchao Shi, Xin Geng 0001, Jianping Fan 0007, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | ReMix: A General and Efficient Framework for Multiple Instance Learning Based Whole Slide Image Classification
Jiawei Yang 0002, Hanbo Chen, Yu Zhao 0009, Fan Yang 0081, Yao Zhang 0010, Lei He 0001, Jianhua Yao 0001 |
MICCAI (2) | 5 |
| 2022 | mmFormer: Multimodal Medical Transformer for Incomplete Multimodal Learning of Brain Tumor Segmentation
Yao Zhang 0010, Nanjun He, Jiawei Yang 0002, Yuexiang Li, Dong Wei 0004, Yawen Huang, Yang Zhang 0002, Zhiqiang He 0002, Yefeng Zheng 0001 |
MICCAI (5) | 1 |
| 2022 | Fast and Low-GPU-memory abdomen CT organ segmentation: The FLARE challengeabstractAutomatic segmentation of abdominal organs in CT scans plays an important role in clinical practice. However, most existing benchmarks and datasets only focus on segmentation accuracy, while the model efficiency and its accuracy on the testing cases from different medical centers have not been evaluated. To comprehensively benchmark abdominal organ segmentation methods, we organized the first Fast and Low GPU memory Abdominal oRgan sEgmentation (FLARE) challenge, where the segmentation methods were encouraged to achieve high accuracy on the testing cases from different medical centers, fast inference speed, and low GPU memory consumption, simultaneously. The winning method surpassed the existing state-of-the-art method, achieving a 19× faster inference speed and reducing the GPU memory consumption by 60% with comparable accuracy. We provide a summary of the top methods, make their code and Docker containers publicly available, and give practical suggestions on building accurate and efficient abdominal organ segmentation models. The FLARE challenge remains open for future submissions through a live platform for benchmarking further methodology developments at https://flare.grand-challenge.org/. Jun Ma 0016, Yao Zhang 0010, Song Gu, Xingle An, Zhihe Wang, Cheng Ge, Yinan Xu 0004, Shuiping Gou, Franz Thaler, Christian Payer, Darko Stern, Edward G. A. Henderson, Dónal M. McSweeney, Andrew Green 0001, Price Jackson, Lachlan McIntosh, Quoc-Cuong Nguyen, Abdul Qayyum 0002, Pierre-Henri Conze, Ziyan Huang, Deng-Ping Fan, Huan Xiong, Guoqiang Dong, Qiongjie Zhu, Xiaoping Yang 0001 |
Medical Image Anal. | 2 |
| 2022 | AbdomenCT-1K: Is Abdominal Organ Segmentation a Solved Problem?abstractWith the unprecedented developments in deep learning, automatic segmentation of main abdominal organs seems to be a solved problem as state-of-the-art (SOTA) methods have achieved comparable results with inter-rater variability on many benchmark datasets. However, most of the existing abdominal datasets only contain single-center, single-phase, single-vendor, or single-disease cases, and it is unclear whether the excellent performance can generalize on diverse datasets. This paper presents a large and diverse abdominal CT organ segmentation dataset, termed AbdomenCT-1K, with more than 1000 (1K) CT scans from 12 medical centers, including multi-phase, multi-vendor, and multi-disease cases. Furthermore, we conduct a large-scale study for liver, kidney, spleen, and pancreas segmentation and reveal the unsolved segmentation problems of the SOTA methods, such as the limited generalization ability on distinct medical centers, phases, and unseen diseases. To advance the unsolved problems, we further build four organ segmentation benchmarks for fully supervised, semi-supervised, weakly supervised, and continual learning, which are currently challenging and active research topics. Accordingly, we develop a simple and effective method for each benchmark, which can be used as out-of-the-box methods and strong baselines. We believe the AbdomenCT-1K dataset will promote future in-depth research towards clinical applicable abdominal organ segmentation methods. Jun Ma 0016, Yao Zhang 0010, Song Gu, Cheng Ge, Yichi Zhang 0007, Xingle An, Shucheng Cao, Qi Zhang 0059, Shangqing Liu, Xiaoping Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | TumorCP: A Simple but Effective Object-Level Data Augmentation for Tumor Segmentation
Jiawei Yang 0002, Yao Zhang 0010, Yuan Liang 0001, Yang Zhang 0002, Lei He 0001, Zhiqiang He 0002 |
MICCAI (1) | 2 |
| 2021 | Modality-Aware Mutual Learning for Multi-modal Medical Image Segmentation
Yao Zhang 0010, Jiawei Yang 0002, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 1 |
| 2021 | The state of the art in kidney and kidney tumor segmentation in contrast-enhanced CT imaging: Results of the KiTS19 challenge
Nicholas Heller, Fabian Isensee, Klaus H. Maier-Hein, Xiaoshuai Hou, Chunmei Xie, Fengyi Li, Yang Nan 0002, Guangrui Mu, Miofei Han, Guang Yao, Yaozong Gao, Yao Zhang 0010, Yixin Wang 0003, Feng Hou, Jiawei Yang 0002, Guangwei Xiong, Jiang Tian, Christopher J. Weight |
Medical Image Anal. | 13 |
| 2021 | Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation: The M&Ms ChallengeabstractThe emergence of deep learning has considerably advanced the state-of-the-art in cardiac magnetic resonance (CMR) segmentation. Many techniques have been proposed over the last few years, bringing the accuracy of automated segmentation close to human performance. However, these models have been all too often trained and validated using cardiac imaging samples from single clinical centres or homogeneous imaging protocols. This has prevented the development and validation of models that are generalizable across different clinical centres, imaging conditions or scanner vendors. To promote further research and scientific benchmarking in the field of generalizable deep learning for cardiac segmentation, this paper presents the results of the Multi-Centre, Multi-Vendor and Multi-Disease Cardiac Segmentation (M&Ms) Challenge, which was recently organized as part of the MICCAI 2020 Conference. A total of 14 teams submitted different solutions to the problem, combining various baseline models, data augmentation strategies, and domain adaptation techniques. The obtained results indicate the importance of intensity-driven data augmentation, as well as the need for further research to improve generalizability towards unseen scanner vendors or new imaging protocols. Furthermore, we present a new resource of 375 heterogeneous CMR datasets acquired by using four different scanner vendors in six hospitals and three different countries (Spain, Canada and Germany), which we provide as open-access for the community to enable future research in the field. Víctor M. Campello, Polyxeni Gkontra, Cristian Izquierdo, Carlos Martín-Isla, Alireza Sojoudi, Peter M. Full, Klaus H. Maier-Hein, Yao Zhang 0010, Zhiqiang He 0002, Jun Ma 0016, Mario Parreño, Alberto Albiol, Fanwei Kong, Shawn C. Shadden, Jorge Corral Acero, Vaanathi Sundaresan, Mina Saber, Mustafa A. Alattar, Hongwei Li 0004, Bjoern Menze, Firas Khader, Christoph Haarburger, Cian M. Scannell, Mitko Veta, Adam Carscadden, Kumaradevan Punithakumar, Xiao Liu 0037, Sotirios A. Tsaftaris, Xiaoqiong Huang, Xin Yang 0009, Lei Li 0020, Xiahai Zhuang, David Viladés, Martín Luís Descalzo, Andrea Guala 0002, Lucia La Mura, Matthias G. W. Friedrich, Ria Garg, Julie Lebel, Filipe Henriques, Mahir Karakas, Ersin Çavus, Steffen E. Petersen, Sergio Escalera, Santi Seguí, Jose Rodriguez-Palomares, Karim Lekadir |
IEEE Trans. Medical Imaging | 8 |
| 2020 | DARN: Deep Attentive Refinement Network for Liver Tumor Segmentation from 3D CT volumeabstractAutomatic liver tumor segmentation from 3D Computed Tomography (CT) is a necessary prerequisite in the interventions of hepatic abnormalities and surgery planning. However, accurate liver tumor segmentation remains challenging due to the large variability of tumor sizes and inhomogeneous texture. Recent advances based on Fully Convolutional Network (FCN) in liver tumor segmentation draw on success of learning discriminative multi-level features. In this paper, we propose a Deep Attentive Refinement Network (DARN) for improved liver tumor segmentation from CT volumes by fully exploiting both low and high level features embedded in different layers of FCN. Different from existing works, we exploit attention mechanism to leverage the relation of different levels of features encoded in different layers of FCN. Specifically, we introduce a Semantic Attention Refinement (SemRef) module to selectively emphasize global semantic information in low level features with the guidance of high level ones, and a Spatial Attention Refinement (SpaRef) module to adaptively enhance spatial details in high level features with the guidance of low level ones. We evaluate our network on the public MICCAI 2017 Liver Tumor Segmentation Challenge dataset (LiTS dataset) and it achieves state-of-the-art performance. The proposed refinement modules are an effective strategy to exploit multi-level features and has great potential to generalize to other medical image segmentation tasks. Yao Zhang 0010, Jiang Tian, Yang Zhang 0002, Zhongchao Shi, Zhiqiang He 0002 |
ICPR | 1 |
| 2020 | Double-Uncertainty Weighted Method for Semi-supervised Learning
Yixin Wang 0003, Yao Zhang 0010, Jiang Tian, Zhongchao Shi, Yang Zhang 0002, Zhiqiang He 0002 |
MICCAI (1) | 2 |
| 2018 | SequentialSegNet: Combination with Sequential Feature for Multi-Organ SegmentationabstractMulti-organ segmentation from computed tomography (CT) images is essential for computer aided diagnosis (CAD), and recent advances in fully convolutional networks (FCNs) for volumetric image segmentation have demonstrated the importance of leveraging spatial information. In this paper, we propose a novel framework called SequentialSegNet, which efficiently combines features within a single CT image (intra-slice) and among multiple adjacent images (inter-slice) for a multi-organ segmentation. Experimental results show that our approach can effectively improve the segmentation performance on both large-size and small-size abdominal organs including liver, spleen and gallbladder. Yao Zhang 0010, Yang Zhang 0002, Zhongchao Shi, Zhensheng Li, Zhiqiang He 0002 |
ICPR | 1 |