EDBT 2026 Demo / reviewers in the wild / expert
Pinzhuo Tian
dblp:226/5207
· DBLP profile ↗
22ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-5472-2469ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 9 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Double Buffer Vaccination: Bolstering Immunity Against Catastrophic Forgetting in Continual Medical Image Segmentation
Kai Chen 0026, Hewei Wang 0001, Pinzhuo Tian, Jing Huo, Taihang Zhen, Yang Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Pay Attention to the Foreground in Object-Centric LearningabstractThe slot attention-based method is widely used in unsupervised object-centric learning, aiming to decompose scenes into interpretable objects and associate them with slots. However, complex backgrounds in the real images can disrupt the model’s focus, leading it to excessively segment background stuff into different regions based on low-level information such as color or texture variations. As a result, the detailed segmentation of foreground objects, which requires shape or geometric information, is often neglected. To address this issue, we introduce a contrastive learning-based indicator designed to differentiate between foreground and background. Integrating this indicator into the existing slot attention-based method enables the model to focus more on segmenting foreground objects while minimizing background distractions. During the testing phase, we utilize a spectral clustering mechanism to refine the results based on the similarity between the slots. Experimental results show that incorporating our method with various state-of-the-art models significantly improves their performance on both simulated data and real-world datasets. Furthermore, multiple sets of ablation experiments confirm the effectiveness of each proposed component. The source code is available at https://github.com/sjyjs09/FG-BG_Indicator. Pinzhuo Tian, Shengjie Yang, Hang Yu 0006, Alex Chichung Kot |
CVPR | 1 |
| 2025 | A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language ModelsabstractWe propose a multi-agent approach (SeM-Agents) based on large language models for medical consultations. This framework incorporates various doctor roles and auxiliary roles, with agents communicating through natural language. Using a residual structure, the system conducts multi-round medical consultations based on the patient’s treatment background and symptoms. In the final summary and output stage of the consultation, it utilizes two experience databases—the Correct Consultation Experience Database and the Chain of Thought (CoT) Experience Database—which evolve with accumulated experience during consultations. This evolution drives the framework’s self-improvement, significantly enhancing the rationality and accuracy of the consultations. To ensure that the conclusions are safe, reliable, and aligned with human values, the final decisions undergo a safety review before being provided to the patient. This framework achieved accuracy rates of 89.2% and 83.1% on the MedQA and PubMedQA datasets, respectively. Kai Chen 0026, Jing Huo, Pinzhuo Tian, Yang Gao 0001 |
ICASSP | 4 |
| 2025 | Exploiting the Relationship within the Unlabelled Samples by Set Matching for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) is an open-world problem in semi-supervised learning, where the model leverages labelled data from known classes to discover both known and unknown classes in unlabelled data. Contrastive learning is key to this process, helping generate discriminative features that separate different categories. However, while labelled data benefits from direct supervision, unlabelled data relies solely on learning through instance discrimination, which focuses on consistency between different views of the same sample. This approach often neglects relationships between samples, limiting performance in GCD tasks. To improve this, we propose a set-matching perspective and formulate the problem as an optimal transport task to learn the coupling between batches of unlabelled data and a candidate pool. Our extensive experiments demonstrate that this approach better preserves the semantic structure of unlabelled data, leading to consistent improvements in baseline models and achieving state-of-the-art performance on various benchmarks. Qiubo Ma, Hang Yu 0006, Yuan Shan, Pinzhuo Tian |
ICASSP | 4 |
| 2025 | GravMAD: Grounded Spatial Value Maps Guided Action Diffusion for Generalized 3D ManipulationabstractRobots' ability to follow language instructions and execute diverse 3D manipulation tasks is vital in robot learning. Traditional imitation learning-based methods perform well on seen tasks but struggle with novel, unseen ones due to variability. Recent approaches leverage large foundation models to assist in understanding novel tasks, thereby mitigating this issue. However, these methods lack a task-specific learning process, which is essential for an accurate understanding of 3D environments, often leading to execution failures. In this paper, we introduce GravMAD, a sub-goal-driven, language-conditioned action diffusion framework that combines the strengths of imitation learning and foundation models. Our approach breaks tasks into sub-goals based on language instructions, allowing auxiliary guidance during both training and inference. During training, we introduce Sub-goal Keypose Discovery to identify key sub-goals from demonstrations. Inference differs from training, as there are no demonstrations available, so we use pre-trained foundation models to bridge the gap and identify sub-goals for the current task. In both phases, GravMaps are generated from sub-goals, providing GravMAD with more flexible 3D spatial guidance compared to fixed 3D positions. Empirical evaluations on RLBench show that GravMAD significantly outperforms state-of-the-art methods, with a 28.63\% improvement on novel tasks and a 13.36\% gain on tasks encountered during training. Evaluations on real-world robotic tasks further show that GravMAD can reason about real-world tasks, associate them with relevant visual information, and generalize to novel tasks. These results demonstrate GravMAD's strong multi-task learning and generalization in 3D manipulation. Video demonstrations are available at: https://gravmad.github.io. Yangtao Chen, Junhui Yin, Jing Huo, Pinzhuo Tian, Jieqi Shi, Yang Gao 0001 |
ICLR | 5 |
| 2025 | REST: A resolution preserving network for photorealistic style transfer via semantic distillation
Jing Huo, Zheng Gu 0001, Jiulin Zhang, Xiangde Liu, Shiyin Jin, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001 |
Comput. Vis. Image Underst. | 6 |
| 2025 | Dictionary Based Generative Adversarial Network for Multi-Collection Style TransferabstractMost collection-based style transfer methods require training a separate model for each individual collection of styles, making the extension to multiple collections of styles less flexible. Besides, the existing collection-based methods are also less flexible in extending to new style collections in a continual manner. To address these issues, we propose a novelMultI-Dictionary Generative Adversarial Network framework (MID-GAN)for multi-collection style transfer. Specifically, we design a multi-dictionary architecture within a GAN, with each dictionary consisting of a set of local style codes for a specific style collection. Benefiting from the local style codes used in the dictionary, a stylization module with aligned skip connections is further proposed, which can better preserve both the local details and the overall image structure. The dictionary design allows a flexible extension to new style collections by readily adding new dictionaries and we propose a continual training strategy that can both preserve the style transfer ability of old styles and achieve good transfer results for newly added styles. Extensive experiments are performed to show that the proposed method is better than existing collection-based style transfer methods. We also demonstrate the proposed method can generate diverse meaningful style transfer results of the same style collection. Jing Huo, Shiyin Jin, Jiashen Li, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | High Specificity Guided Cross-Domain Few-Shot SegmentationabstractCross-domain few-shot segmentation (CD-FSS) is a challenging vision task that involves segmenting novel classes from unseen domains using only a few annotated examples. Recent methods typically rely on powerful fundamental models for feature extraction, combined with complex, parameter-based decoders for segmentation. Due to the scarcity of annotated samples in the target domain, training a large number of parameters from scratch only on the source domain can lead to overfitting, which harms generalization in cross-domain set tings. These methods often assume that the fundamental feature extractor provides a sufficiently robust feature space, and freeze it during training to reduce the number of parameters. However, there is an inherent discrepancy between the tasks of training the fundamental model and performing segmentation, leading to a feature space mismatch. This misalignment results in features that, while useful for general tasks, may not fully satisfy the specific requirements of the segmentation task. To address this issue, we propose a metric-based approach called the High Specificity Guided Prototype Network (HSGNet). Our method is lightweight and focuses on fine-tuning the embedding network to align the feature space for the segmentation task. Specifically, we introduce a novel Feature Enrichment Module with extremely few parameters, which enhances the embedding network's ability to better align with the segmentation requirements. Instead of using a parameter-based decoder, our approach employs a non-parameter, similarity-based, high-specificity segmentation strategy. Additionally, we introduce a non-parameter test-time refinement mechanism to further improve prediction accuracy. Extensive experiments on cross-domain benchmarks demonstrate that our method achieves state-of-the-art performance with minimal additional parameters. Pinzhuo Tian, Hang Yu 0006, Shaorong Xie |
IEEE Trans. Multim. | 1 |
| 2025 | ReCL: A Plug-and-Play Module for Enhancing Generalized Category Discovery Using Transport-Based Method to Uncover the Relationship in SamplesabstractDeep learning systems excel in closed-set environments but face challenges in open-set settings due to mismatched label spaces between training and test data. Generalized category discovery (GCD) is one of such real-world open-set learning problems. In GCD, given a dataset, only a subset of samples is labeled. The model is expected to simultaneously classify samples from labeled and unlabeled classes. Contrastive learning plays a critical role in solving the GCD problem, used to learn discriminative features for samples. However, in contrast to labeled data, due to the absence of label information, unlabeled samples rely solely on unsupervised contrastive loss to learn discriminated features by keeping different views of the same data consistent. Unfortunately, this approach often overlooks the relationships within unlabeled samples. In this article, we propose a relationship-based contrastive learning (ReCL) module. In ReCL, we use a transport-based assignment method to find appropriate samples for each unlabeled data point. Then, a prototype-based fusion method is applied to merge these selected samples, creating a positive anchor in contrastive learning that helps pull the unlabeled samples closer to the corresponding positive anchor. Extensive experimental evaluation across different domains demonstrates that our method can be seamlessly integrated with various existing GCD models and further improve them to achieve the state-of-the-art performance across different benchmarks. Notably, we also analyze the sample selection process between our transport-based method and the cosine similarity-based method. The results show that our method provides samples that contain semantic similarity while offering greater diversity. Pinzhuo Tian, Qiubo Ma, Hang Yu 0006, Jie Lu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Learning from Orthogonal Space with Multimodal Large Models for Generalized Few-shot SegmentationabstractGeneralized Few-shot Segmentation (GFSS) aims to segment both base and novel classes in a query image, conditioning on richly annotated data of base classes and limited exemplars from novel classes. The learning of novel classes undoubtedly faces a disadvantage in this competition due to the highly unbalanced data, which skews the learned feature space toward the base classes. In this article, we present an innovative idea termed as “learning from orthogonal space” to avoid the conflict in the process of learning novel classes. Specifically, we first utilize textual modal information from labels to provide more distinguishable initial prototypes for different categories, ensuring that the prototypes for base and novel classes have distinct initial separations. Then, a simple but effective Feature Separating Module (FSM) is introduced to enhance the model’s ability to differentiate between base and novel classes through learning the novel features from orthogonal space. In addition, we propose a Trigger-Promoting Framework (TPF) during the testing stage to further boost performance. The prediction results from the FSM serve as a multimodal prompt to leverage information residing in large models, such as CLIP and SAM, to enhance performance. Comprehensive experiments on two benchmarks demonstrate that our method achieves superior performance on novel classes without sacrificing accuracy on base classes. Notably, our Feature Separating with Trigger-promoting Network (FS-TPNet) outperforms the current state-of-the-art method by \(12.8\%\) overall IoU on novel classes on PASCAL- \(5^{i}\) under the 1-shot scenario. Our codes will be available at https://github.com/returnZXJ/FS-TPNet . Hang Yu 0006, Shengjie Yang, Jing Huo, Pinzhuo Tian |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Modeling Rationality: Toward Better Performance Against Unknown Agents in Sequential GamesabstractOpponent modeling is necessary for autonomous agents to capture the intents of others during strategic interactions. Most previous works assume that they can access enough interaction history to build the model. However, it may not be realistic. To solve this problem, we present a novel rationality-consistent opponent modeling (ROM) method for games with imperfect information. In our approach, a game-theoretical concept of consistence about rationality is proposed to take advantage of the characteristic of imperfect information sequential games that rational behavior at disjoint information sets is correlated through anticipated opponent's behavior. With the correlation between different information sets, agents could infer the opponents' strategies at information sets correlated to observed behavior. To exploit the correlation, ROM attempts to conduct reasoning from the opponent's perspective and rationalize its past behavior. In this way, ROM acquires the ability to better adapt to different opponents and achieves a more accurate opponent model with insufficient observation history, which is verified by experiments in different settings. A heuristic adaptation approach is also applied in ROM, which updates the opponent model in an online manner and significantly reduces the computation cost. We evaluate ROM in both a grid world game and a poker game. Compared with other opponent modeling methods, ROM shows better performance and has more accurate predictions in both games against different types of opponents with limited action interactions. Experimental results also show that ROM's time cost is significantly reduced through heuristic adaptation. Zhenxing Ge, Shangdong Yang, Pinzhuo Tian, Yang Gao 0001 |
IEEE Trans. Cybern. | 3 |
| 2023 | An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form GamesabstractOne of the most popular methods for learning Nash equilibrium (NE) in large-scale imperfect information extensive-form games (IIEFGs) is the neural variants of counterfactual regret minimization (CFR). CFR is a special case of Follow-The-Regularized-Leader (FTRL). At each iteration, the neural variants of CFR update the agent's strategy via the estimated counterfactual regrets. Then, they use neural networks to approximate the new strategy, which incurs an approximation error. These approximation errors will accumulate since the counterfactual regrets at iteration t are estimated using the agent's past approximated strategies. Such accumulated approximation error causes poor performance. To address this accumulated approximation error, we propose a novel FTRL algorithm called FTRL-ORW, which does not utilize the agent's past strategies to pick the next iteration strategy. More importantly, FTRL-ORW can update its strategy via the trajectories sampled from the game, which is suitable to solve large-scale IIEFGs since sampling multiple actions for each information set is too expensive in such games. However, it remains unclear which algorithm to use to compute the next iteration strategy for FTRL-ORW when only such sampled trajectories are revealed at iteration t. To address this problem and scale FTRL-ORW to large-scale games, we provide a model-free method called Deep FTRL-ORW, which computes the next iteration strategy using model-free Maximum Entropy Deep Reinforcement Learning. Experimental results on two-player zero-sum IIEFGs show that Deep FTRL-ORW significantly outperforms existing model-free neural methods and OS-MCCFR. Linjian Meng, Zhenxing Ge, Pinzhuo Tian, Bo An 0001, Yang Gao 0001 |
AAAI | 3 |
| 2023 | Incremental event detection via an improved knowledge distillation based model
Changhua Xu, Hang Yu 0006, Pinzhuo Tian, Xiangfeng Luo |
Neurocomputing | 4 |
| 2023 | Can we improve meta-learning model in few-shot learning by aligning data distributions?
Pinzhuo Tian, Hang Yu 0006 |
Knowl. Based Syst. | 1 |
| 2023 | LibFewShot: A Comprehensive Library for Few-Shot LearningabstractFew-shot learning, especially few-shot image classification, has received increasing attention and witnessed significant advances in recent years. Some recent studies implicitly show that many generic techniques or "tricks", such as data augmentation, pre-training, knowledge distillation, and self-supervision, may greatly boost the performance of a few-shot learning method. Moreover, different works may employ different software platforms, backbone architectures and input image sizes, making fair comparisons difficult and practitioners struggle with reproducibility. To address these situations, we propose a comprehensive library for few-shot learning (LibFewShot) by re-implementing eighteen state-of-the-art few-shot learning methods in a unified framework with the same single codebase in PyTorch. Furthermore, based on LibFewShot, we provide comprehensive evaluations on multiple benchmarks with various backbone architectures to evaluate common pitfalls and effects of different training tricks. In addition, with respect to the recent doubts on the necessity of meta- or episodic-training mechanism, our evaluation results confirm that such a mechanism is still necessary especially when combined with pre-training. We hope our work can not only lower the barriers for beginners to enter the area of few-shot learning but also elucidate the effects of nontrivial tricks to facilitate intrinsic research on few-shot learning. Wenbin Li 0006, Xuesong Yang, Chuanqi Dong, Pinzhuo Tian, Tiexin Qin, Jing Huo, Yinghuan Shi, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | An Adversarial Meta-Training Framework for Cross-Domain Few-Shot LearningabstractMeta-learning provides a promising way for deep learning models to efficiently learn in few-shot learning. With this capacity, many deep learning systems can be applied in many real applications. However, many existing meta-learning based few-shot learning systems suffer from vulnerable generalization when new tasks are from unseen domains (a.k.a, cross-domain few-shot learning). In this work, we consider this problem from the perspective of designing a model-agnostic meta-training framework to improve the generalization of existing meta-learning methods in cross-domain few-shot learning. In this way, compared with focusing on elaborately designing modules for a specific meta-learning model, our method is endowed with the ability to be compatible with different meta-learning models in various few-shot problems. To achieve this goal, a novel adversarial meta-training framework is proposed. The proposed framework utilizes max-min episodic iteration. In the episode of maximization, our framework focuses on how to dynamically generate appropriate pseudo tasks which benefit learning cross-domain knowledge. In the episode of minimization, our method aims to solve how to help meta-learning model learn cross-task and robust meta-knowledge. To comprehensively evaluate our framework, experiments are conducted on two few-shot learning settings, three meta-learning models, and eight datasets. These results demonstrate that our method is applicable to various meta-learning models in different few-shot learning problems. The superiority of our method is verified compared with existing state-of-the-art methods. Pinzhuo Tian, Shaorong Xie |
IEEE Trans. Multim. | 1 |
| 2022 | DDMA: Discrepancy-Driven Multi-agent Reinforcement Learning
Yujing Hu, Pinzhuo Tian, Shaokang Dong, Yang Gao 0001 |
PRICAI (3) | 3 |
| 2022 | Improving meta-learning model via meta-contrastive loss
Pinzhuo Tian, Yang Gao 0001 |
Frontiers Comput. Sci. | 1 |
| 2022 | Consistent Meta-Regularization for Better Meta-Knowledge in Few-Shot LearningabstractRecently, meta-learning provides a powerful paradigm to deal with the few-shot learning problem. However, existing meta-learning approaches ignore the prior fact that good meta-knowledge should alleviate the data inconsistency between training and test data, caused by the extremely limited data, in each few-shot learning task. Moreover, legitimately utilizing the prior understanding of meta-knowledge can lead us to design an efficient method to improve the meta-learning model. Under this circumstance, we consider the data inconsistency from the distribution perspective, making it convenient to bring in the prior fact, and propose a new consistent meta-regularization (Con-MetaReg) to help the meta-learning model learn how to reduce the data-distribution discrepancy between the training and test data. In this way, the ability of meta-knowledge on keeping the training and test data consistent is enhanced, and the performance of the meta-learning model can be further improved. The extensive analyses and experiments demonstrate that our method can indeed improve the performances of different meta-learning models in few-shot regression, classification, and fine-grained classification. Pinzhuo Tian, Wenbin Li 0006, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Differentiable Meta-Learning Model for Few-Shot Semantic SegmentationabstractTo address the annotation scarcity issue in some cases of semantic segmentation, there have been a few attempts to develop the segmentation model in the few-shot learning paradigm. However, most existing methods only focus on the traditional 1-way segmentation setting (i.e., one image only contains a single object). This is far away from practical semantic segmentation tasks where the K-way setting (K > 1) is usually required by performing the accurate multi-object segmentation. To deal with this issue, we formulate the few-shot semantic segmentation task as a learning-based pixel classification problem, and propose a novel framework called MetaSegNet based on meta-learning. In MetaSegNet, an architecture of embedding module consisting of the global and local feature branches is developed to extract the appropriate meta-knowledge for the few-shot segmentation. Moreover, we incorporate a linear model into MetaSegNet as a base learner to directly predict the label of each pixel for the multi-object segmentation. Furthermore, our MetaSegNet can be trained by the episodic training mechanism in an end-to-end manner from scratch. Experiments on two popular semantic segmentation datasets, i.e., PASCAL VOC and COCO, reveal the effectiveness of the proposed MetaSegNet in the K-way few-shot semantic segmentation task. Pinzhuo Tian, Zhangkai Wu, Lei Qi 0001, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 1 |
| 2020 | Consistent MetaReg: Alleviating Intra-task Discrepancy for Better Meta-knowledgeabstractIn the few-shot learning scenario, the data-distribution discrepancy between training data and test data in a task usually exists due to the limited data. However, most existing meta-learning approaches seldom consider this intra-task discrepancy in the meta-training phase which might deteriorate the performance. To overcome this limitation, we develop a new consistent meta-regularization method to reduce the intra-task data-distribution discrepancy. Moreover, the proposed meta-regularization method could be readily inserted into existing optimization-based meta-learning models to learn better meta-knowledge. Particularly, we provide the theoretical analysis to prove that using the proposed meta-regularization, the conventional gradient-based meta-learning method can reach the lower regret bound. The extensive experiments also demonstrate the effectiveness of our method, which indeed improves the performances of the state-of-the-art gradient-based meta-learning models in the few-shot classification task. Pinzhuo Tian, Lei Qi 0001, Shaokang Dong, Yinghuan Shi, Yang Gao 0001 |
IJCAI | 1 |
| 2018 | A Novel Image-Specific Transfer Approach for Prostate Segmentation in MR ImagesabstractProstate segmentation in Magnetic Resonance (MR) Images is a significant yet challenging task for prostate cancer treatment. Most of the existing works attempted to design a global classifier for all MR images, which neglect the discrepancy of images across different patients. To this end, we propose a novel transfer approach for prostate segmentation in MR images. Firstly, an image-specific classifier is built for each training image. Secondly, a pair of dictionaries and a mapping matrix are jointly obtained by a novel Semi-Coupled Dictionary Transfer Learning (SCDTL). Finally, the classifiers on the source domain could be selectively transferred to the target domain (i.e. testing images) by the dictionaries and the mapping matrix. The evaluation demonstrates that our approach has a competitive performance compared with the state-of-the-art transfer learning methods. Moreover, the proposed transfer approach outperforms the conventional deep neural network based method. Pinzhuo Tian, Lei Qi 0001, Yinghuan Shi, Luping Zhou, Yang Gao 0001, Dinggang Sheri |
ICASSP | 1 |