EDBT 2026 Demo / reviewers in the wild / expert
Jintai Chen
dblp:249/3929
· DBLP profile ↗
48ranked-venue papers
11as first author
43since 2021 · last 2027
0000-0002-3199-2597ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 7 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Offline decision tree-based microscale evolutionary algorithm for multimodal multiobjective optimization
Fangqing Liu, Yinghan Hong, Jintai Chen, Han Huang 0002 |
Expert Syst. Appl. | 3 |
| 2026 | Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward ModelingabstractQiyuan Chen, Hongsen Huang, Jiahe Chen, Qian Shao, Jintai Chen, Hongxia Xu, Renjie Hua, Ren Chuan, Jian Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Qiyuan Chen 0003, Hongsen Huang, Qian Shao, Jintai Chen, Renjie Hua, Ren Chuan, Jian Wu 0001 |
ACL (1) | 5 |
| 2026 | Federated Learning Meets Test-Time Adaptation: Methods, Challenges, and Future Directions
Ge Su, Huaxiao Zhou, Lu Hao, Feng Zhu 0004, Jintai Chen, Jianwei Yin |
Int. J. Comput. Vis. | 6 |
| 2026 | Versatile and Risk-Sensitive Cardiac Diagnosis via Graph-Based ECG Signal RepresentationabstractDespite the rapid advancements of electrocardiogram (ECG) signal diagnosis and analysis methods through deep learning, two major hurdles still limit their clinical adoption: the lack of versatility in processing ECG signals with diverse configurations, and the inadequate detection of risk signals due to sample imbalances. Addressing these challenges, we introduceVersAtile andRisk-Sensitive cardiac diagnosis (VARS), an innovative approach that employs a graph-based representation to uniformly model heterogeneous ECG signals. VARS stands out by transforming ECG signals into versatile graph structures that capture critical diagnostic features, irrespective of signal diversity in the lead count, sampling frequency, and duration. This graph-centric formulation also enhances diagnostic sensitivity, enabling precise localization and identification of abnormal ECG patterns that often elude standard analysis methods. To facilitate representation transformation, our approach integrates denoising reconstruction with contrastive learning to preserve raw ECG information while highlighting pathognomonic patterns. We rigorously evaluate the efficacy of VARS on three distinct ECG datasets, encompassing a range of structural variations. The results demonstrate that VARS not only consistently surpasses existing state-of-the-art models across all these datasets but also exhibits substantial improvement in identifying risk signals. Additionally, VARS offers interpretability by pinpointing the exact waveforms that lead to specific model outputs, thereby assisting clinicians in making informed decisions. These findings suggest that our VARS will likely emerge as an invaluable tool for comprehensive cardiac health assessment. Yuyang Xu, Renjun Hu, Fanqi Shen, Hanyun Jiang, Jun Wang 0072, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001, Haochao Ying |
IEEE Trans. Big Data | 7 |
| 2025 | ProtCLIP: Function-Informed Protein Multi-Modal LearningabstractMulti-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model. Hanjing Zhou, Mingze Yin, Wei Wu 0045, Kun Fu 0002, Jintai Chen, Jian Wu 0001, Zheng Wang 0027 |
AAAI | 6 |
| 2025 | Scalable Autoregressive Monocular Depth EstimationabstractThis paper proposes a new autoregressive model as an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core designs. First, our depth autoregressive model (DAR) treats the depth map of different resolutions as a set of tokens, and conducts the low-to-high resolution autoregressive objective with a patch-wise causal mask. Second, our DAR recursively discretizes the entire depth range into more compact intervals, and attains the coarse-to-fine granularity autoregressive objective in an ordinal-regression manner. By coupling these two autoregressive objectives, our DAR establishes new state-of-the-art (SOTA) on KITTI and NYU Depth v2 by clear margins. Further, our scalable approach allows us to scale the model up to 2.0B and achieve the best RMSE of 1.799 on the KITTI dataset (5% improvement) compared to 1.896 by the current SOTA (Depth Anything). DAR further showcases zero-shot generalization ability on unseen datasets. These results suggest that DAR yields superior performance with an autoregressive prediction paradigm, providing a promising approach to equip modern autoregressive large models (e.g., GPT-4o) with depth estimation capabilities. Project page: https://depth-ar.github.io/. Dongqi Tang, Weiqiang Wang 0002, Danny Ziyi Chen, Jintai Chen, Jian Wu 0001 |
CVPR | 7 |
| 2025 | Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent RegulationabstractQiyuan Chen, Hongsen Huang, Qian Shao, Jiahe Chen, Jintai Chen, Hongxia Xu, Renjie Hua, Ren Chuan, Jian Wu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Qiyuan Chen 0003, Hongsen Huang, Qian Shao, Jintai Chen, Renjie Hua, Ren Chuan, Jian Wu 0001 |
EMNLP | 5 |
| 2025 | Proxy-Bridged Game Transformer for Interactive Extreme Motion Prediction
Yanwen Fang, Wenqi Jia 0001, Peng-Tao Jiang, Jintai Chen |
ICCV | 6 |
| 2025 | OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM
Shuo Tong, Dongqi Tang, Weiqiang Wang 0002, Danny Ziyi Chen, Jintai Chen, Jian Wu 0001 |
ICCV | 9 |
| 2025 | Small Models are LLM Knowledge Triggers for Medical Tabular PredictionabstractRecent development in large language models (LLMs) has demonstrated impressive domain proficiency on unstructured textual or multi-modal tasks. However, despite with intrinsic world knowledge, their application on structured tabular data prediction still lags behind, primarily due to the numerical insensitivity and modality discrepancy that brings a gap between LLM reasoning and statistical tabular learning. Unlike textual or vision data (e.g., electronic clinical notes or medical imaging data), tabular data is often presented in heterogeneous numerical values (e.g., CBC reports). This ubiquitous data format requires intensive expert annotation, and its numerical nature limits LLMs' capability to effectively transfer untapped domain expertise. In this paper, we propose SERSAL, a general self-prompting method by synergy learning with small models to enhance LLM tabular prediction in an unsupervised manner. Specifically, SERSAL utilizes the LLM's prior outcomes as original soft noisy annotations, which are dynamically leveraged to teach a better small student model. Reversely, the outcomes from the trained small model are used to teach the LLM to further refine its real capability. This process can be repeatedly applied to gradually distill refined knowledge for continuous progress. Comprehensive experiments on widely used medical domain tabular datasets show that, without access to gold labels, applying SERSAL to OpenAI GPT reasoning process attains substantial improvement compared to linguistic prompting methods, which serves as an orthogonal direction for tabular LLM, and increasing prompting bonus is observed as more powerful LLMs appear. Codes are available at https://github.com/jyansir/sersal. Jiahuan Yan, Jintai Chen, Chaowen Hu, Bo Zheng 0011, Yaojun Hu, Jimeng Sun 0001, Jian Wu 0001 |
ICLR | 2 |
| 2025 | Group-On: Boosting One-Shot Segmentation with Supportive QueryabstractOne-shot semantic segmentation aims to segment query images given only ONE annotated support image of the same class. This task is challenging because target objects in the support and query images can be largely different in appearance and pose (i.e., intra-class variation). Prior works suggested that incorporating more annotated support images in few-shot settings boosts performances but increases costs due to additional manual labeling. In this paper, we propose a novel and effective approach for ONE-shot semantic segmentation, called Group-On, which packs multiple query images in batches for the benefit of mutual knowledge support within the same category. Specifically, after coarse segmentation masks of the batch of queries are predicted, query-mask pairs act as pseudo support data to enhance mask predictions mutually. To effectively steer such process, we construct an innovative MoME module, where a flexible number of mask experts are guided by a scene-driven router and work together to make comprehensive decisions, fully promoting mutual benefits of queries. Comprehensive experiments on three standard benchmarks show that, in the ONE-shot setting, Group-On significantly outperforms previous works by considerable margins. With only one annotated support image, Group-On can be even competitive with the counterparts using 5 annotated images. Hanjing Zhou, Mingze Yin, Danny Ziyi Chen, Jian Wu 0001, Jintai Chen |
ICME | 5 |
| 2025 | Toward Human Deictic Gesture Target EstimationabstractHumans have a remarkable ability to use co-speech deictic gestures, such as pointing and showing, to enrich verbal communication and support social interaction. These gestures are so fundamental that infants begin to use them even before they acquire spoken language, which highlights their central role in human communication. Understanding the intended targets of another individual's deictic gestures enables inference of their intentions, comprehension of their current actions, and prediction of upcoming behaviors. Despite its significance, gesture target estimation remains an underexplored task within the computer vision community. In this paper, we introduce GestureTarget, a novel task designed specifically for comprehensive evaluation of social deictic gesture semantic target estimation. To address this task, we propose TransGesture, a set of Transformer-based gesture target prediction models. Given an input image and the spatial location of a person, our models predict the intended target of their gesture within the scene. Critically, our gaze-aware joint cross attention fusion model demonstrates how incorporating gaze-following cues significantly improves gesture target mask prediction IoU by 6% and gesture existence prediction accuracy by 10%. Our results underscore the complexity and importance of integrating gaze cues into deictic gesture intention understanding, advocating for increased research attention to this emerging area. All data, code will be made publicly available upon acceptance. Code of TransGesture is available at GitHub.com/IrohXu/TransGesture. Pranav Virupaksha, Sangmin Lee 0001, Bolin Lai, Wenqi Jia 0001, Jintai Chen, James M. Rehg |
NeurIPS | 6 |
| 2025 | A Progressively-Passing-Then-Disentangling Approach to Recipe RecommendationabstractThe increasing popularity of online food blogs and food ordering services has made personalized recipe recommendation a vital aspect of our emotional well-being. However, existing solutions, mainly based on graph neural networks, still face significant challenges, such as (a) focusing on exploiting the user-recipe interactions while neglecting other crucial pairwise and high-order relationships, and (b) failing to explicitly distinguish the distinct factors, e.g., hedonic and healthy, that influence recipe selection. To address these issues, we propose a progressively-passing-then-disentangling approach named P2D. Our approach utilizes a three-stage progressive message-passing mechanism for better representation learning. Specifically, we incorporate the extra pairwise relationships between recipes and nutrients, ingredients, and visual contents to create fine-grained and multimodal recipe representations. We next refine these representations via message passing between high-order recipe relationships to learn people's shared food preferences. Based on them, we could derive comprehensive user representations, which are subsequently transformed into disentangled forms that correspond to various decision factors through contrastive and mutual information regularization. Experimental results demonstrate both the superiority and the rationality of our method: (a) P2D outperforms the state-of-the-art recipe recommendation methods by a large margin under various metrics, (b) ablation studies confirm the positive impact of each of its components, and (c) our visualization analysis empirically supports the advantage of explicitly disentangling decision factors. Chunlai Dong, Haochao Ying, Renjun Hu, Yuyang Xu, Jintai Chen, Fuzhen Zhuang, Jian Wu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Multi-rater Prompting for Ambiguous Medical Image SegmentationabstractMulti-rater annotations commonly occur when medical images are independently annotated by multiple experts (raters). In this paper, we tackle two challenges arisen in multi-rater annotations for medical image segmentation (called ambiguous medical image segmentation): (1) How to train a deep learning model when a group of raters produces a set of diverse but plausible annotations, and (2) how to fine-tune the model efficiently when computation resources are not available for retraining the entire model on a different dataset domain. We propose a multi-rater prompt-based approach to address these two challenges altogether. Specifically, we introduce a series of rater-aware prompts that can be plugged into the U-Net model for uncertainty estimation to handle multi-annotation cases. During the prompt-based fine-tuning process, only 0.3% of learnable parameters are required to be updated comparing to training the entire model. Further, in order to integrate expert consensus and disagreement, we explore different multi-rater incorporation strategies and design a mix-training strategy for comprehensive insight learning. Extensive experiments verify the effectiveness of our new approach for ambiguous medical image segmentation on two public datasets while alleviating the heavy burden of model re-training. Code will be made available. Jintai Chen, Danny Ziyi Chen, Jian Wu 0001 |
BIBM | 3 |
| 2024 | Unraveling Babel: Exploring Multilingual Activation Patterns of LLMs and Their ApplicationsabstractRecently, large language models (LLMs) have achieved tremendous breakthroughs in the field of NLP, but still lack understanding of their internal neuron activities when processing different languages. We designed a method to convert dense LLMs into fine-grained MoE architectures, and then visually studied the multilingual activation patterns of LLMs through expert activation frequency heatmaps. Through comprehensive experiments on different model families, different model sizes, and different variants, we analyzed the similarities and differences in the internal neuron activation patterns of LLMs when processing different languages. Specifically, we investigated the distribution of high-frequency activated experts, multilingual shared experts, whether multilingual activation patterns are related to language families, and the impact of instruction tuning on activation patterns. We further explored leveraging the discovered differences in expert activation frequencies to guide sparse activation and pruning. Experimental results demonstrated that our method significantly outperformed random expert pruning and even exceeded the performance of unpruned models in some languages. Additionally, we found that configuring different pruning rates for different layers based on activation level differences could achieve better results. Our findings reveal the multilingual processing mechanisms within LLMs and utilize these insights to offer new perspectives for applications such as sparse activation and model pruning. Weize Liu, Yinlong Xu 0002, Jintai Chen, Xuming Hu, Jian Wu 0001 |
EMNLP | 4 |
| 2024 | Making Pre-trained Language Models Great on Tabular PredictionabstractThe transferability of deep neural networks (DNNs) has made significant progress in image and language processing. However, due to the heterogeneity among tables, such DNN bonus is still far from being well exploited on tabular data prediction (e.g., regression or classification tasks). Condensing knowledge from diverse domains, language models (LMs) possess the capability to comprehend feature names from various tables, potentially serving as versatile learners in transferring knowledge across distinct tables and diverse prediction tasks, but their discrete text representation space is inherently incompatible with numerical feature values in tables. In this paper, we present TP-BERTa, a specifically pre-trained LM for tabular data prediction. Concretely, a novel relative magnitude tokenization converts scalar numerical feature values to finely discrete, high-dimensional tokens, and an intra-feature attention approach integrates feature values with the corresponding feature names. Comprehensive experiments demonstrate that our pre-trained TP-BERTa leads the performance among tabular DNNs and is competitive with Gradient Boosted Decision Tree models in typical tabular data regime. Jiahuan Yan, Bo Zheng 0011, Yiheng Zhu 0002, Danny Ziyi Chen, Jimeng Sun 0001, Jian Wu 0001, Jintai Chen |
ICLR | 8 |
| 2024 | Personalized Heart Disease Detection via ECG Digital Twin Generation
Yaojun Hu, Jintai Chen, Lianting Hu, Dantong Li, Jiahuan Yan, Haochao Ying, Huiying Liang, Jian Wu 0001 |
IJCAI | 2 |
| 2024 | Can a Deep Learning Model be a Sure Bet for Tabular Prediction?abstractData organized in tabular format is ubiquitous in real-world applications, and users often craft tables with biased feature definitions and flexibly set prediction targets of their interests. Thus, a rapid development of a robust, effective, dataset-versatile, user-friendly tabular prediction approach is highly desired. While Gradient Boosting Decision Trees (GBDTs) and existing deep neural networks (DNNs) have been extensively utilized by professional users, they present several challenges for casual users, particularly: (i) the dilemma of model selection due to their different dataset preferences, and (ii) the need for heavy hyperparameter searching, failing which their performances are deemed inadequate. In this paper, we delve into this question: Can we develop a deep learning model that serves as a sure bet solution for a wide range of tabular prediction tasks, while also being user-friendly for casual users? We delve into three key drawbacks of deep tabular models, encompassing: (P1) lack of rotational variance property, (P2) large data demand, and (P3) over-smooth solution. We propose ExcelFormer, addressing these challenges through a semi-permeable attention module that effectively constrains the influence of less informative features to break the DNNs' rotational invariance property (for P1), data augmentation approaches tailored for tabular data (for P2), and attentive feedforward network to boost the model fitting capability (for P3). These designs collectively make ExcelFormer a sure bet solution for diverse tabular datasets. Extensive and stratified experiments conducted on real-world datasets demonstrate that our model outperforms previous approaches across diverse tabular data prediction tasks, and this framework can be friendly to casual users, offering ease of use without the heavy hyperparameter tuning. The codes are available at https://github.com/whatashot/excelformer. Jintai Chen, Jiahuan Yan, Qiyuan Chen 0003, Danny Ziyi Chen, Jian Wu 0001, Jimeng Sun 0001 |
KDD | 1 |
| 2024 | Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPsabstractTabular datasets play a crucial role in various applications.Thus, developing efficient, effective, and widely compatible prediction algorithms for tabular data is important.Currently, two prominent model types, Gradient Boosted Decision Trees (GBDTs) and Deep Neural Networks (DNNs), have demonstrated performance advantages on distinct tabular prediction tasks.However, selecting an effective model for a specific tabular dataset is challenging, often demanding time-consuming hyperparameter tuning.To address this model selection dilemma, this paper proposes a new framework that amalgamates the advantages of both GBDTs and DNNs, resulting in a DNN algorithm that is as efficient as GBDTs and is competitively effective regardless of dataset preferences for GBDTs or DNNs.Our idea is rooted in an observation that deep learning (DL) offers a larger parameter space that can represent a well-performing GBDT model, yet the current back-propagation optimizer struggles to efficiently discover such optimal functionality.On the other hand, during GBDT development, hard tree pruning, entropy-driven feature gate, and model ensemble have proved to be more adaptable to tabular data.By combining these key components, we present a Tree-hybrid simple MLP (T-MLP).In our framework, a tensorized, rapidly trained GBDT feature gate, a DNN architecture pruning approach, as well as a vanilla back-propagation optimizer collaboratively train a randomly initialized MLP model.Comprehensive experiments show that T-MLP is competitive with extensively tuned DNNs and GBDTs in their dominating tabular benchmarks (88 datasets) respectively, all achieved with compact model storage and significantly reduced training duration.The codes and full experiment results are available at https://github.com/jyansir/tmlp. Jiahuan Yan, Jintai Chen, Qianxing Wang, Danny Ziyi Chen, Jian Wu 0001 |
KDD | 2 |
| 2024 | 🐍 LKM-UNet: Large Kernel Vision Mamba UNet for Medical Image Segmentation
Jintai Chen, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (8) | 2 |
| 2024 | TeleOR: Real-Time Telemedicine System for Full-Scene Operating Room
Kaiyuan Hu, Qian Shao, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (6) | 4 |
| 2024 | Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language ModelsabstractWeize Liu, Guocong Li, Kai Zhang, Bang Du, Qiyuan Chen, Xuming Hu, Hongxia Xu, Jintai Chen, Jian Wu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Weize Liu, Guocong Li, Kai Zhang 0053, Bang Du, Qiyuan Chen 0003, Xuming Hu, Jintai Chen, Jian Wu 0001 |
NAACL-HLT | 8 |
| 2024 | A Corresponding Region Fusion Framework for Multi-Modal Cervical Lesion DetectionabstractCervical lesion detection (CLD) using colposcopic images of multi-modality (acetic and iodine) is critical to computer-aided diagnosis (CAD) systems for accurate, objective, and comprehensive cervical cancer screening. To robustly capture lesion features and conform with clinical diagnosis practice, we propose a novel corresponding region fusion network (CRFNet) for multi-modal CLD. CRFNet first extracts feature maps and generates proposals for each modality, then performs proposal shifting to obtain corresponding regions under large position shifts between modalities, and finally fuses those region features with a new corresponding channel attention to detect lesion regions on both modalities. To evaluate CRFNet, we build a large multi-modal colposcopic image dataset collected from our collaborative hospital. We show that our proposed CRFNet surpasses known single-modal and multi-modal CLD methods and achieves state-of-the-art performance, especially in terms of Average Precision. Tingting Chen 0002, Heping Hu, Chunhua Luo, Jintai Chen, Chunnv Yuan, Weiguo Lu, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | A Protein-Context Enhanced Master Slave Framework for Zero-Shot Drug Target Interaction PredictionabstractDrug Target Interaction (DTI) prediction plays a crucial role in in-silico drug discovery, especially for deep learning (DL) models. Along this line, existing methods usually first extract features from drugs and target proteins, and use drug-target pairs to train DL models. However, these DL-based methods essentially rely on similar structures and patterns defined by the homologous proteins from a large amount of data. When few drug-target interactions are known for a newly discovered protein and its homologous proteins, prediction performance can suffer notable reduction. In this paper, we propose a novel Protein-Context enhanced Master/Slave Framework (PCMS), for zero-shot DTI prediction. This framework facilitates the efficient discovery of ligands for newly discovered target proteins, addressing the challenge of predicting interactions without prior data. Specifically, the PCMS framework consists of two main components: a Master Learner and a Slave Learner. The Master Learner first learns the target protein context information, and then adaptively generates the corresponding parameters for the Slave Learner. The Slave Learner then perform zero-shot DTI prediction in different protein contexts. Extensive experiments verify the effectiveness of our PCMS compared to state-of-the-art methods in various metrics on two public datasets. Yuyang Xu, Jingbo Zhou 0003, Haochao Ying, Jintai Chen, Wei Chen 0001, Danny Ziyi Chen, Jian Wu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Polygonal Approximation Learning for Convex Object Segmentation in Biomedical Images With Bounding Box SupervisionabstractAs a common and critical medical image analysis task, deep learning based biomedical image segmentation is hindered by the dependence on costly fine-grained annotations. To alleviate this data dependence, in this article, a novel approach, called Polygonal Approximation Learning (PAL), is proposed for convex object instance segmentation with only bounding-box supervision. The key idea behind PAL is that the detection model for convex objects already contains the necessary information for segmenting them since their convex hulls, which can be generated approximately by the intersection of bounding boxes, are equivalent to the masks representing the objects. To extract the essential information from the detection model, a repeated detection approach is employed on biomedical images where various rotation angles are applied and a dice loss with the projection of the rotated detection results is utilized as a supervised signal in training our segmentation model. In biomedical imaging tasks involving convex objects, such as nuclei instance segmentation, PAL outperforms the known models (e.g., BoxInst) that rely solely on box supervision. Furthermore, PAL achieves comparable performance with mask-supervised models including Mask R-CNN and Cascade Mask R-CNN. Interestingly, PAL also demonstrates remarkable performance on non-convex object instance segmentation tasks, for example, surgical instrument and organ instance segmentation. Jintai Chen, Kai Zhang 0053, Jiahuan Yan, Bang Du, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | T2G-FORMER: Organizing Tabular Features into Relation Graphs Promotes Heterogeneous Feature InteractionabstractRecent development of deep neural networks (DNNs) for tabular learning has largely benefited from the capability of DNNs for automatic feature interaction. However, the heterogeneity nature of tabular features makes such features relatively independent, and developing effective methods to promote tabular feature interaction still remains an open problem. In this paper, we propose a novel Graph Estimator, which automatically estimates the relations among tabular features and builds graphs by assigning edges between related features. Such relation graphs organize independent tabular features into a kind of graph data such that interaction of nodes (tabular features) can be conducted in an orderly fashion. Based on our proposed Graph Estimator, we present a bespoke Transformer network tailored for tabular learning, called T2G-Former, which processes tabular data by performing tabular feature interaction guided by the relation graphs. A specific Cross-level Readout collects salient features predicted by the layers in T2G-Former across different levels, and attains global semantics for final prediction. Comprehensive experiments show that our T2G-Former achieves superior performance among DNNs and is competitive with non-deep Gradient Boosted Decision Tree models. The code and detailed results are available at https://github.com/jyansir/t2g-former. Jiahuan Yan, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001 |
AAAI | 2 |
| 2023 | Ord2Seq: Regarding Ordinal Regression as Label Sequence PredictionabstractOrdinal regression refers to classifying object instances into ordinal categories. It has been widely studied in many scenarios, such as medical disease grading and movie rating. Known methods focused only on learning inter-class ordinal relationships, but still incur limitations in distinguishing adjacent categories thus far. In this paper, we propose a simple sequence prediction framework for ordinal regression called Ord2Seq, which, for the first time, transforms each ordinal category label into a special label sequence and thus regards an ordinal regression task as a sequence prediction process. In this way, we decompose an ordinal regression task into a series of recursive binary classification steps, so as to subtly distinguish adjacent categories. Comprehensive experiments show the effectiveness of distinguishing adjacent categories for performance improvement and our new approach exceeds state-of-the-art performances in four different scenarios. Codes are available at https://github.com/wjh892521292/Ord2Seq. Jintai Chen, Tingting Chen 0002, Danny Ziyi Chen, Jian Wu 0001 |
ICCV | 3 |
| 2023 | TabCaps: A Capsule Neural Network for Tabular Data Classification with BoW Routing
Jintai Chen, Kuanlun Liao, Yanwen Fang, Danny Ziyi Chen, Jian Wu 0001 |
ICLR | 1 |
| 2023 | Cross-Layer Retrospective Retrieving via Layer Attention
Yanwen Fang, Yuxi Cai, Jintai Chen, Jingyu Zhao 0001, Guangjian Tian |
ICLR | 3 |
| 2023 | GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta LabelsabstractSince annotating medical images for segmentation tasks commonly incurs expensive costs, it is highly desirable to design an annotation-efficient method to alleviate the annotation burden. Recently, contrastive learning has exhibited a great potential in learning robust representations to boost downstream tasks with limited labels. In medical imaging scenarios, ready-made meta labels (i.e., specific attribute information of medical images) inherently reveal semantic relationships among images, which have been used to define positive pairs in previous work. However, the multi-perspective semantics revealed by various meta labels are usually incompatible and can incur intractable "semantic contradiction" when combining different meta labels. In this paper, we tackle the issue of "semantic contradiction" in a gradient-guided manner using our proposed Gradient Mitigator method, which systematically unifies multi-perspective meta labels to enable a pre-trained model to attain a better high-level semantic recognition ability. Moreover, we emphasize that the fine-grained discrimination ability is vital for segmentation-oriented pre-training, and develop a novel method called Gradient Filter to dynamically screen pixel pairs with the most discriminating power based on the magnitude of gradients. Comprehensive experiments on four medical image segmentation datasets verify that our new method GCL: (1) learns informative image representations and considerably boosts segmentation performance with limited labels, and (2) shows promising generalizability on out-of-distribution datasets. Jintai Chen, Jiahuan Yan, Yiheng Zhu 0002, Danny Ziyi Chen, Jian Wu 0001 |
ACM Multimedia | 2 |
| 2023 | Robust Training of Graph Neural Networks via Noise GovernanceabstractGraph Neural Networks (GNNs) have become widely-used models for semi-supervised learning. However, the robustness of GNNs in the presence of label noise remains a largely under-explored problem. In this paper, we consider an important yet challenging scenario where labels on nodes of graphs are not only noisy but also scarce. In this scenario, the performance of GNNs is prone to degrade due to label noise propagation and insufficient learning. To address these issues, we propose a novel RTGNN (Robust Training of Graph Neural Networks via Noise Governance) framework that achieves better robustness by learning to explicitly govern label noise. More specifically, we introduce self-reinforcement and consistency regularization as supplemental supervision. The self-reinforcement supervision is inspired by the memorization effects of deep neural networks and aims to correct noisy labels. Further, the consistency regularization prevents GNNs from overfitting to noisy labels via mimicry loss in both the inter-view and intra-view perspectives. To leverage such supervisions, we divide labels into clean and noisy types, rectify inaccurate labels, and further generate pseudo-labels on unlabeled nodes. Supervision for nodes with different types of labels is then chosen adaptively. This enables sufficient learning from clean labels while limiting the impact of noisy ones. We conduct extensive experiments to evaluate the effectiveness of our RTGNN framework, and the results validate its consistent superior performance over state-of-the-art methods with two types of label noises and various noise rates. Siyi Qian, Haochao Ying, Renjun Hu, Jingbo Zhou 0003, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001 |
WSDM | 5 |
| 2023 | D-former: a U-shaped Dilated Transformer for 3D medical image segmentation
Kuanlun Liao, Jintai Chen, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001 |
Neural Comput. Appl. | 3 |
| 2023 | Identifying Electrocardiogram Abnormalities Using a Handcrafted-Rule-Enhanced Neural NetworkabstractA large number of people suffer from life-threatening cardiac abnormalities, and electrocardiogram (ECG) analysis is beneficial to determining whether an individual is at risk of such abnormalities. Automatic ECG classification methods, especially the deep learning based ones, have been proposed to detect cardiac abnormalities using ECG records, showing good potential to improve clinical diagnosis and help early prevention of cardiovascular diseases. However, the predictions of the known neural networks still do not satisfactorily meet the needs of clinicians, and this phenomenon suggests that some information used in clinical diagnosis may not be well captured and utilized by these methods. In this paper, we introduce some rules into convolutional neural networks, which help present clinical knowledge to deep learning based ECG analysis, in order to improve automated ECG diagnosis performance. Specifically, we propose a Handcrafted-Rule-enhanced Neural Network (called HRNN) for ECG classification with standard 12-lead ECG input, which consists of a rule inference module and a deep learning module. Experiments on two large-scale public ECG datasets show that our new approach considerably outperforms existing state-of-the-art methods. Further, our proposed approach not only can improve the diagnosis performance, but also can assist in detecting mislabelled ECG samples. Yuexin Bian, Jintai Chen, Xiaoxian Yang, Danny Ziyi Chen, Jian Wu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | DANets: Deep Abstract Networks for Tabular Data Classification and RegressionabstractTabular data are ubiquitous in real world applications. Although many commonly-used neural components (e.g., convolution) and extensible neural networks (e.g., ResNet) have been developed by the machine learning community, few of them were effective for tabular data and few designs were adequately tailored for tabular data structures. In this paper, we propose a novel and flexible neural component for tabular data, called Abstract Layer (AbstLay), which learns to explicitly group correlative input features and generate higher-level features for semantics abstraction. Also, we design a structure re-parameterization method to compress the trained AbstLay, thus reducing the computational complexity by a clear margin in the reference phase. A special basic block is built using AbstLays, and we construct a family of Deep Abstract Networks (DANets) for tabular data classification and regression by stacking such blocks. In DANets, a special shortcut path is introduced to fetch information from raw tabular features, assisting feature interactions across different levels. Comprehensive experiments on seven real-world tabular datasets show that our AbstLay and DANets are effective for tabular data classification and regression, and the computational complexity is superior to competitive methods. Besides, we evaluate the performance gains of DANet as it goes deep, verifying the extendibility of our method. Our code is available at https://github.com/WhatAShot/DANet. Jintai Chen, Kuanlun Liao, Yao Wan 0001, Danny Ziyi Chen, Jian Wu 0001 |
AAAI | 1 |
| 2022 | ME-GAN: Learning Panoptic Electrocardio Representations for Multi-view ECG Synthesis Conditioned on Heart DiseasesabstractElectrocardiogram (ECG) is a widely used non-invasive diagnostic tool for heart diseases. Many studies have devised ECG analysis models (e.g., classifiers) to assist diagnosis. As an upstream task, researches have built generative models to synthesize ECG data, which are beneficial to providing training samples, privacy protection, and annotation reduction. However, previous generative methods for ECG often neither synthesized multi-view data, nor dealt with heart disease conditions. In this paper, we propose a novel disease-aware generative adversarial network for multi-view ECG synthesis called ME-GAN, which attains panoptic electrocardio representations conditioned on heart diseases and projects the representations onto multiple standard views to yield ECG signals. Since ECG manifestations of heart diseases are often localized in specific waveforms, we propose a new "mixup normalization" to inject disease information precisely into suitable locations. In addition, we propose a "view discriminator" to revert disordered ECG views into a pre-determined order, supervising the generator to obtain ECG representing correct view characteristics. Besides, a new metric, rFID, is presented to assess the quality of the synthesized ECG signals. Comprehensive experiments verify that our ME-GAN performs well on multi-view ECG signal synthesis with trusty morbid manifestations. Jintai Chen, Kuanlun Liao, Haochao Ying, Danny Ziyi Chen, Jian Wu 0001 |
ICML | 1 |
| 2022 | Self-learning and One-Shot Learning Based Single-Slice Annotation for 3D Medical Image Segmentation
Bo Zheng 0011, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (8) | 3 |
| 2022 | ChroNet: A multi-task learning based approach for prediction of multiple chronic diseases
Ruiwei Feng, Xuechen Liu 0004, Tingting Chen 0002, Jintai Chen, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001 |
Multim. Tools Appl. | 5 |
| 2021 | A Receptor Skeleton for Capsule Neural NetworksabstractIn previous Capsule Neural Networks (CapsNets), routing algorithms often performed clustering processes to assemble the child capsules’ representations into parent capsules. Such routing algorithms were typically implemented with iterative processes and incurred high computing complexity. This paper presents a new capsule structure, which contains a set of optimizable receptors and a transmitter is devised on the capsule’s representation. Specifically, child capsules’ representations are sent to the parent capsules whose receptors match well the transmitters of the child capsules’ representations, avoiding applying computationally complex routing algorithms. To ensure the receptors in a CapsNet work cooperatively, we build a skeleton to organize the receptors in different capsule layers in a CapsNet. The receptor skeleton assigns a share-out objective for each receptor, making the CapsNet perform as a hierarchical agglomerative clustering process. Comprehensive experiments verify that our approach facilitates efficient clustering processes, and CapsNets with our approach significantly outperform CapsNets with previous routing algorithms on image classification, affine transformation generalization, overlapped object recognition, and representation semantic decoupling. Jintai Chen, Hongyun Yu, Chengde Qian, Danny Ziyi Chen, Jian Wu 0001 |
ICML | 1 |
| 2021 | Electrocardio Panorama: Synthesizing New ECG views with Self-supervisionabstractMulti-lead electrocardiogram (ECG) provides clinical information of heartbeats from several fixed viewpoints determined by the lead positioning. However, it is often not satisfactory to visualize ECG signals in these fixed and limited views, as some clinically useful information is represented only from a few specific ECG viewpoints. For the first time, we propose a new concept, Electrocardio Panorama, which allows visualizing ECG signals from any queried viewpoints. To build Electrocardio Panorama, we assume that an underlying electrocardio field exists, representing locations, magnitudes, and directions of ECG signals. We present a Neural electrocardio field Network (Nef-Net), which first predicts the electrocardio field representation by using a sparse set of one or few input ECG views and then synthesizes Electrocardio Panorama based on the predicted representations. Specially, to better disentangle electrocardio field information from viewpoint biases, a new Angular Encoding is proposed to process viewpoint angles. Also, we propose a self-supervised learning approach called Standin Learning, which helps model the electrocardio field without direct supervision. Further, with very few modifications, Nef-Net can synthesize ECG signals from scratch. Experiments verify that our Nef-Net performs well on Electrocardio Panorama synthesis, and outperforms the previous work on the auxiliary tasks (ECG view transformation and ECG synthesis from scratch). The codes and the division labels of cardiac cycles and ECG deflections on Tianchi ECG and PTB datasets are available at https://github.com/WhatAShot/Electrocardio-Panorama. Jintai Chen, Xiangshang Zheng, Hongyun Yu, Danny Ziyi Chen, Jian Wu 0001 |
IJCAI | 1 |
| 2021 | A semi-supervised deep convolutional framework for signet ring cell detection
Haochao Ying, Qingyu Song 0004, Jintai Chen, Tingting Liang, Jingjing Gu, Fuzhen Zhuang, Danny Ziyi Chen, Jian Wu 0001 |
Neurocomputing | 3 |
| 2021 | A Transfer Learning Based Super-Resolution Microscopy for Biopsy Slice Images: The Joint Methods PerspectiveabstractHigher-resolution biopsy slice images reveal many details, which are widely used in medical practice. However, taking high-resolution slice images is more costly than taking low-resolution ones. In this paper, we propose a joint framework containing a novel transfer learning strategy and a deep super-resolution framework to generate high-resolution slice images from low-resolution ones. The super-resolution framework called SRFBN+ is proposed by modifying a state-of-the-art framework SRFBN. Specifically, the structure of the feedback block of SRFBN was modified to be more flexible. Besides, it is challenging to use typical transfer learning strategies directly for the tasks on slice images, as the patterns on different types of biopsy slice images are varying. To this end, we propose a novel transfer learning strategy, called Channel Fusion Transfer Learning (CF-Trans). CF-Trans builds a middle domain by fusing the data manifolds of the source domain and the target domain, serving as a springboard for knowledge transfer. Thus, in the transfer learning setting, SRFBN+ can be trained on the source domain and then the middle domain and finally the target domain. Experiments on biopsy slice images validate SRFBN+ works well in generating super-resolution slice images, and CF-Trans is an efficient transfer learning strategy. Jintai Chen, Haochao Ying, Xuechen Liu 0004, Jingjing Gu, Ruiwei Feng, Tingting Chen 0002, Honghao Gao, Jian Wu 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | A Deep Learning Approach for Colonoscopy Pathology WSI Analysis: Accurate Segmentation and ClassificationabstractColorectal cancer (CRC) is one of the most life-threatening malignancies. Colonoscopy pathology examination can identify cells of early-stage colon tumors in small tissue image slices. But, such examination is time-consuming and exhausting on high resolution images. In this paper, we present a new framework for colonoscopy pathology whole slide image (WSI) analysis, including lesion segmentation and tissue diagnosis. Our framework contains an improved U-shape network with a VGG net as backbone, and two schemes for training and inference, respectively (the training scheme and inference scheme). Based on the characteristics of colonoscopy pathology WSI, we introduce a specific sampling strategy for sample selection and a transfer learning strategy for model training in our training scheme. Besides, we propose a specific loss function, class-wise DSC loss, to train the segmentation network. In our inference scheme, we apply a sliding-window based sampling strategy for patch generation and diploid ensemble (data ensemble and model ensemble) for the final prediction. We use the predicted segmentation mask to generate the classification probability for the likelihood of WSI being malignant. To our best knowledge, DigestPath 2019 is the first challenge and the first public dataset available on colonoscopy tissue screening and segmentation, and our proposed framework yields good performance on this dataset. Our new framework achieved a DSC of 0.7789 and AUC of 1 on the online test dataset, and we won the [Formula: see text] place in the DigestPath 2019 Challenge (task 2). Our code is available at https://github.com/bhfs9999/colonoscopy_tissue_screen_and_segmentation. Ruiwei Feng, Xuechen Liu 0004, Jintai Chen, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Interactive Few-Shot Learning: Limited Supervision, Better Medical Image SegmentationabstractMany known supervised deep learning methods for medical image segmentation suffer an expensive burden of data annotation for model training. Recently, few-shot segmentation methods were proposed to alleviate this burden, but such methods often showed poor adaptability to the target tasks. By prudently introducing interactive learning into the few-shot learning strategy, we develop a novel few-shot segmentation approach called Interactive Few-shot Learning (IFSL), which not only addresses the annotation burden of medical image segmentation models but also tackles the common issues of the known few-shot segmentation methods. First, we design a new few-shot segmentation structure, called Medical Prior-based Few-shot Learning Network (MPrNet), which uses only a few annotated samples (e.g., 10 samples) as support images to guide the segmentation of query images without any pre-training. Then, we propose an Interactive Learning-based Test Time Optimization Algorithm (IL-TTOA) to strengthen our MPrNet on the fly for the target task in an interactive fashion. To our best knowledge, our IFSL approach is the first to allow few-shot segmentation models to be optimized and strengthened on the target tasks in an interactive and controllable manner. Experiments on four few-shot segmentation tasks show that our IFSL approach outperforms the state-of-the-art methods by more than 20% in the DSC metric. Specifically, the interactive optimization algorithm (IL-TTOA) further contributes ~10% DSC improvement for the few-shot segmentation models. Ruiwei Feng, Xiangshang Zheng, Tianxiang Gao, Jintai Chen, Wenzhe Wang, Danny Ziyi Chen, Jian Wu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Flow-Mixup: Classifying Multi-labeled Medical Images with Corrupted LabelsabstractIn clinical practice, medical image interpretation often involves multi-labeled classification, since the affected parts of a patient tend to present multiple symptoms or comorbidities. Recently, deep learning based frameworks have attained expertlevel performance on medical image interpretation, which can be attributed partially to large amounts of accurate annotations. However, manually annotating massive amounts of medical images is impractical, while automatic annotation is fast but imprecise (possibly introducing corrupted labels). In this work, we propose a new regularization approach, called Flow-Mixup, for multi-labeled medical image classification with corrupted labels. Flow-Mixup guides the models to capture robust features for each abnormality, thus helping handle corrupted labels effectively and making it possible to apply automatic annotation. Specifically, Flow-Mixup decouples the extracted features by adding constraints to the hidden states of the models. Also, FlowMixup is more stable and effective comparing to other known regularization methods, as shown by theoretical and empirical analyses. Experiments on two electrocardiogram datasets and a chest X-ray dataset containing corrupted labels verify that FlowMixup is effective and insensitive to corrupted labels. Jintai Chen, Hongyun Yu, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001 |
BIBM | 1 |
| 2020 | A Hierarchical Graph Network for 3D Object Detection on Point Cloudsabstract3D object detection on point clouds finds many applications. However, most known point cloud object detection methods did not adequately accommodate the characteristics (e.g., sparsity) of point clouds, and thus some key semantic information (e.g., shape information) is not well captured. In this paper, we propose a new graph convolution (GConv) based hierarchical graph network (HGNet) for 3D object detection, which processes raw point clouds directly to predict 3D bounding boxes. HGNet effectively captures the relationship of the points and utilizes the multi-level semantics for object detection. Specially, we propose a novel shape-attentive GConv (SA-GConv) to capture the local shape features, by modelling the relative geometric positions of points to describe object shapes. An SA-GConv based U-shape network captures the multi-level features, which are mapped into an identical feature space by an improved voting module and then further utilized to generate proposals. Next, a new GConv based Proposal Reasoning Module reasons on the proposals considering the global scene semantics, and the bounding boxes are then predicted. Consequently, our new framework outperforms state-of-the-art methods on two large-scale point cloud datasets, by ~4% mean average precision (mAP) on SUN RGB-D and by ~3% mAP on ScanNet-V2. Jintai Chen, Biwen Lei, Qingyu Song 0004, Haochao Ying, Danny Ziyi Chen, Jian Wu 0001 |
CVPR | 1 |
| 2020 | Doctor Imitator: A Graph-Based Bone Age Assessment Framework Using Hand Radiographs
Jintai Chen, Bohan Yu, Biwen Lei, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (6) | 1 |
| 2019 | Multi-view Learning with Feature Level Fusion for Cervical Dysplasia Diagnosis
Tingting Chen 0002, Xinjun Ma, Xuechen Liu 0004, Wenzhe Wang, Ruiwei Feng, Jintai Chen, Chunnv Yuan, Weiguo Lu, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (1) | 6 |
| 2019 | LSRC: A Long-Short Range Context-Fusing Framework for Automatic 3D Vertebra Localization
Jintai Chen, Ruoqian Guo, Bohan Yu, Tingting Chen 0002, Wenzhe Wang, Ruiwei Feng, Danny Ziyi Chen, Jian Wu 0001 |
MICCAI (6) | 1 |