EDBT 2026 Demo / reviewers in the wild / expert
Jiahuan Yan
dblp:334/7537
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-2002-2579ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Generative modeling · 30% Kernel, tree and ensemble methods · 24% Segmentation and scene understanding · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 64% Medical and health informatics · 36% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods
tabular prediction |
2.3 | 3 | 2024 | Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPs · KDD 2024 Can a Deep Learning Model be a Sure Bet for Tabular Prediction? · KDD 2024 Making Pre-trained Language Models Great on Tabular Prediction · ICLR 2024 |
Machine learning › Generative modeling › diffusion model
diffusion bridge |
0.8 | 1 | 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov Bridges · NeurIPS 2024 |
Robotics › Robot manipulation
digital twin generation |
0.8 | 1 | 2024 | Personalized Heart Disease Detection via ECG Digital Twin Generation · IJCAI 2024 |
Machine learning › Generative modeling › generative model › continuous-time generative model
markov bridge |
0.8 | 1 | 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov Bridges · NeurIPS 2024 |
Machine learning › Generative modeling › synthetic data generation
tabular data augmentation |
0.8 | 1 | 2024 | Can a Deep Learning Model be a Sure Bet for Tabular Prediction? · KDD 2024 |
Bioinformatics and computational biology › protein design
inverse protein folding |
0.8 | 1 | 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov Bridges · NeurIPS 2024 |
Bioinformatics and computational biology
protein design |
0.8 | 1 | 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov Bridges · NeurIPS 2024 |
Computer vision › Segmentation and scene understanding
annotation-efficient segmentation |
0.7 | 1 | 2023 | GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta Labels · ACM Multimedia 2023 |
Machine learning › Generative modeling
generative flow networks |
0.7 | 1 | 2023 | Sample-efficient Multi-objective Molecular Optimization with GFlowNets · NeurIPS 2023 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.7 | 1 | 2023 | GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta Labels · ACM Multimedia 2023 |
Machine learning › Deep learning architectures and training
tabular data learning |
0.7 | 1 | 2023 | T2G-FORMER: Organizing Tabular Features into Relation Graphs Promotes Heterogeneous Feature Interaction · AAAI 2023 |
Bioinformatics and computational biology › molecular informatics
molecular design |
0.7 | 1 | 2023 | Sample-efficient Multi-objective Molecular Optimization with GFlowNets · NeurIPS 2023 |
Bioinformatics and computational biology › molecular informatics › cheminformatics › molecule generation
multi-objective molecular optimization |
0.7 | 1 | 2023 | Sample-efficient Multi-objective Molecular Optimization with GFlowNets · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.3 | 1 | 2025 | Small Models are LLM Knowledge Triggers for Medical Tabular Prediction · ICLR 2025 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.2 | 1 | 2024 | Can a Deep Learning Model be a Sure Bet for Tabular Prediction? · KDD 2024 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.2 | 1 | 2024 | Making Pre-trained Language Models Great on Tabular Prediction · ICLR 2024 |
Data mining › predictive modeling › tree-based models
gradient boosting decision tree |
0.2 | 1 | 2024 | Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPs · KDD 2024 |
Wearable and physiological sensing › vital sign monitoring
ECG monitoring |
0.2 | 1 | 2024 | Personalized Heart Disease Detection via ECG Digital Twin Generation · IJCAI 2024 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.2 | 1 | 2023 | GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta Labels · ACM Multimedia 2023 |
Methods — techniques the papers use, named apart from their topics
generative model · 2.3digital twin · 2.3synergy learning · 1.7self-prompting · 1.7knowledge distillation · 1.7gradient-boosted decision trees · 1.4structure encoder · 0.8relative magnitude tokenization · 0.8protein language model · 0.8intra-feature attention · 0.8gradient boosted decision trees · 0.8feature gate · 0.8data augmentation · 0.8backpropagation · 0.8attentive feedforward network · 0.8architecture pruning · 0.8multi-objective bayesian optimization · 0.7hypernetwork · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Small Models are LLM Knowledge Triggers for Medical Tabular PredictionabstractRecent development in large language models (LLMs) has demonstrated impressive domain proficiency on unstructured textual or multi-modal tasks. However, despite with intrinsic world knowledge, their application on structured tabular data prediction still lags behind, primarily due to the numerical insensitivity and modality discrepancy that brings a gap between LLM reasoning and statistical tabular learning. Unlike textual or vision data (e.g., electronic clinical notes or medical imaging data), tabular data is often presented in heterogeneous numerical values (e.g., CBC reports). This ubiquitous data format requires intensive expert annotation, and its numerical nature limits LLMs' capability to effectively transfer untapped domain expertise. In this paper, we propose SERSAL, a general self-prompting method by synergy learning with small models to enhance LLM tabular prediction in an unsupervised manner. Specifically, SERSAL utilizes the LLM's prior outcomes as original soft noisy annotations, which are dynamically leveraged to teach a better small student model. Reversely, the outcomes from the trained small model are used to teach the LLM to further refine its real capability. This process can be repeatedly applied to gradually distill refined knowledge for continuous progress. Comprehensive experiments on widely used medical domain tabular datasets show that, without access to gold labels, applying SERSAL to OpenAI GPT reasoning process attains substantial improvement compared to linguistic prompting methods, which serves as an orthogonal direction for tabular LLM, and increasing prompting bonus is observed as more powerful LLMs appear. Codes are available at https://github.com/jyansir/sersal. Jiahuan Yan, Jintai Chen, Chaowen Hu, Bo Zheng 0011, Yaojun Hu, Jimeng Sun 0001, Jian Wu 0001 |
ICLR | 1 |
| 2024 | Making Pre-trained Language Models Great on Tabular PredictionabstractThe transferability of deep neural networks (DNNs) has made significant progress in image and language processing. However, due to the heterogeneity among tables, such DNN bonus is still far from being well exploited on tabular data prediction (e.g., regression or classification tasks). Condensing knowledge from diverse domains, language models (LMs) possess the capability to comprehend feature names from various tables, potentially serving as versatile learners in transferring knowledge across distinct tables and diverse prediction tasks, but their discrete text representation space is inherently incompatible with numerical feature values in tables. In this paper, we present TP-BERTa, a specifically pre-trained LM for tabular data prediction. Concretely, a novel relative magnitude tokenization converts scalar numerical feature values to finely discrete, high-dimensional tokens, and an intra-feature attention approach integrates feature values with the corresponding feature names. Comprehensive experiments demonstrate that our pre-trained TP-BERTa leads the performance among tabular DNNs and is competitive with Gradient Boosted Decision Tree models in typical tabular data regime. Jiahuan Yan, Bo Zheng 0011, Yiheng Zhu 0002, Danny Ziyi Chen, Jimeng Sun 0001, Jian Wu 0001, Jintai Chen |
ICLR | 1 |
| 2024 | Personalized Heart Disease Detection via ECG Digital Twin Generation
Yaojun Hu, Jintai Chen, Lianting Hu, Dantong Li, Jiahuan Yan, Haochao Ying, Huiying Liang, Jian Wu 0001 |
IJCAI | 5 |
| 2024 | Can a Deep Learning Model be a Sure Bet for Tabular Prediction?abstractData organized in tabular format is ubiquitous in real-world applications, and users often craft tables with biased feature definitions and flexibly set prediction targets of their interests. Thus, a rapid development of a robust, effective, dataset-versatile, user-friendly tabular prediction approach is highly desired. While Gradient Boosting Decision Trees (GBDTs) and existing deep neural networks (DNNs) have been extensively utilized by professional users, they present several challenges for casual users, particularly: (i) the dilemma of model selection due to their different dataset preferences, and (ii) the need for heavy hyperparameter searching, failing which their performances are deemed inadequate. In this paper, we delve into this question: Can we develop a deep learning model that serves as a sure bet solution for a wide range of tabular prediction tasks, while also being user-friendly for casual users? We delve into three key drawbacks of deep tabular models, encompassing: (P1) lack of rotational variance property, (P2) large data demand, and (P3) over-smooth solution. We propose ExcelFormer, addressing these challenges through a semi-permeable attention module that effectively constrains the influence of less informative features to break the DNNs' rotational invariance property (for P1), data augmentation approaches tailored for tabular data (for P2), and attentive feedforward network to boost the model fitting capability (for P3). These designs collectively make ExcelFormer a sure bet solution for diverse tabular datasets. Extensive and stratified experiments conducted on real-world datasets demonstrate that our model outperforms previous approaches across diverse tabular data prediction tasks, and this framework can be friendly to casual users, offering ease of use without the heavy hyperparameter tuning. The codes are available at https://github.com/whatashot/excelformer. Jintai Chen, Jiahuan Yan, Qiyuan Chen 0003, Danny Ziyi Chen, Jian Wu 0001, Jimeng Sun 0001 |
KDD | 2 |
| 2024 | Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPsabstractTabular datasets play a crucial role in various applications.Thus, developing efficient, effective, and widely compatible prediction algorithms for tabular data is important.Currently, two prominent model types, Gradient Boosted Decision Trees (GBDTs) and Deep Neural Networks (DNNs), have demonstrated performance advantages on distinct tabular prediction tasks.However, selecting an effective model for a specific tabular dataset is challenging, often demanding time-consuming hyperparameter tuning.To address this model selection dilemma, this paper proposes a new framework that amalgamates the advantages of both GBDTs and DNNs, resulting in a DNN algorithm that is as efficient as GBDTs and is competitively effective regardless of dataset preferences for GBDTs or DNNs.Our idea is rooted in an observation that deep learning (DL) offers a larger parameter space that can represent a well-performing GBDT model, yet the current back-propagation optimizer struggles to efficiently discover such optimal functionality.On the other hand, during GBDT development, hard tree pruning, entropy-driven feature gate, and model ensemble have proved to be more adaptable to tabular data.By combining these key components, we present a Tree-hybrid simple MLP (T-MLP).In our framework, a tensorized, rapidly trained GBDT feature gate, a DNN architecture pruning approach, as well as a vanilla back-propagation optimizer collaboratively train a randomly initialized MLP model.Comprehensive experiments show that T-MLP is competitive with extensively tuned DNNs and GBDTs in their dominating tabular benchmarks (88 datasets) respectively, all achieved with compact model storage and significantly reduced training duration.The codes and full experiment results are available at https://github.com/jyansir/tmlp. Jiahuan Yan, Jintai Chen, Qianxing Wang, Danny Ziyi Chen, Jian Wu 0001 |
KDD | 1 |
| 2024 | Bridge-IF: Learning Inverse Protein Folding with Markov BridgesabstractInverse protein folding is a fundamental task in computational protein design, which aims to design protein sequences that fold into the desired backbone structures. While the development of machine learning algorithms for this task has seen significant success, the prevailing approaches, which predominantly employ a discriminative formulation, frequently encounter the error accumulation issue and often fail to capture the extensive variety of plausible sequences. To fill these gaps, we propose Bridge-IF, a generative diffusion bridge model for inverse folding, which is designed to learn the probabilistic dependency between the distributions of backbone structures and protein sequences. Specifically, we harness an expressive structure encoder to propose a discrete, informative prior derived from structures, and establish a Markov bridge to connect this prior with native sequences. During the inference stage, Bridge-IF progressively refines the prior sequence, culminating in a more plausible design. Moreover, we introduce a reparameterization perspective on Markov bridge models, from which we derive a simplified loss function that facilitates more effective training. We also modulate protein language models (PLMs) with structural conditions to precisely approximate the Markov bridge process, thereby significantly enhancing generation performance while maintaining parameter-efficient training. Extensive experiments on well-established benchmarks demonstrate that Bridge-IF predominantly surpasses existing baselines in sequence recovery and excels in the design of plausible proteins with high foldability. The code is available at https://github.com/violet-sto/Bridge-IF. Jialu Wu, Qiuyi Li, Jiahuan Yan, Mingze Yin, Jieping Ye |
NeurIPS | 4 |
| 2024 | Polygonal Approximation Learning for Convex Object Segmentation in Biomedical Images With Bounding Box SupervisionabstractAs a common and critical medical image analysis task, deep learning based biomedical image segmentation is hindered by the dependence on costly fine-grained annotations. To alleviate this data dependence, in this article, a novel approach, called Polygonal Approximation Learning (PAL), is proposed for convex object instance segmentation with only bounding-box supervision. The key idea behind PAL is that the detection model for convex objects already contains the necessary information for segmenting them since their convex hulls, which can be generated approximately by the intersection of bounding boxes, are equivalent to the masks representing the objects. To extract the essential information from the detection model, a repeated detection approach is employed on biomedical images where various rotation angles are applied and a dice loss with the projection of the rotated detection results is utilized as a supervised signal in training our segmentation model. In biomedical imaging tasks involving convex objects, such as nuclei instance segmentation, PAL outperforms the known models (e.g., BoxInst) that rely solely on box supervision. Furthermore, PAL achieves comparable performance with mask-supervised models including Mask R-CNN and Cascade Mask R-CNN. Interestingly, PAL also demonstrates remarkable performance on non-convex object instance segmentation tasks, for example, surgical instrument and organ instance segmentation. Jintai Chen, Kai Zhang 0053, Jiahuan Yan, Bang Du, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | T2G-FORMER: Organizing Tabular Features into Relation Graphs Promotes Heterogeneous Feature InteractionabstractRecent development of deep neural networks (DNNs) for tabular learning has largely benefited from the capability of DNNs for automatic feature interaction. However, the heterogeneity nature of tabular features makes such features relatively independent, and developing effective methods to promote tabular feature interaction still remains an open problem. In this paper, we propose a novel Graph Estimator, which automatically estimates the relations among tabular features and builds graphs by assigning edges between related features. Such relation graphs organize independent tabular features into a kind of graph data such that interaction of nodes (tabular features) can be conducted in an orderly fashion. Based on our proposed Graph Estimator, we present a bespoke Transformer network tailored for tabular learning, called T2G-Former, which processes tabular data by performing tabular feature interaction guided by the relation graphs. A specific Cross-level Readout collects salient features predicted by the layers in T2G-Former across different levels, and attains global semantics for final prediction. Comprehensive experiments show that our T2G-Former achieves superior performance among DNNs and is competitive with non-deep Gradient Boosted Decision Tree models. The code and detailed results are available at https://github.com/jyansir/t2g-former. Jiahuan Yan, Jintai Chen, Danny Ziyi Chen, Jian Wu 0001 |
AAAI | 1 |
| 2023 | GCL: Gradient-Guided Contrastive Learning for Medical Image Segmentation with Multi-Perspective Meta LabelsabstractSince annotating medical images for segmentation tasks commonly incurs expensive costs, it is highly desirable to design an annotation-efficient method to alleviate the annotation burden. Recently, contrastive learning has exhibited a great potential in learning robust representations to boost downstream tasks with limited labels. In medical imaging scenarios, ready-made meta labels (i.e., specific attribute information of medical images) inherently reveal semantic relationships among images, which have been used to define positive pairs in previous work. However, the multi-perspective semantics revealed by various meta labels are usually incompatible and can incur intractable "semantic contradiction" when combining different meta labels. In this paper, we tackle the issue of "semantic contradiction" in a gradient-guided manner using our proposed Gradient Mitigator method, which systematically unifies multi-perspective meta labels to enable a pre-trained model to attain a better high-level semantic recognition ability. Moreover, we emphasize that the fine-grained discrimination ability is vital for segmentation-oriented pre-training, and develop a novel method called Gradient Filter to dynamically screen pixel pairs with the most discriminating power based on the magnitude of gradients. Comprehensive experiments on four medical image segmentation datasets verify that our new method GCL: (1) learns informative image representations and considerably boosts segmentation performance with limited labels, and (2) shows promising generalizability on out-of-distribution datasets. Jintai Chen, Jiahuan Yan, Yiheng Zhu 0002, Danny Ziyi Chen, Jian Wu 0001 |
ACM Multimedia | 3 |
| 2023 | Sample-efficient Multi-objective Molecular Optimization with GFlowNetsabstractMany crucial scientific problems involve designing novel molecules with desired properties, which can be formulated as a black-box optimization problem over the *discrete* chemical space. In practice, multiple conflicting objectives and costly evaluations (e.g., wet-lab experiments) make the *diversity* of candidates paramount. Computational methods have achieved initial success but still struggle with considering diversity in both objective and search space. To fill this gap, we propose a multi-objective Bayesian optimization (MOBO) algorithm leveraging the hypernetwork-based GFlowNets (HN-GFN) as an acquisition function optimizer, with the purpose of sampling a diverse batch of candidate molecular graphs from an approximate Pareto front. Using a single preference-conditioned hypernetwork, HN-GFN learns to explore various trade-offs between objectives. We further propose a hindsight-like off-policy strategy to share high-performing molecules among different preferences in order to speed up learning for HN-GFN. We empirically illustrate that HN-GFN has adequate capacity to generalize over preferences. Moreover, experiments in various real-world MOBO settings demonstrate that our framework predominantly outperforms existing methods in terms of candidate quality and sample efficiency. The code is available at https://github.com/violet-sto/HN-GFN. Yiheng Zhu 0002, Jialu Wu, Chaowen Hu, Jiahuan Yan, Chang-Yu Hsieh, Tingjun Hou, Jian Wu 0001 |
NeurIPS | 4 |