EDBT 2026 Demo / reviewers in the wild / expert
Cheng Jin 0003
dblp:84/4201-3
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-3522-3592ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning with less supervision: A survey of label-efficient learning for medical image analysis
Cheng Jin 0003, Zhengrui Guo, Yi Lin 0009, Luyang Luo, Hao Chen 0011 |
Medical Image Anal. | 1 |
| 2026 | LLM-Driven Medical Report Generation via Communication-Efficient Heterogeneous Federated LearningabstractLarge Language Models (LLMs) have demonstrated significant potential in Medical Report Generation (MRG), yet their development requires large amounts of medical image-report pairs, which are commonly scattered across multiple centers. Centralizing these data is exceptionally challenging due to privacy regulations, thereby impeding model development and broader adoption of LLM-driven MRG models. To address this challenge, we present FedMRG, the first framework that leverages Federated Learning (FL) to enable privacy-preserving, multi-center development of LLM-driven MRG models, specifically designed to overcome the critical challenge of communication-efficient LLM training under multi-modal data heterogeneity. To start with, our framework tackles the fundamental challenge of communication overhead in federated LLM tuning by employing low-rank factorization to efficiently decompose parameter updates, significantly reducing gradient transmission costs and making LLM-driven MRG feasible in bandwidth-constrained FL settings. Furthermore, we observed the dual heterogeneity in MRG under the FL scenario: varying image characteristics across medical centers, as well as diverse reporting styles and terminology preferences. To address the data heterogeneity, we further enhance FedMRG with (1) client-aware contrastive learning in the MRG encoder, coupled with diagnosis-driven prompts, which capture both globally generalizable and locally distinctive features while maintaining diagnostic accuracy; and (2) a dual-adapter mutual boosting mechanism in the MRG decoder that harmonizes generic and specialized adapters to address variations in reporting styles and terminology. Through extensive evaluation of our established FL-MRG benchmark, we demonstrate the generalizability and adaptability of FedMRG, underscoring its potential in harnessing multi-center data and generating clinically accurate reports while maintaining communication efficiency. Haoxuan Che, Haibo Jin, Zhengrui Guo, Yi Lin 0009, Cheng Jin 0003, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | GameGen-X: Interactive Open-world Game Video GenerationabstractWe introduce GameGen-$\mathbb{X}$, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos.
This model facilitates high-quality, open-domain generation by approximating various game elements, such as innovative characters, dynamic environments, complex actions, and diverse events.
Additionally, it provides interactive controllability, predicting and altering future content based on the current clip, thus allowing for gameplay simulation.
To realize this vision, we first collected and built an Open-World Video Game Dataset (OGameData) from scratch.
It is the first and largest dataset for open-world game video generation and control, which comprises over one million diverse gameplay video clips with informative captions.
GameGen-$\mathbb{X}$ undergoes a two-stage training process, consisting of pre-training and instruction tuning.
Firstly, the model was pre-trained via text-to-video generation and video continuation, enabling long-sequence open-domain game video generation with improved fidelity and coherence.
Further, to achieve interactive controllability, we designed InstructNet to incorporate game-related multi-modal control signal experts.
This allows the model to adjust latent representations based on user inputs, advancing the integration of character interaction and scene content control in video generation.
During instruction tuning, only the InstructNet is updated while the pre-trained foundation model is frozen, enabling the integration of interactive controllability without loss of diversity and quality of generated content.
GameGen-$\mathbb{X}$ contributes to advancements in open-world game design using generative models.
It demonstrates the potential of generative models to serve as auxiliary tools to traditional rendering techniques, demonstrating the potential for merging creative generation with interactive capabilities.
The project will be available at https://github.com/GameGen-X/GameGen-X. Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin 0003, Hao Chen 0011 |
ICLR | 4 |
| 2025 | HMIL: Hierarchical Multi-Instance Learning for Fine-Grained Whole Slide Image ClassificationabstractFine-grained classification of whole slide images (WSIs) is essential in precision oncology, enabling precise cancer diagnosis and personalized treatment strategies. The core of this task involves distinguishing subtle morphological variations within the same broad category of gigapixel-resolution images, which presents a significant challenge. While the multi-instance learning (MIL) paradigm alleviates the computational burden of WSIs, existing MIL methods often overlook hierarchical label correlations, treating fine-grained classification as a flat multi-class classification task. To overcome these limitations, we introduce a novel hierarchical multi-instance learning (HMIL) framework. By facilitating on the hierarchical alignment of inherent relationships between different hierarchy of labels at instance and bag level, our approach provides a more structured and informative learning process. Specifically, HMIL incorporates a class-wise attention mechanism that aligns hierarchical information at both the instance and bag levels. Furthermore, we introduce supervised contrastive learning to enhance the discriminative capability for fine-grained classification and a curriculum-based dynamic weighting module to adaptively balance the hierarchical feature during training. Extensive experiments on our large-scale cytology cervical cancer (CCC) dataset and two public histology datasets, BRACS and PANDA, demonstrate the state-of-the-art class-wise and overall performance of our HMIL framework. Our source code is available at https://github.com/ChengJin-git/HMIL. Cheng Jin 0003, Luyang Luo, Huangjing Lin, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Shapley Values-Enabled Progressive Pseudo Bag Augmentation for Whole-Slide Image ClassificationabstractIn computational pathology, whole-slide image (WSI) classification presents a formidable challenge due to its gigapixel resolution and limited fine-grained annotations. Multiple-instance learning (MIL) offers a weakly supervised solution, yet refining instance-level information from bag-level labels remains challenging. While most of the conventional MIL methods use attention scores to estimate instance importance scores (IIS) which contribute to the prediction of the slide labels, these often lead to skewed attention distributions and inaccuracies in identifying crucial instances. To address these issues, we propose a new approach inspired by cooperative game theory: employing Shapley values to assess each instance's contribution, thereby improving IIS estimation. The computation of the Shapley value is then accelerated using attention, meanwhile retaining the enhanced instance identification and prioritization. We further introduce a framework for the progressive assignment of pseudo bags based on estimated IIS, encouraging more balanced attention distributions in MIL models. Our extensive experiments on CAMELYON-16, BRACS, TCGA-LUNG, and TCGA-BRCA datasets show our method's superiority over existing state-of-the-art approaches, offering enhanced interpretability and class-wise insights. Our source code is available at https://github.com/RenaoYan/PMIL. Renao Yan, Qiehe Sun, Cheng Jin 0003, Yonghong He, Tian Guan, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 3 |
| 2023 | WYTIWYR: A User Intent-Aware Framework with Multi-modal Inputs for Visualization RetrievalabstractAbstract Retrieving charts from a large corpus is a fundamental task that can benefit numerous applications such as visualization recommendations. The retrieved results are expected to conform to both explicit visual attributes (e.g., chart type, colormap) and implicit user intents (e.g., design style, context information) that vary upon application scenarios. However, existing example‐based chart retrieval methods are built upon non‐decoupled and low‐level visual features that are hard to interpret, while definition‐based ones are constrained to pre‐defined attributes that are hard to extend. In this work, we propose a new framework, namelyWYTIWYR (What‐You‐Think‐Is‐What‐You‐Retrieve), that integrates user intents into the chart retrieval process. The framework consists of two stages: first, theAnnotationstage disentangles the visual attributes within the query chart; and second, theRetrievalstage embeds the user's intent with customized text prompt as well as bitmap query chart, to recall targeted retrieval result. We develop aprototypeWYTIWYRsystem leveraging a contrastive language‐image pre‐training (CLIP) model to achieve zero‐shot classification as well as multi‐modal input encoding, and test the prototype on a large corpus with charts crawled from the Internet. Quantitative experiments, case studies, and qualitative interviews are conducted. The results demonstrate the usability and effectiveness of our proposed framework. Shishi Xiao, Yihan Hou, Cheng Jin 0003, Wei Zeng 0004 |
Comput. Graph. Forum | 3 |
| 2022 | SpanConv: A New Convolution via Spanning Kernel Space for Lightweight PansharpeningabstractStandard convolution operations can effectively perform feature extraction and representation but result in high computational cost, largely due to the generation of the original convolution kernel corresponding to the channel dimension of the feature map, which will cause unnecessary redundancy. In this paper, we focus on kernel generation and present an interpretable span strategy, named SpanConv, for the effective construction of kernel space. Specifically, we first learn two navigated kernels with single channel as bases, then extend the two kernels by learnable coefficients, and finally span the two sets of kernels by their linear combination to construct the so-called SpanKernel. The proposed SpanConv is realized by replacing plain convolution kernel by SpanKernel. To verify the effectiveness of SpanConv, we design a simple network with SpanConv. Experiments demonstrate the proposed network significantly reduces parameters comparing with benchmark networks for remote sensing pansharpening, while achieving competitive performance and excellent generalization. Code is available at https://github.com/zhi-xuan-chen/IJCAI-2022 SpanConv. Zhi-Xuan Chen, Cheng Jin 0003, Tianjing Zhang, Liang-Jian Deng |
IJCAI | 2 |
| 2021 | Weighted Shallow-Deep Feature Fusion Network for PansharpeningabstractIn this paper, we propose a novel weighted shallow-deep feature fusion convolutional neural network (WSDFNet) for the task of multispectral image pansharpening. This network could effectively overcome the drawback of the common identity skip connection (ISC), and propagate shallow features scaled by a novel adaptive skip weighter (ASW) to deeper layers. By the technique, it could favor the feature fusion in different network depths adequately, as well as yield a promising outcome. Experimental results on reduced- and full-resolution WorldView-3 dataset demonstrate the superiority of the WSDFNet compared with recent state-of-the-art (SOTA) pansharpening approaches. Moreover, WSDFNet is also verified as a lightweight network. Zi-Rong Jin, Tianjing Zhang, Cheng Jin 0003, Liang-Jian Deng |
IGARSS | 3 |
| 2021 | Progressive Band-Separated Convolutional Neural Network for Multispectral PansharpeningabstractRecently, convolutional neural networks (CNNs) have been introduced to pansharpening for enhancing fusion accuracy and overcoming the drawbacks of the conventional methods. However, most of methods based on CNN fail to distinguish the difference of multispectral bands, and only use a uniform set of convolutional kernels to extract features. In this paper, we design a progressive, band-separated convolutional network architecture for discriminatively learning the features and relation among spectral bands, aiming to address the problem mentioned before. More specifically, the proposed architecture mainly consists of three aspects. First, to accurately preserve the spectral peculiarities, we divide the multispectral input image in terms of its bands into several groups. Second, our original panchromatic and multispectral inputs are filtered by a high-pass operation to further yield more spatial details. Third, we use a spectral fusion module (SFM) for each group and associate them to progressively assemble the whole architecture. It is worth mentioning that the architecture could be integrated into any other competitive CNNs to improve the performance. Both visual and quantitative experiments have demonstrated that our proposed method outperforms recent state-of-the-art pansharpening techniques. Shishi Xiao, Cheng Jin 0003, Tianjing Zhang, Ran Ran 0001, Liang-Jian Deng |
IGARSS | 2 |
| 2021 | Detail Injection-Based Deep Convolutional Neural Networks for PansharpeningabstractThe fusion of high spatial resolution panchromatic (PAN) data with simultaneously acquired multispectral (MS) data with the lower spatial resolution is a hot topic, which is often called pansharpening. In this article, we exploit the combination of machine learning techniques and fusion schemes introduced to address the pansharpening problem. In particular, deep convolutional neural networks (DCNNs) are proposed to solve this issue. The latter is combined first with the traditional component substitution and multiresolution analysis fusion schemes in order to estimate the nonlinear injection models that rule the combination of the upsampled low-resolution MS image with the extracted details exploiting the two philosophies. Furthermore, inspired by these two approaches, we also developed another DCNN for pansharpening. This is fed by the direct difference between the PAN image and the upsampled low-resolution MS image. Extensive experiments conducted both at reduced and full resolutions demonstrate that this latter convolutional neural network outperforms both the other detail injection-based proposals and several state-of-the-art pansharpening methods. Liang-Jian Deng, Gemine Vivone, Cheng Jin 0003, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 3 |