EDBT 2026 Demo / reviewers in the wild / expert
Chong Ma 0004
dblp:189/0575-4
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0002-5068-8814ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning better contrastive view from radiologist's gaze
Sheng Wang 0014, Zihao Zhao 0002, Zixu Zhuang, Xi Ouyang, Lichi Zhang, Zheren Li, Chong Ma 0004, Tianming Liu 0001, Dinggang Shen, Qian Wang 0001 |
Pattern Recognit. | 7 |
| 2025 | Exploring the Trade-Offs: Unified Large Language Models vs Local Fine-Tuned Models for Highly-Specific Radiology NLI TaskabstractRecently, ChatGPT and GPT-4 have emerged and gained immense global attention due to their unparalleled performance in language processing. Despite demonstrating impressive capability in various open-domain tasks, their adequacy in highly specific fields like radiology remains untested. Radiology presents unique linguistic phenomena distinct from open-domain data due to its specificity and complexity. Assessing the performance of large language models (LLMs) in such specific domains is crucial not only for a thorough evaluation of their overall performance but also for providing valuable insights into future model design directions: whether model design should be generic or domain-specific. To this end, in this study, we evaluate the performance of ChatGPT/GPT-4 on a radiology natural language inference (NLI) task and compare it to other models fine-tuned specifically on task-related data samples. We also conduct a comprehensive investigation on ChatGPT/GPT-4’s reasoning ability by introducing varying levels of inference difficulty. Our results show that 1) ChatGPT and GPT-4 outperform other LLMs in the radiology NLI task and 2) other specifically fine-tuned Bert-based models require significant amounts of data samples to achieve comparable performance to ChatGPT/GPT-4. These findings not only demonstrate the feasibility and promise of constructing a generic model capable of addressing various tasks across different domains, but also highlight several key factors crucial for developing a unified model, particularly in a medical context, paving the way for future artificial general intelligence (AGI) systems. We release our code and data to the research community. Zihao Wu 0001, Lu Zhang 0050, Xiaowei Yu 0001, Zhengliang Liu, Lin Zhao 0004, Yiwei Li 0002, Haixing Dai, Chong Ma 0004, Gang Li 0001, Wei Liu 0146, Quanzheng Li, Dinggang Shen, Xiang Li 0001, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Big Data | 9 |
| 2025 | Integrating Eye Tracking With Grouped Fusion Networks for Semantic Segmentation on Mammogram ImagesabstractMedical image segmentation has seen great progress in recent years, largely due to the development of deep neural networks. However, unlike in computer vision, high-quality clinical data is relatively scarce, and the annotation process is often a burden for clinicians. As a result, the scarcity of medical data limits the performance of existing medical image segmentation models. In this paper, we propose a novel framework that integrates eye tracking information from experienced radiologists during the screening process to improve the performance of deep neural networks with limited data. Our approach, a grouped hierarchical network, guides the network to learn from its faults by using gaze information as weak supervision. We demonstrate the effectiveness of our framework on mammogram images, particularly for handling segmentation classes with large scale differences. We evaluate the impact of gaze information on medical image segmentation tasks and show that our method achieves better segmentation performance compared to state-of-the-art models. A robustness study is conducted to investigate the influence of distraction or inaccuracies in gaze collection. We also develop a convenient system for collecting gaze data without interrupting the normal clinical workflow. Our work offers novel insights into the potential benefits of integrating gaze information into medical image segmentation tasks. Jiaming Xie, Zhiming Cui 0001, Chong Ma 0004, Wenping Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2025 | ChatABL: Abductive Learning via Natural Language Interaction With ChatGPTabstractLarge language models (LLMs) such as ChatGPT have recently demonstrated significant potential in mathematical abilities, providing a valuable reasoning paradigm consistent with human natural language. However, LLMs currently have difficulty in bridging perception, language understanding, and reasoning (PLR) capabilities due to incompatibility of the underlying information flow among them, making their reasoning ability not fully elicited and challenging to accomplish complicated reasoning tasks autonomously. To resolve the above problem, a novel method called ChatABL is proposed by integrating LLMs into an abductive learning (ABL) framework, capable of unifying the three abilities effectively in a more user-friendly and understandable manner. Initially, the proposed method uses LLMs to correct the incomplete logical facts for optimizing the perception module, by summarizing and reorganizing domain knowledge represented in natural language format. Then, the perception module also provides necessary logical reasoning materials for feeding LLMs. Finally, these parts are integrated into a dynamic closed-loop system by introducing the feedback form and automatic learning strategies to mutually promote their performance. As a testbed, the variable-length handwritten equation decipherment (HED), an abstract expression of the Mayan calendar decoding, is used to demonstrate that ChatABL has reasoning ability beyond most existing state-of-the-art methods, which has been well-supported by comparative studies. To the best of authors' knowledge, the proposed ChatABL is the first attempt to explore a possible and novel avenue to approaching human-level cognitive ability via natural language interaction by means of ChatGPT. Tianyang Zhong, Yi Pan 0001, Yutong Zhang 0019, Yaonai Wei, Zhengliang Liu, Xiaozheng Wei, Wenjun Li 0001, Chong Ma 0004, Xi Jiang 0001, Dinggang Shen, Junwei Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 10 |
| 2024 | Eye-gaze Guided Multi-modal Alignment for Medical Representation LearningabstractIn the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework. Chong Ma 0004, Hanqi Jiang, Wenting Chen, Yiwei Li 0002, Zihao Wu 0001, Xiaowei Yu 0001, Zhengliang Liu, Lei Guo 0002, Dajiang Zhu, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
NeurIPS | 1 |
| 2024 | Brain Structural Connectivity Guided Vision Transformers for Identification of Functional Connectivity Characteristics in Preterm NeonatesabstractPreterm birth is the leading cause of death in children under five years old, and is associated with a wide sequence of complications in both short and long term. In view of rapid neurodevelopment during the neonatal period, preterm neonates may exhibit considerable functional alterations compared to term ones. However, the identified functional alterations in previous studies merely achieve moderate classification performance, while more accurate functional characteristics with satisfying discrimination ability for better diagnosis and therapeutic treatment is underexplored. To address this problem, we propose a novel brain structural connectivity (SC) guided Vision Transformer (SCG-ViT) to identify functional connectivity (FC) differences among three neonatal groups: preterm, preterm with early postnatal experience, and term. Particularly, inspired by the neuroscience-derived information, a novel patch token of SC/FC matrix is defined, and the SC matrix is then adopted as an effective mask into the ViT model to screen out input FC patch embeddings with weaker SC, and to focus on stronger ones for better classification and identification of FC differences among the three groups. The experimental results on multi-modal MRI data of 437 neonatal brains from publicly released Developing Human Connectome Project (dHCP) demonstrate that SCG-ViT achieves superior classification ability compared to baseline models, and successfully identifies holistically different FC patterns among the three groups. Moreover, these different FCs are significantly correlated with the differential gene expressions of the three groups. In summary, SCG-ViT provides a powerfully brain-guided pipeline of adopting large-scale and data-intensive deep learning models for medical imaging-based diagnosis. Yuzhong Chen 0002, Zhenxiang Xiao, Yusong Sun, Jingchao Zhou, Weitong Guo, Chong Ma 0004, Lin Zhao 0004, Keith M. Kendrick, Benjamin Becker, Tianming Liu 0001, Xi Jiang 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2024 | Rectify ViT Shortcut Learning by Visual SaliencyabstractShortcut learning in deep learning models occurs when unintended features are prioritized, resulting in degenerated feature representations and reduced generalizability and interpretability. However, shortcut learning in the widely used vision transformer (ViT) framework is largely unknown. Meanwhile, introducing domain-specific knowledge is a major approach to rectifying the shortcuts that are predominated by background-related factors. For example, eye-gaze data from radiologists are effective human visual prior knowledge that has the great potential to guide the deep learning models to focus on meaningful foreground regions. However, obtaining eye-gaze data can still sometimes be time-consuming, labor-intensive, and even impractical. In this work, we propose a novel and effective saliency-guided ViT (SGT) model to rectify shortcut learning in ViT with the absence of eye-gaze data. Specifically, a computational visual saliency model (either pretrained or fine-tuned) is adopted to predict saliency maps for input image samples. Then, the saliency maps are used to filter the most informative image patches. Considering that this filter operation may lead to global information loss, we further introduce a residual connection that calculates the self-attention across all the image patches. The experiment results on natural and medical image datasets show that our SGT framework can effectively learn and leverage human prior knowledge without eye-gaze data and achieves much better performance than baselines. Meanwhile, it successfully rectifies the harmful shortcut learning and significantly improves the interpretability of the ViT model, demonstrating the promise of transferring human prior knowledge derived visual saliency in rectifying shortcut learning. Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Lei Guo 0002, Xintao Hu, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | FMRI-Guided Time-Symmetric Joint Model for Visual Attention PredictionabstractVisual attention prediction is linked to brain activity, cognition, and behavior. Despite the availability of brain activity features, previous studies have not fully utilized them, resulting in saliency maps predicted by models primarily based on image features that do not accurately reflect visual attention in the human brain. This inspires us to use functional Magnetic Resonance Imaging (fMRI) signals as a "brain observer" to supervise the training of developing models that integrate top-down image attention-dependent cues and supervise information from saliency maps generated from gaze movement patterns under natural stimuli. Hence, this paper presents an FMRI-Guided Time-Symmetric Joint Model to predict saliency maps from movie clips, which captures the dynamic aspects of human brain cognition and attention, enabling the combination of image features with brain features. Furthermore, we generalize the model to the MS-COCO challenge, evaluating its performance on non-movie data. Our model outperforms other brain-feature-free methods in focusing on visual attention regions of humans in both movie and non-movie datasets. Additionally, incorporating brain features improves model performance, indicating their ability to bridge the semantic gap between human cognition and visual images, allowing for more accurate capture of visual attention regions. Yaonai Wei, Chong Ma 0004, Tianyang Zhong, Lei Du 0001, Songyao Zhang, Tianming Liu 0001, Muheng Shang, Junwei Han 0001 |
BIBM | 2 |
| 2023 | Chat2Brain: A Method for Mapping Open-Ended Semantic Queries to Brain Activation MapsabstractOver decades, neuroscience has accumulated a wealth of research results in the text modality that can be used to explore cognitive processes. Meta-analysis is a typical method that successfully establishes a link from text queries to brain activation maps using these research results, but it still relies on an ideal query environment. In practical applications, text queries used for meta-analyses may encounter issues such as semantic redundancy and ambiguity, resulting in an inaccurate mapping to brain images. On the other hand, large language models (LLMs) like ChatGPT have shown great potential in tasks such as context understanding and reasoning, displaying a high degree of consistency with human natural language. Hence, LLMs could improve the connection between text modality and neuroscience, resolving existing challenges of meta-analyses. In this study, we propose a method called Chat2Brain that combines LLMs to basic text-2-image model, known as Text2Brain, to map open-ended semantic queries to brain activation maps in data-scarce and complex query environments. By utilizing the understanding and reasoning capabilities of LLMs, the performance of the mapping model is optimized by transferring text queries to semantic queries. We demonstrate that Chat2Brain can synthesize anatomically plausible neural activation patterns for more complex tasks of text queries. Yaonai Wei, Tianyang Zhong, Songyao Zhang, Xiao Li 0024, Lin Zhao 0004, Zhengliang Liu, Muheng Shang, Tianming Liu 0001, Chong Ma 0004, Lei Du 0001, Junwei Han 0001 |
BIBM | 11 |
| 2023 | Mammo-Net: Integrating Gaze Supervision and Interactive Information in Multi-view Mammogram Classification
Changkai Ji, Changde Du, Sheng Wang 0014, Chong Ma 0004, Jiaming Xie, Huiguang He, Dinggang Shen |
MICCAI (7) | 5 |
| 2023 | A Small-Sample Method with EEG Signals Based on Abductive Learning for Motor Imagery Decoding
Tianyang Zhong, Xiaozheng Wei, Enze Shi, Jiaxing Gao, Chong Ma 0004, Yaonai Wei, Songyao Zhang, Lei Guo 0002, Junwei Han 0001, Tianming Liu 0001 |
MICCAI (1) | 5 |
| 2023 | Eye-Gaze-Guided Vision Transformer for Rectifying Shortcut LearningabstractLearning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned representation. The situation becomes even more serious in medical image analysis, where the clinical data are limited and scarce while the reliability, generalizability and transparency of the learned model are highly required. To rectify the harmful shortcuts in medical imaging applications, in this paper, we propose a novel eye-gaze-guided vision transformer (EG-ViT) model which infuses the visual attention from radiologists to proactively guide the vision transformer (ViT) model to focus on regions with potential pathology rather than spurious correlations. To do so, the EG-ViT model takes the masked image patches that are within the radiologists' interest as input while has an additional residual connection to the last encoder layer to maintain the interactions of all patches. The experiments on two medical imaging datasets demonstrate that the proposed EG-ViT model can effectively rectify the harmful shortcut learning and improve the interpretability of the model. Meanwhile, infusing the experts' domain knowledge can also improve the large-scale ViT model's performance over all compared baseline methods with limited samples available. In general, EG-ViT takes the advantages of powerful deep neural networks while rectifies the harmful shortcut learning with human expert's prior knowledge. This work also opens new avenues for advancing current artificial intelligence paradigms by infusing human intelligence. Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Sheng Wang 0014, Lei Guo 0002, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 1 |