Chu-ran Wang

dblp:248/5699 · also Churan Wang · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-7699-0894ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Autoregressive Sequence Modeling for 3D Medical Image Representation
abstract
Three-dimensional (3D) medical images, such as Computed Tomography (CT) and Magnetic Resonance Imaging (MRI), are essential for clinical applications. However, the need for diverse and comprehensive representations is particularly pronounced when considering the variability across different organs, diagnostic tasks, and imaging modalities. How to effectively interpret the intricate contextual information and extract meaningful insights from these images remains an open challenge to the community. While current self-supervised learning methods have shown potential, they often consider an image as a whole thereby overlooking the extensive, complex relationships among local regions from one or multiple images. In this work, we introduce a pioneering method for learning 3D medical image representations through an autoregressive pre-training framework. Our approach sequences various 3D medical images based on spatial, contrast, and semantic correlations, treating them as interconnected visual tokens within a token sequence. By employing an autoregressive sequence modeling task, we predict the next visual token in the sequence, which allows our model to deeply understand and integrate the contextual information inherent in 3D medical images. Additionally, we implement a random startup strategy to avoid overestimating token relationships and to enhance the robustness of learning. The effectiveness of our approach is demonstrated by the superior performance over others on nine downstream tasks in public datasets.
Chu-ran Wang, Lixian Su, Fandong Zhang, Yizhou Wang 0001, Yizhou Yu
AAAI2
2025 UnrealZoo: Enriching Photo-Realistic Virtual Worlds for Embodied AI
abstract
We introduce UnrealZoo, a collection of over 100 photo-realistic 3D virtual worlds built on Unreal Engine, designed to reflect the complexity and variability of open-world environments. We also provide a rich variety of playable entities, including humans, animals, robots, and vehicles for embodied AI research. We extend UnrealCV with optimized APIs and tools for data collection, environment augmentation, distributed training, and benchmarking. These improvements achieve significant improvements in the efficiency of rendering and communication, enabling advanced applications such as multi-agent interactions. Our experimental evaluation across visual navigation and tracking tasks reveals two key insights: 1) environmental diversity provides substantial benefits for developing generalizable reinforcement learning (RL) agents, and 2) current embodied agents face persistent challenges in open-world scenarios, including navigation in unstructured terrain, adaptation to unseen morphologies, and managing latency in the close-loop control systems for interacting in highly dynamic objects. UnrealZoo thus serves as both a comprehensive testing ground and a pathway toward developing more capable embodied AI systems for real-world deployment.
Fangwei Zhong, Kui Wu 0007, Chu-ran Wang, Hao Chen 0062, Hai Ci, Zhoujun Li 0001, Yizhou Wang 0001
ICCV3
2025 VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
abstract
We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracking failure. Our approach combines the off-the-shelf active tracking methods with VLMs’ reasoning capabilities, deploying a fast visual policy for normal tracking and activating VLM reasoning only upon failure detection. The framework features a memory-augmented self-reflection mechanism that enables the VLM to progressively improve by learning from past experiences, effectively addressing VLMs’ limitations in 3D spatial reasoning. Experimental results demonstrate significant performance improvements, with our framework boosting success rates by 72% with state-of-the-art RL-based approaches and 220% with PID-based methods in challenging environments. This work represents the first integration of VLM-based reasoning to assist EVT agents in proactive failure recovery, offering substantial advances for real-world robotic applications that require continuous target monitoring in dynamic, unstructured environments. Project website: https://sites.google.com/view/evt-recovery-assistant.
Kui Wu 0007, Shuhang Xu, Hao Chen 0062, Chu-ran Wang, Zhoujun Li 0001, Yizhou Wang 0001, Fangwei Zhong
IROS4
2024 Empowering Embodied Visual Tracking with Visual Foundation Models and Offline RL
Fangwei Zhong, Kui Wu 0007, Hai Ci, Chu-ran Wang, Hao Chen 0062
ECCV (73)4
2024 Cross-dimensional Medical Self-supervised Representation Learning Based on a Pseudo-3D Transformation
Fandong Zhang, Yizhou Wang 0001, Chu-ran Wang, Yizhou Yu
MICCAI (11)6
2023 Learning Domain-Agnostic Representation for Disease Diagnosis
Chu-ran Wang, Jing Li 0091, Xinwei Sun 0001, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
ICLR1
2022 Disentangling Disease-related Representation from Obscure for Disease Prediction
abstract
Disease-related representations play a crucial role in image-based disease prediction such as cancer diagnosis, due to its considerable generalization capacity. However, it is still a challenge to identify lesion characteristics in obscured images, as many lesions are obscured by other tissues. In this paper, to learn the representations for identifying obscured lesions, we propose a disentanglement learning strategy under the guidance of alpha blending generation in an encoder-decoder framework (DAB-Net). Specifically, we take mammogram mass benign/malignant classification as an example. In our framework, composite obscured mass images are generated by alpha blending and then explicitly disentangled into disease-related mass features and interference glands features. To achieve disentanglement learning, features of these two parts are decoded to reconstruct the mass and the glands with corresponding reconstruction losses, and only disease-related mass features are fed into the classifier for disease prediction. Experimental results on one public dataset DDSM and three in-house datasets demonstrate that the proposed strategy can achieve state-of-the-art performance. DAB-Net achieves substantial improvements of 3.9%~4.4% AUC in obscured cases. Besides, the visualization analysis shows the model can better disentangle the mass and glands in the obscured image, suggesting the effectiveness of our solution in exploring the hidden characteristics in this challenging problem.
Chu-ran Wang, Fandong Zhang, Fangwei Zhong, Yizhou Yu, Yizhou Wang 0001
ICML1
2021 DAE-GCN: Identifying Disease-Related Features for Disease Prediction
Chu-ran Wang, Xinwei Sun 0001, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
MICCAI (5)1
2021 Bilateral Asymmetry Guided Counterfactual Generating Network for Mammogram Classification
abstract
Mammogram benign or malignant classification with only image-level labels is challenging due to the absence of lesion annotations. Motivated by the symmetric prior that the lesions on one side of breasts rarely appear in the corresponding areas on the other side, we explore to answer a counterfactual question to identify the lesion areas. This counterfactual question means: given an image with lesions, how would the features have behaved if there were no lesions in the image? To answer this question, we derive a new theoretical result based on the symmetric prior. Specifically, by building a causal model that entails such a prior for bilateral images, we identify to optimize the distances in distribution between i) the counterfactual features and the target side's features in lesion-free areas; and ii) the counterfactual features and the reference side's features in lesion areas. To realize these optimizations for better benign/malignant classification, we propose a counterfactual generative network, which is mainly composed of Generator Adversarial Network and a prediction feedback mechanism, they are optimized jointly and prompt each other. Specifically, the former can further improve the classi?cation performance by generating counterfactual features to calculate lesion areas. On the other hand, the latter helps counterfactual generation by the supervision of classification loss. The utility of our method and the effectiveness of each module in our model can be verified by state-of-the-art performance on INBreast and an in-house dataset and ablation studies.
Chu-ran Wang, Jing Li 0091, Fandong Zhang, Xinwei Sun 0001, Hao Dong 0003, Yizhou Yu, Yizhou Wang 0001
IEEE Trans. Image Process.1
2020 BR-GAN: Bilateral Residual Generating Adversarial Network for Mammogram Classification
Chu-ran Wang, Fandong Zhang, Yizhou Yu, Yizhou Wang 0001
MICCAI (2)1
2019 DDRM-CapsNet: Capsule Network Based on Deep Dynamic Routing Mechanism for Complex Data
Jianwei Liu 0006, Runkun Lu, Yuanfeng Lian, Dianzhong Wang, Xionglin Luo, Chu-ran Wang
ICANN (1)7
2019 FSC-CapsNet: Fractionally-Strided Convolutional Capsule Network for complex data
abstract
Recently, a novel neural network called CapsNet has attracted the attention of many researchers. It is a great attempt to overcome the drawback of convolutional neural networks (CNNs) and achieves state-of-the-art performance on some simple datasets like MNIST. However, this network architecture is built specifically for MNIST and gets a poor performance on more complex datasets like CIFAR-10. To address this problem, aiming at complex data, we propose a new CapsNet architecture called Fractionally-Strided Convolutional Capsule Network (FSC-CapsNet). We modify both network structures of the encoder and decoder of CapsNet. For the purpose of extracting better features, we increase the number of convolutional layers before capsule layer in the encoder and improve the reconstruction performance by adopting two fractionally-strided convolutional layers in the decoder. In addition, no pooling layers are used in our architecture. To assess the performance of our proposed network on complex data, we conduct experiments with a single model without using any ensembled methods and data augmentation techniques on five real-world datasets, which are of higher dimensionality and larger size than MNIST. The experimental results demonstrate that our proposed method achieves better performance and improves the reconstruction performance compared with the normal CapsNet.
Jianwei Liu 0006, Runkun Lu, Yuanfeng Lian, Dianzhong Wang, Xionglin Luo, Chu-ran Wang
IJCNN7