VLDB 2026 Research / reviewers in the wild / expert
Chaoyi Zhang
dblp:116/7852
· DBLP profile ↗
45ranked-venue papers
7as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 13 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM DecodingabstractExisting vision-language models (VLMs) often suffer from visual hallucination, where the generated responses contain inaccuracies that are not grounded in the visual input. Efforts to address this issue without model finetuning primarily mitigate hallucination by contrastively reducing language biases or amplifying the weights of visual embedding during decoding. However, these approaches remain limited in their ability to capture fine-grained visual details. In this work, we propose the Perception Magnifier (PM), a novel visual decoding method that iteratively isolates relevant visual tokens based on attention and magnifies the corresponding regions, spurring the model to concentrate on fine-grained visual details during decoding. By magnifying critical regions while preserving the structural and contextual information at each decoding step, PM allows the VLM to enhance its scrutiny of the visual input, hence producing more accurate and faithful responses. Extensive experimental results demonstrate that PM not only achieves superior hallucination mitigation but also enhances language generation while preserving strong reasoning capabilities. Code can be found at https://github.com/ShunqiM/PM. Shunqi Mao, Chaoyi Zhang, Tom Weidong Cai |
ACL (1) | 2 |
| 2025 | RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation ParadigmabstractAfter pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of multimodal interleaved documents remains underutilized for contrastive vision-language representation learning. To fully leverage these unpaired documents, we initially establish a Real-World Data Extraction pipeline to extract high-quality images and texts. Then we design a hierarchical retrieval method to efficiently associate each image with multiple semantically relevant realistic texts. To further enhance fine-grained visual information, we propose an image semantic augmented generation module for synthetic text production. Furthermore, we employ a semantic balance sampling strategy to improve dataset diversity, enabling better learning of long-tail concepts. Based on these innovations, we construct RealSyn, a dataset combining realistic and synthetic texts, available in three scales: 15M, 30M, and 100M. We compare our dataset with other widely used datasets of equivalent scale for CLIP training. Models pre-trained on RealSyn consistently achieve state-of-the-art performance across various downstream tasks, including linear probe, zero-shot transfer, zero-shot robustness, and zero-shot retrieval. Furthermore, extensive experiments confirm that RealSyn significantly enhances contrastive vision-language representation learning and demonstrates robust scalability. The code will be released in https://garygutc.github.io/RealSyn. Tiancheng Gu, Kaicheng Yang 0002, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Tom Weidong Cai, Jiankang Deng |
ACM Multimedia | 3 |
| 2025 | Model Merging in Pre-training of Large Language ModelsabstractModel merging has emerged as a promising technique for enhancing large language models, though its application in large-scale pre-training remains relatively unexplored. In this paper, we present a comprehensive investigation of model merging techniques during the pre-training process. Through extensive experiments with both dense and Mixture-of-Experts (MoE) architectures ranging from millions to over 100 billion parameters, we demonstrate that merging checkpoints trained with constant learning rates not only achieves significant performance improvements but also enables accurate prediction of annealing behavior. These improvements lead to both more efficient model development and significantly lower training costs. Our detailed ablation studies on merging strategies and hyperparameters provide new insights into the underlying mechanisms while uncovering novel applications. Through comprehensive experimental analysis, we offer the open-source community practical pre-training guidelines for effective model merging. Yunshui Li, Yiyuan Ma, Chaoyi Zhang, Jianqiao Lu, Ziwen Xu, Mengzhao Chen, Minrui Wang, Shiyi Zhan, Xunhao Lai, Yao Luo, Xingyan Bin, Hongbin Ren, Mingji Han, Wenhao Hao, Bairen Yi, LingJun Liu, Bole Ma, Xiaoying Jia 0005 |
NeurIPS | 4 |
| 2025 | TractGraphFormer: Anatomically informed hybrid graph CNN-transformer network for interpretable sex and age prediction from diffusion MRI tractography
Yuqian Chen, Fan Zhang 0013, Leo R. Zekelman, Suheyla Cetin Karayumak, Tengfei Xue, Chaoyi Zhang, Yang Song 0001, Jarrett Rushmore, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Lauren O'Donnell |
Medical Image Anal. | 7 |
| 2025 | A Physical Human-Robot Interaction Framework for Trajectory Adaptation Based on Human Motion Prediction and Adaptive Impedance ControlabstractPhysical human-robot interaction (pHRI) plays an important role in robotic. In order for a human operator to be able to easily adapt to interact with a robot, a minimal interaction force in pHRI should be achieved. In this paper, a pHRI framework is proposed to allow the robot to regulate its trajectory adaptively for minimizing the interaction force with small position-tracking errors. The trajectory of the robot is first adjusted by the interaction force which is updated by the performance evaluation index. Then, the human hand motion is predicted based on the autoregressive (AR) model to further adapt the trajectory. Thirdly, an adaptive impedance control method is developed to update the stiffness in the robot impedance controller using surface electromyography (sEMG) signals for robot compliant interaction with the environment. This method allows the human operator to interact with the robot by the interaction force, the hand motion and muscle contraction. By investigating the performance of the proposed method, the interaction force is decreased and a good position tracking accuracy is achieved. Comparative experiments demonstrate the enhanced performance of the proposed method. Note to Practitioners—This paper focuses on developing a novel method that can allow the robot to compliantly interact with the human operator while simultaneously taking into account the trajectory-tracking accuracy and the interaction force in pHRI scenarios. The proposed method has a large application potential in a variety of pHRI tasks, such as human-robot collaborative transporting, curing, assembly, cutting, and so on. In addition, the proposed method can allow the human operator to physically interact with the robot in an easier and more intuitive manner, by taking advantage of human motion prediction and adaptive impedance control. Therefore, it is also potentially utilized for rehabilitation and assistive robots, and robot learning skills from human physical demonstration. Jing Luo 0005, Chaoyi Zhang, Weiyong Si, Yiming Jiang 0001, Chenguang Yang 0001, Chao Zeng 0002 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Enhancing Robustness to Noise Corruption for Point Cloud Recognition via Spatial Sorting and Set-Mixing Aggregation Module
Dingxin Zhang 0001, Jianhui Yu, Tengfei Xue, Chaoyi Zhang, Dongnan Liu, Tom Weidong Cai |
ACCV (9) | 4 |
| 2024 | MM-Narrator: Narrating Long-form Videos with Multimodal In-Context LearningabstractWe present MM-Narrator, a novel system leveraging GPT-4 with multimodal in-context learning for the gener-ation of audio descriptions (AD). Unlike previous methods that primarily focused on downstream fine-tuning with short video clips, MM-Narrator excels in generating precise audio descriptions for videos of extensive lengths, even be-yond hours, in an autoregressive manner. This capability is made possible by the proposed memory-augmented generation process, which effectively utilizes both the short-term textual context and long-term visual memory through an efficient register-and-recall mechanism. These contextual memories compile pertinent past information, including storylines and character identities, ensuring an accurate tracking and depicting of story-coherent and character-centric audio descriptions. Maintaining the training-free design of MM-Narrator, we further propose a complexity-based demonstration selection strategy to largely enhance its multi-step reasoning capability via few-shot multimodal in-context learning (MM-ICL). Experimental results on MAD-eval dataset demonstrate that MM-Narrator consistently outperforms both the existing fine-tuning-based approaches and LLM-based approaches in most scenarios, as measured by standard evaluation metrics. Additionally, we introduce the first segment-based evaluator for recurrent text generation. Empowered by GPT-4, this evaluator comprehensively reasons and marks AD generation performance in various extendable dimensions. Chaoyi Zhang, Zhengyuan Yang, Chung-Ching Lin, Zicheng Liu 0001 |
CVPR | 1 |
| 2024 | Controllable Contextualized Image Captioning: Directing the Visual Narrative Through User-Defined Highlights
Shunqi Mao, Chaoyi Zhang, Hwanjun Song, Igor Shalyminov, Tom Weidong Cai |
ECCV (50) | 2 |
| 2024 | Enhancing Advanced Visual Reasoning Ability of Large Language ModelsabstractRecent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models’ advanced reasoning ability. Traditional Vision-Language models (VLMs) perform well in visual perception tasks while struggling with complex reasoning scenarios. Conversely, Large Language Models (LLMs) demonstrate robust text reasoning capabilities; however, they lack visual acuity. To bridge this gap, we propose Complex Visual Reasoning Large Language Models (CVR-LLM), capitalizing on VLMs’ visual perception proficiency and LLMs’ extensive reasoning capability. Unlike recent multimodal large language models (MLLMs) that require a projection layer, our approach transforms images into detailed, context-aware descriptions using an iterative self-refinement loop and leverages LLMs’ text knowledge for accurate predictions without extra training. We also introduce a novel multi-modal in-context learning (ICL) methodology to enhance LLMs’ contextual understanding and reasoning. Additionally, we introduce Chain-of-Comparison (CoC), a step-by-step comparison technique enabling contrasting various aspects of predictions. Our CVR-LLM presents the first comprehensive study across a wide array of complex visual reasoning tasks and achieves SOTA performance among all. Dongnan Liu, Chaoyi Zhang, Heng Wang 0007, Tengfei Xue, Tom Weidong Cai |
EMNLP | 3 |
| 2024 | Learning to Synthesize Graphics Programs for Geometric Artworks
Qi Bing, Chaoyi Zhang, Tom Weidong Cai |
ICPR (18) | 2 |
| 2024 | Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
Dongnan Liu, Heng Wang 0007, Chaoyi Zhang, Tom Weidong Cai |
MMAsia | 4 |
| 2024 | Predicting Channel Delay State Information in 5G-TSN Systems Using Extreme Learning Machine Autoencoder (ELM-AE) Model Based on Intelligent Deep Extreme Learning Machine (DELM)abstractThis paper investigates the joint scheduling of cross-channel traffic resources in 5G-Time-Sensitive Networks (TSN) and proposes an adaptive prediction method suitable for cross-domain Channel State Information (CSI). Firstly, we analyze the 5G-TSN cross-domain data forwarding mechanism by leveraging the architecture of the 5G-TSN bridging network and combining the functions of 5G and TSN network elements. Secondly, we propose a representation method for the 5G-TSN cross-network wireless CSI, specifically the data transmission delay information, as a dataset for channel quality prediction. This serves as a data foundation for subsequent intelligent prediction. Next, to make better use of the local information of channel state and achieve fast convergence, we employ an Extreme Learning Machine Auto-Encoder (ELM-AE) prediction logic based on Deep Extreme Learning Machine (DELM) and introduce the Dung Beetle Optimizer (DBO) algorithm to improve the DELM regression prediction. We perform prediction and analysis of the 5G channel delay and TSN domain data transmission delay. Then, we use the 5G-TSN CSI, collected in practice, as the data source to train and test the wireless channel delay indicators, which helps form the 5G-TSN channel model. Finally, we build a laboratory transmission prototype test bed for 5G-TSN cross-network transmission and conduct end-to-end transmission delay testing based on the proposed offline-generated channel model. The results demonstrate that the channel prediction model enables the end-to-end delay to decrease to less than 5 ms, the cross-network time synchronization accuracy to reduce to less than 100 ns, and the relevant performance indicators to reach industry-leading levels. Chaoyi Zhang, Jianquan Wang 0001, Meixia Fu |
IEEE Internet Things J. | 1 |
| 2024 | TractGeoNet: A geometric deep learning framework for pointwise analysis of tract microstructure to predict language assessment performance
Yuqian Chen, Leo R. Zekelman, Chaoyi Zhang, Tengfei Xue, Yang Song 0001, Nikos Makris, Yogesh Rathi, Alexandra J. Golby, Tom Weidong Cai, Fan Zhang 0013, Lauren O'Donnell |
Medical Image Anal. | 3 |
| 2024 | Exploiting Structural Consistency of Chest Anatomy for Unsupervised Anomaly Detection in Radiography ImagesabstractRadiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. Exploiting this structured information could potentially ease the detection of anomalies from radiography images. To this end, we propose a Simple Space-Aware Memory Matrix for In-painting and Detecting anomalies from radiography images (abbreviated as SimSID). We formulate anomaly detection as an image reconstruction task, consisting of a space-aware memory matrix and an in-painting block in the feature space. During the training, SimSID can taxonomize the ingrained anatomical structures into recurrent visual patterns, and in the inference, it can identify anomalies (unseen/modified visual patterns) from the test image. Our SimSID surpasses the state of the arts in unsupervised anomaly detection by +8.0%, +5.0%, and +9.9% AUC scores on ZhangLab, COVIDx, and CheXpert benchmark datasets, respectively. Tiange Xiang, Yixiao Zhang 0001, Yongyi Lu, Alan L. Yuille, Chaoyi Zhang, Tom Weidong Cai, Zongwei Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Self-Supervised Deep Unrolled Reconstruction Using Regularization by DenoisingabstractDeep learning methods have been successfully used in various computer vision tasks. Inspired by that success, deep learning has been explored in magnetic resonance imaging (MRI) reconstruction. In particular, integrating deep learning and model-based optimization methods has shown considerable advantages. However, a large amount of labeled training data is typically needed for high reconstruction quality, which is challenging for some MRI applications. In this paper, we propose a novel reconstruction method, named DURED-Net, that enables interpretable self-supervised learning for MR image reconstruction by combining a self-supervised denoising network and a plug-and-play method. We aim to boost the reconstruction performance of Noise2Noise in MR reconstruction by adding an explicit prior that utilizes imaging physics. Specifically, the leverage of a denoising network for MRI reconstruction is achieved using Regularization by Denoising (RED). Experiment results demonstrate that the proposed method requires a reduced amount of training data to achieve high reconstruction quality among the state-of-the-art approaches utilizing Noise2Noise. Peizhou Huang, Chaoyi Zhang, Xiaoliang Zhang 0001, Dong Liang 0001, Leslie Ying |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Rethinking Rotation Invariance with Point Cloud RegistrationabstractRecent investigations on rotation invariance for 3D point clouds have been devoted to devising rotation-invariant feature descriptors or learning canonical spaces where objects are semantically aligned. Examinations of learning frameworks for invariance have seldom been looked into. In this work, we review rotation invariance (RI) in terms of point cloud registration (PCR) and propose an effective framework for rotation invariance learning via three sequential stages, namely rotation-invariant shape encoding, aligned feature integration, and deep feature registration. We first encode shape descriptors constructed with respect to reference frames defined over different scales, e.g., local patches and global topology, to generate rotation-invariant latent shape codes. Within the integration stage, we propose an Aligned Integration Transformer (AIT) to produce a discriminative feature representation by integrating point-wise self- and cross-relations established within the shape codes. Meanwhile, we adopt rigid transformations between reference frames to align the shape codes for feature consistency across different scales. Finally, the deep integrated feature is registered to both rotation-invariant shape codes to maximize their feature similarities, such that rotation invariance of the integrated feature is preserved and shared semantic information is implicitly extracted from shape codes. Experimental results on 3D shape classification, part segmentation, and retrieval tasks prove the feasibility of our framework. Our project page is released at: https://rotation3d.github.io/. Jianhui Yu, Chaoyi Zhang, Tom Weidong Cai |
AAAI | 2 |
| 2023 | PaRot: Patch-Wise Rotation-Invariant Network via Feature Disentanglement and Pose RestorationabstractRecent interest in point cloud analysis has led rapid progress in designing deep learning methods for 3D models. However, state-of-the-art models are not robust to rotations, which remains an unknown prior to real applications and harms the model performance. In this work, we introduce a novel Patch-wise Rotation-invariant network (PaRot), which achieves rotation invariance via feature disentanglement and produces consistent predictions for samples with arbitrary rotations. Specifically, we design a siamese training module which disentangles rotation invariance and equivariance from patches defined over different scales, e.g., the local geometry and global shape, via a pair of rotations. However, our disentangled invariant feature loses the intrinsic pose information of each patch. To solve this problem, we propose a rotation-invariant geometric relation to restore the relative pose with equivariant information for patches defined over different scales. Utilising the pose information, we propose a hierarchical module which implements intra-scale and inter-scale feature aggregation for 3D shape learning. Moreover, we introduce a pose-aware feature propagation process with the rotation-invariant relative pose information embedded. Experiments show that our disentanglement module extracts high-quality rotation-robust features and the proposed lightweight model achieves competitive results in rotated 3D object classification and part segmentation tasks. Dingxin Zhang 0001, Jianhui Yu, Chaoyi Zhang, Tom Weidong Cai |
AAAI | 3 |
| 2023 | SQUID: Deep Feature In-Painting for Unsupervised Anomaly DetectionabstractRadiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. To exploit this structured information, we propose the use of Space-aware Memory Queues for In-painting and Detecting anomalies from radiography images (abbreviated as SQUID). We show that SQUID can taxonomize the ingrained anatomical structures into recurrent patterns; and in the inference, it can identify anomalies (unseen/modified patterns) in the image. SQUID surpasses 13 state-of-the-art methods in unsupervised anomaly detection by at least 5 points on two chest X-ray benchmark datasets measured by the Area Under the Curve (AUC). Additionally, we have created a new dataset (DigitAnatomy), which synthesizes the spatial correlation and consistent shape in chest anatomy. We hope DigitAnatomy can prompt the development, evaluation, and interpretability of anomaly detection methods. Tiange Xiang, Yixiao Zhang 0001, Yongyi Lu, Alan L. Yuille, Chaoyi Zhang, Tom Weidong Cai, Zongwei Zhou |
CVPR | 5 |
| 2023 | TractCloud: Registration-Free Tractography Parcellation with a Novel Local-Global Streamline Point Cloud Representation
Tengfei Xue, Yuqian Chen, Chaoyi Zhang, Alexandra J. Golby, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Fan Zhang 0013, Lauren O'Donnell |
MICCAI (8) | 3 |
| 2023 | PointNeuron: 3D Neuron Reconstruction via Geometry and Topology Learning of Point CloudsabstractDigital neuron reconstruction from 3D microscopy images is an essential technique for investigating brain connectomics and neuron morphology. Existing reconstruction frameworks use convolution-based segmentation networks to partition the neuron from noisy backgrounds before applying the tracing algorithm. The tracing results are sensitive to the raw image quality and segmentation accuracy. In this paper, we propose a novel framework for 3D neuron reconstruction. Our key idea is to use the geometric representation power of the point cloud to better explore the intrinsic structural information of neurons. Our proposed framework adopts one graph convolutional network to predict the neural skeleton points and another one to produce the connectivity of these points. We finally generate the target SWC file through the interpretation of the predicted point coordinates, radius, and connections. Evaluated on the Janelia-Fly dataset from the BigNeuron project, we show that our framework achieves competitive neuron reconstruction performance. Our geometry and topology learning of point clouds could further benefit 3D medical image analysis, such as cardiac surface reconstruction. Our code is available at https://github.com/RunkaiZhao/PointNeuron. Runkai Zhao, Heng Wang 0007, Chaoyi Zhang, Tom Weidong Cai |
WACV | 3 |
| 2023 | Region-based fully convolutional networks with deformable convolution and attention fusion for steel surface defect detection in industrial Internet of ThingsabstractAbstract Next‐generation 6G networks will fully drive the development of the industrial Internet of Things. Steel surface defect detection as an important application in industrial Internet of Things has recently received increasing attention from the military industry, the aviation industry and other fields, which is closely related to the quality of industrial production products. However, many typical convolutional neural networks‐based methods are insensitive to the problem of unclear boundaries. In this article, the authors develop a region‐based fully convolutional networks with deformable convolution and attention fusion to adaptively learn salient features for steel surface defect detection. Specifically, deformable convolution is applied into selectively replace the standard convolution in the backbone of the region‐based fully convolutional networks, which performs significantly in scenarios with unclear defect boundaries. Moreover, convolutional block attention module is utilised in region proposal network to further enhance detection accuracy. The proposed architecture is demonstrated on two popular steel defect detection benchmarks, including NEU‐DET and GC10‐DET, which can effectively present the performance of steel surface defect detection by abundant experiments. The mean average precision on two datasets reaches 80.9% and 66.2%. The average precision of defect crazing, inclusion, patches, pitted‐surface, rolled‐in scale and scratches on NEU‐DET is 58.2%, 82.3%, 95.7%, 85.6%, 75.9%, and 87.9% respectively. Meixia Fu, Qu Wang, Lei Sun 0012, Zhangchao Ma, Chaoyi Zhang, Wanqing Guan, Wei Li 0037, Na Chen 0004, Danshi Wang, Jianquan Wang 0001 |
IET Signal Process. | 6 |
| 2023 | Superficial white matter analysis: An efficient point-cloud-based deep learning framework with supervised contrastive learning for consistent tractography parcellation across populations and dMRI acquisitions
Tengfei Xue, Fan Zhang 0013, Chaoyi Zhang, Yuqian Chen, Yang Song 0001, Alexandra J. Golby, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Lauren O'Donnell |
Medical Image Anal. | 3 |
| 2023 | Decompose to Adapt: Cross-Domain Object Detection Via Feature DisentanglementabstractRecent advances in unsupervised domain adaptation (UDA) techniques have witnessed great success in cross-domain computer vision tasks, enhancing the generalization ability of data-driven deep learning architectures by bridging the domain distribution gaps. For the UDA-based cross-domain object detection methods, the majority of them alleviate the domain bias by inducing the domain-invariant feature generation via adversarial learning strategy. However, their domain discriminators have limited classification ability due to the unstable adversarial training process. Therefore, the extracted features induced by them cannot be perfectly domain-invariant and still contain domain-private factors, bringing obstacles to further alleviate the cross-domain discrepancy. To tackle this issue, we design a Domain Disentanglement Faster-RCNN (DDF) to eliminate the source-specific information in the features for detection task learning. Our DDF method facilitates the feature disentanglement at the global and local stages, with a Global Triplet Disentanglement (GTD) module and an Instance Similarity Disentanglement (ISD) module, respectively. By outperforming state-of-the-art methods on four benchmark UDA object detection tasks, our DDF method is demonstrated to be effective with wide applicability. Dongnan Liu, Chaoyi Zhang, Yang Song 0001, Heng Huang 0001, Chenyu Wang 0001, Michael Barnett 0006, Tom Weidong Cai |
IEEE Trans. Multim. | 2 |
| 2022 | Unsupervised Domain Adaptive Fundus Image Segmentation with Few Labeled Source Data
Qianbi Yu, Dongnan Liu, Chaoyi Zhang, Xinwen Zhang, Tom Weidong Cai |
BMVC | 3 |
| 2022 | Channel-Position Self-Attention with Query Refinement Skeleton Graph Neural Network in Human Pose EstimationabstractHuman Pose Estimation (HPE) is a long-standing yet challenging task in computer vision. The nature of the problem requires comprehensive global contextual reasoning among joints in different locations. In this work, we explore how to incorporate two popular and effective concepts, self-attention and Graph Neural Network (GNN), to model long-range information in HPE. Three different ways to implement self-attention in 3D feature maps are studied, where the best result is achieved via the channel-position version. Accuracy is further improved by refining the queries via an efficient channel-wise parallel GNN that explicitly models the human joint graphical relationships. We are able to improve prediction accuracy on strong baseline models and achieve state-of-the-art results. Shek Wai Chu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
ICIP | 2 |
| 2022 | Spatiality-guided Transformer for 3D Dense Captioning on Point CloudsabstractDense captioning in 3D point clouds is an emerging vision-and-language task involving object-level 3D scene understanding. Apart from coarse semantic class prediction and bounding box regression as in traditional 3D object detection, 3D dense captioning aims at producing a further and finer instance-level label of natural language description on visual appearance and spatial relations for each scene object of interest. To detect and describe objects in a scene, following the spirit of neural machine translation, we propose a transformer-based encoder-decoder architecture, namely SpaCap3D, to transform objects into descriptions, where we especially investigate the relative spatiality of objects in 3D scenes and design a spatiality-guided encoder via a token-to-token spatial relation learning objective and an object-centric decoder for precise and spatiality-enhanced object caption generation. Evaluated on two benchmark datasets, ScanRefer and ReferIt3D, our proposed SpaCap3D outperforms the baseline method Scan2Cap by 4.94% and 9.61% in [email protected], respectively. Our project page with source code and supplementary files is available at https://SpaCap3D.github.io/. Heng Wang 0007, Chaoyi Zhang, Jianhui Yu, Tom Weidong Cai |
IJCAI | 2 |
| 2022 | White Matter Tracts are Point Clouds: Neuropsychological Score Prediction and Critical Region Localization via Geometric Deep Learning
Yuqian Chen, Fan Zhang 0013, Chaoyi Zhang, Tengfei Xue, Leo R. Zekelman, Jianzhong He 0001, Yang Song 0001, Nikos Makris, Yogesh Rathi, Alexandra J. Golby, Tom Weidong Cai, Lauren O'Donnell |
MICCAI (1) | 3 |
| 2022 | Towards bi-directional skip connections in encoder-decoder architectures and beyond
Tiange Xiang, Chaoyi Zhang, Xinyi Wang 0015, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai |
Medical Image Anal. | 2 |
| 2022 | Multiple Sclerosis Lesion Analysis in Brain Magnetic Resonance Images: Techniques and Clinical ApplicationsabstractMultiple sclerosis (MS) is a chronic inflammatory and degenerative disease of the central nervous system, characterized by the appearance of focal lesions in the white and gray matter that topographically correlate with an individual patient's neurological symptoms and signs. Magnetic resonance imaging (MRI) provides detailed in-vivo structural information, permitting the quantification and categorization of MS lesions that critically inform disease management. Traditionally, MS lesions have been manually annotated on 2D MRI slices, a process that is inefficient and prone to inter-/intra-observer errors. Recently, automated statistical imaging analysis techniques have been proposed to detect and segment MS lesions based on MRI voxel intensity. However, their effectiveness is limited by the heterogeneity of both MRI data acquisition techniques and the appearance of MS lesions. By learning complex lesion representations directly from images, deep learning techniques have achieved remarkable breakthroughs in the MS lesion segmentation task. Here, we provide a comprehensive review of state-of-the-art automatic statistical and deep-learning MS segmentation methods and discuss current and future clinical applications. Further, we review technical strategies, such as domain adaptation, to enhance MS lesion segmentation in real-world clinical settings. Chaoyi Zhang, Mariano Cabezas, Yang Song 0001, Zihao Tang 0002, Dongnan Liu, Tom Weidong Cai, Michael Barnett 0006, Chenyu Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | DSNet: A Dual-Stream Framework for Weakly-Supervised Gigapixel Pathology Image AnalysisabstractWe present a novel weakly-supervised framework for classifying whole slide images (WSIs). WSIs, due to their gigapixel resolution, are commonly processed by patch-wise classification with patch-level labels. However, patch-level labels require precise annotations, which is expensive and usually unavailable on clinical data. With image-level labels only, patch-wise classification would be sub-optimal due to inconsistency between the patch appearance and image-level label. To address this issue, we posit that WSI analysis can be effectively conducted by integrating information at both high magnification (local) and low magnification (regional) levels. We auto-encode the visual signals in each patch into a latent embedding vector representing local information, and down-sample the raw WSI to hardware-acceptable thumbnails representing regional information. The WSI label is then predicted with a Dual-Stream Network (DSNet), which takes the transformed local patch embeddings and multi-scale thumbnail images as inputs and can be trained by the image-level label only. Experiments conducted on three large-scale public datasets demonstrate that our method outperforms all recent state-of-the-art weakly-supervised WSI classification methods. Tiange Xiang, Yang Song 0001, Chaoyi Zhang, Dongnan Liu, Fan Zhang 0013, Heng Huang 0001, Lauren O'Donnell, Tom Weidong Cai |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Exploiting Edge-Oriented Reasoning for 3D Point-Based Scene Graph AnalysisabstractScene understanding is a critical problem in computer vision. In this paper, we propose a 3D point-based scene graph generation (SGGpoint) framework to effectively bridge perception and reasoning to achieve scene under-standing via three sequential stages, namely scene graph construction, reasoning, and inference. Within the reasoning stage, an EDGE-oriented Graph Convolutional Network (EdgeGCN) is created to exploit multi-dimensional edge features for explicit relationship modeling, together with the exploration of two associated twinning interaction mechanisms between nodes and edges for the independent evolution of scene graph representations. Overall, our integrated SGGpointframework is established to seek and infer scene structures of interest from both real-world and synthetic 3D point-based scenes. Our experimental results show promising edge-oriented reasoning effects on scene graph generation studies. We also demonstrate our method advantage on several traditional graph representation learning benchmark datasets, including the node-wise classification on citation networks and whole-graph recognition problems for molecular analysis. Chaoyi Zhang, Jianhui Yu, Yang Song 0001, Tom Weidong Cai |
CVPR | 1 |
| 2021 | Walk in the Cloud: Learning Curves for Point Clouds Shape AnalysisabstractDiscrete point cloud objects lack sufficient shape descriptors of 3D geometries. In this paper, we present a novel method for aggregating hypothetical curves in point clouds. Sequences of connected points (curves) are initially grouped by taking guided walks in the point clouds, and then subsequently aggregated back to augment their pointwise features. We provide an effective implementation of the proposed aggregation strategy including a novel curve grouping operator followed by a curve aggregation operator. Our method was benchmarked on several point cloud analysis tasks where we achieved the state-of-the-art classification accuracy of 94.2% on the ModelNet40 classification task, instance IoU of 86.8% on the ShapeNetPart segmentation task and cosine error of 0.11 on the ModelNet40 normal estimation task. Our project page with source code is available at: https://curvenet.github.io/. Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Jianhui Yu, Tom Weidong Cai |
ICCV | 2 |
| 2021 | Iterative Subnetwork With Linear Hierarchical Ordering for Human Pose EstimationabstractHuman pose estimation is a long-standing and challenging problem in computer vision. Many recent advancements in the field have relied on complex structure refinement and specific human joint graphical relations. However, progress has been saturated in terms of accuracy. Each time, new state-of-the-art approaches only improve accuracy by less than 0.3% in the MPII test set despite using complicated model structures. Most recent developments can be summarized into two main ideas: 1) refinement subnetwork to improve predictions iteratively and 2) exploitation of human joint graphical relations. In this work, we present how efficient and simple iterative subnetworks with linear hierarchical ordering based on the aforementioned ideas can help to improve accuracy on strong backbone models. Different versions of iterative subnetwork are examined. Significant improvements on difficult body part predictions such as wrists and ankles using simple convolution subnetwork are observed. Further improvements can be made by using a large receptive field subnetwork such as axial-transformer [1]. Shek Wai Chu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
ICIP | 2 |
| 2021 | ICE-GAN: Identity-Aware and Capsule-Enhanced GAN with Graph-Based Reasoning for Micro-Expression Recognition and SynthesisabstractMicro-expressions are reflections of people's true feelings and motives, which attract an increasing number of researchers into the study of automatic facial micro-expression recognition. The short detection window, the subtle facial muscle movements, and the limited training samples make micro-expression recognition challenging. To this end, we propose a novel Identity-aware and Capsule-Enhanced Generative Adversarial Network with graph-based reasoning (ICE-GAN), introducing micro-expression synthesis as an auxiliary task to assist recognition. The generator produces synthetic faces with controllable micro-expressions and identity-aware features, whose long-ranged dependencies are captured through the graph reasoning module (GRM), and the discriminator detects the image authenticity and expression classes. Our ICE-GAN was evaluated on Micro-Expression Grand Challenge 2019 (MEGC2019) with a significant improvement (12.9%) over the winner and surpassed other state-of-the-art methods. Jianhui Yu, Chaoyi Zhang, Yang Song 0001, Tom Weidong Cai |
IJCNN | 2 |
| 2021 | Deep Fiber Clustering: Anatomically Informed Unsupervised Deep Learning for Fast and Effective White Matter Parcellation
Yuqian Chen, Chaoyi Zhang, Yang Song 0001, Nikos Makris, Yogesh Rathi, Tom Weidong Cai, Fan Zhang 0013, Lauren O'Donnell |
MICCAI (7) | 2 |
| 2021 | BiX-NAS: Searching Efficient Bi-directional Architecture for Medical Image Segmentation
Xinyi Wang 0015, Tiange Xiang, Chaoyi Zhang, Yang Song 0001, Dongnan Liu, Heng Huang 0001, Tom Weidong Cai |
MICCAI (1) | 3 |
| 2021 | Divergence-Free Fitting-Based Incompressible Deformation Quantification of LiverabstractLiver is an incompressible organ that maintains its volume during the respiration-induced deformation. Quantifying this deformation with the incompressible constraint is significant for liver tracking. The constraint can be accomplished with retaining the divergence-free field obtained by the deformation decomposition. However, the decomposition process is time-consuming, and the removal of non-divergence-free field weakens the deformation. In this study, a divergence-free fitting-based registration method is proposed to quantify the incompressible deformation rapidly and accurately. First, the deformation to be estimated is mapped to the velocity in a diffeomorphic space. Then, this velocity is decomposed by a fast Fourier-based Hodge-Helmholtz decomposition to obtain the divergence-free, curl-free, and harmonic fields. The curl-free field is replaced and fitted by the obtained harmonic field with a translation field to generate a new divergence-free velocity. By optimizing this velocity, the final incompressible deformation is obtained. Moreover, a deep learning framework (DLF) is constructed to accelerate the incompressible deformation quantification. An incompressible respiratory motion model is built for the DLF by using the proposed registration method and is then used to augment the training data. An encoder-decoder network is introduced to learn appearance-velocity correlation at patch scale. In the experiment, we compare the proposed registration with three state-of-the-art methods. The results show that the proposed method can accurately achieve the incompressible registration of liver with a mean liver overlap ratio of 95.33%. Moreover, the time consumed by DLF is nearly 15 times shorter than that by other methods. Tianyu Fu 0003, Jingfan Fan, Dingkun Liu, Hong Song 0003, Chaoyi Zhang, Danni Ai, Zhigang Cheng, Jian Yang 0009 |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Shape-Oriented Convolution Neural Network for Point Cloud AnalysisabstractPoint cloud is a principal data structure adopted for 3D geometric information encoding. Unlike other conventional visual data, such as images and videos, these irregular points describe the complex shape features of 3D objects, which makes shape feature learning an essential component of point cloud analysis. To this end, a shape-oriented message passing scheme dubbed ShapeConv is proposed to focus on the representation learning of the underlying shape formed by each local neighboring point. Despite this intra-shape relationship learning, ShapeConv is also designed to incorporate the contextual effects from the inter-shape relationship through capturing the long-ranged dependencies between local underlying shapes. This shape-oriented operator is stacked into our hierarchical learning architecture, namely Shape-Oriented Convolutional Neural Network (SOCNN), developed for point cloud analysis. Extensive experiments have been performed to evaluate its significance in the tasks of point cloud classification and part segmentation. Chaoyi Zhang, Yang Song 0001, Lina Yao 0001, Tom Weidong Cai |
AAAI | 1 |
| 2020 | BiO-Net: Learning Recurrent Bi-directional Connections for Encoder-Decoder Architecture
Tiange Xiang, Chaoyi Zhang, Dongnan Liu, Yang Song 0001, Heng Huang 0001, Tom Weidong Cai |
MICCAI (1) | 2 |
| 2020 | DeepAntigen: a novel method for neoantigen prioritization via 3D genome and deep sparse learningabstractMOTIVATION: The mutations of cancers can encode the seeds of their own destruction, in the form of T-cell recognizable immunogenic peptides, also known as neoantigens. It is computationally challenging, however, to accurately prioritize the potential neoantigen candidates according to their ability of activating the T-cell immunoresponse, especially when the somatic mutations are abundant. Although a few neoantigen prioritization methods have been proposed to address this issue, advanced machine learning model that is specifically designed to tackle this problem is still lacking. Moreover, none of the existing methods considers the original DNA loci of the neoantigens in the perspective of 3D genome which may provide key information for inferring neoantigens' immunogenicity. RESULTS: In this study, we discovered that DNA loci of the immunopositive and immunonegative MHC-I neoantigens have distinct spatial distribution patterns across the genome. We therefore used the 3D genome information along with an ensemble pMHC-I coding strategy, and developed a group feature selection-based deep sparse neural network model (DNN-GFS) that is optimized for neoantigen prioritization. DNN-GFS demonstrated increased neoantigen prioritization power comparing to existing sequence-based approaches. We also developed a webserver named deepAntigen (http://yishi.sjtu.edu.cn/deepAntigen) that implements the DNN-GFS as well as other machine learning methods. We believe that this work provides a new perspective toward more accurate neoantigen prediction which eventually contribute to personalized cancer immunotherapy. AVAILABILITY AND IMPLEMENTATION: Data and implementation are available on webserver: http://yishi.sjtu.edu.cn/deepAntigen. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi Shi 0007, Zehua Guo 0004, Xianbin Su, Luming Meng, Minhua Zheng, Xueyin Shang, Wangqiu Cheng, Yaoliang Yu, Yujia Cai, Chaoyi Zhang, Tom Weidong Cai, Guang He, Zeguang Han |
Bioinform. | 14 |
| 2019 | Nuclei Segmentation via a Deep Panoptic Model with Semantic Feature FusionabstractAutomated detection and segmentation of individual nuclei in histopathology images is important for cancer diagnosis and prognosis. Due to the high variability of nuclei appearances and numerous overlapping objects, this task still remains challenging. Deep learning based semantic and instance segmentation models have been proposed to address the challenges, but these methods tend to concentrate on either the global or local features and hence still suffer from information loss. In this work, we propose a panoptic segmentation model which incorporates an auxiliary semantic segmentation branch with the instance branch to integrate global and local features. Furthermore, we design a feature map fusion mechanism in the instance branch and a new mask generator to prevent information loss. Experimental results on three different histopathology datasets demonstrate that our method outperforms the state-of-the-art nuclei segmentation methods and popular semantic and instance segmentation models by a large margin. Dongnan Liu, Donghao Zhang 0004, Yang Song 0001, Chaoyi Zhang, Fan Zhang 0013, Lauren O'Donnell, Tom Weidong Cai |
IJCAI | 4 |
| 2019 | Vessel-Net: Retinal Vessel Segmentation Under Multi-path Supervision
Yicheng Wu 0001, Yong Xia 0001, Yang Song 0001, Donghao Zhang 0004, Dongnan Liu, Chaoyi Zhang, Tom Weidong Cai |
MICCAI (1) | 6 |
| 2018 | Whole Slide Image Classification via Iterative Patch LabellingabstractBrain tumor can be a fatal disease in the world. With the aim of improving survival rates, many computerized algorithms have been proposed to assist the pathologists to make a diagnosis' using Whole Slide Pathology Images (WSI). Most methods focus on performing patch-level classification and aggregating the patch-level results to obtain the image classification. Since not all patches carry diagnostic information, it is thus important for our algorithm to recognize discriminative and non-discriminative patches. In this study, we propose an iterative patch labelling algorithm based on the Convolutional Neural Network (CNN), with a well-designed thresholding scheme, a training policy and a novel discriminative model architecture, to distinguish patches and use the discriminative ones to achieve WSI -classification. Our method is evaluated on the MICCAI 2015 Challenge Dataset, and shows a large improvement over the baseline approaches. Chaoyi Zhang, Yang Song 0001, Donghao Zhang 0004, Sidong Liu, Tom Weidong Cai |
ICIP | 1 |
| 2013 | A relay selection algorithm for radio and television services based on time-delay and bandwidthabstractThis paper presents a relay routing method for Radio and TV services, this method through obtaining node’s timedelay and power information, obtains the value of system interrupt decisions, and as a decision threshold to select relay node. While in consideration of link priorities and fairness, we design a relay routing protocol that can dynamically change the route when network is changed. Simulation results show that this protocol can expand coverage, reduce communication blind spots, increase system throughput and enhance the quality of service. Chaoyi Zhang, Muqing Wu, Linlin Luan, Chunxiu Xu |
ICMV | 1 |
| 2013 | An optimal PSO distributed precoding algorithm in QRD-based multi-relay system
Chaoyi Zhang, Muqing Wu, Linlin Luan |
Future Gener. Comput. Syst. | 1 |