Chao Tong 0001

dblp:18/5379-1 · DBLP profile ↗
← Back
33ranked-venue papers
14as first author
16since 2021 · last 2026
0000-0003-4414-4965ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 9 since 2021Computer networks · 10 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MoMIL: Multi-order enhanced multiple instance learning for computational pathology
Jiakai Wang, Baoyu Liang, Yuancheng Yang, Chao Tong 0001
Image Vis. Comput.6
2026 Prioritized scanning: Combining spatial information multiple instance learning for computational pathology
Jiakai Wang, Baoyu Liang, Yuancheng Yang, Siyang Wu, Chao Tong 0001
Pattern Recognit.6
2026 HGroupScene: Hierarchical Grouping and Similar Aggregation for 3D Semantic Scene Completion
abstract
3D Semantic Scene Completion (SSC) aims to infer voxel-level occupancy and semantics from partial 2D observations. However, existing methods often rely on global attention or uniform voxel modeling, which may cause semantic interference across unrelated regions and degrade performance under occlusion. To address this, we propose HGroupScene, a unified framework that integrates spatial priors and region-constrained reasoning for robust SSC. HGroupScene introduces: 1) a Hierarchical Grouping Module that partitions the voxel space into subregions and performs semantic aggregation via differentiable Gumbel-Softmax attention; 2) a dual-branch architecture composed of an Explicit Constraint Branch for extracting region-level structural features and an Implicit Diffusion Branch for fine-grained semantic reasoning; and 3) a Region-Constrained Feature Diffusion Mechanism that enables controllable feature propagation under the guidance of structural region priors. Extensive experiments on SemanticKITTI and SSCBench-KITTI360 under both single-frame and multi-frame settings demonstrate that HGroupScene achieves competitive or superior performance compared to state-of-the-art methods, validating the effectiveness of spatially structured semantic modeling.
Baoyu Liang, Shoutai Zhu, Zhenyu Xing, Chao Tong 0001
IEEE Trans. Image Process.5
2026 Lightweight Deep Learning Framework for Real-Time Traffic Sign Detection on In-Vehicle Embedded Computing Platforms
abstract
Traffic sign detection is crucial for both autonomous and manual driving, but high-resolution images from in-vehicle cameras pose challenges due to the small size of distant objects and the substantial computational cost. Due to the limited computational and parallel processing capabilities of in-vehicle computing platforms, achieving real-time, high-precision traffic sign detection in practical scenarios based on high-resolution images becomes impractical. Therefore, we propose a lightweight subgraph-based deep learning traffic sign detection framework tailored for in-vehicle embedded computing devices. The framework comprises two sub-models: 1) a Subgraph Filtering Model based on lightweight Convolutional Neural Networks, and 2) a Lightweight Deep Learning Traffic Sign Detection Model. The proposed framework enables to input high-resolution images while significantly reducing computational costs and produces high-precision traffic sign detection results. Experimental results on the TT100K dataset show an [email protected] of 0.954 with only 0.93M parameters, demonstrating the framework’s ability to perform real-time, high-precision traffic sign detection on in-vehicle embedded devices.
Yuancheng Yang, Chao Tong 0001
IEEE Trans. Intell. Transp. Syst.3
2025 VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
abstract
Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose the task of video event understanding that extracts event scripts and makes predictions with these scripts from videos. To support this task, we introduce VidEvent, a large-scale dataset containing over 23,000 well-labeled events, featuring detailed event structures, broad hierarchies, and logical relations extracted from movie recap videos. The dataset was created through a meticulous annotation process, ensuring high-quality and reliable event data. We also provide comprehensive baseline models offering detailed descriptions of their architecture and performance metrics. These models serve as benchmarks for future research, facilitating comparisons and improvements. Our analysis of VidEvent and the baseline models highlights the dataset's potential to advance video event understanding and encourages the exploration of innovative algorithms and models.
Baoyu Liang, Qile Su, Shoutai Zhu, Chao Tong 0001
AAAI5
2025 Separable and Flexible Classification on Pseudo-Features for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) presents significant challenges due to the constraints of the sequential learning framework, which limits the availability of training samples for new classes and prohibits revisiting samples from previously learned classes. To address these challenges, we propose a novel approach that leverages dynamic classifier weights and gradient backpropagation from pseudo-features of all encountered classes to classify incrementally appearing content. By directly synthesizing features for previously learned classes, we address the issue of catastrophic forgetting. Moreover, generating features for current classes helps mitigate overfitting, especially given the scarcity of original features. Our method utilizes pseudo-features to form deformable weights for each class, recognizing that optimal weights can vary across sessions, even for the same class. Unlike prior approaches that rely on fixed classification weights per class or generative models with pre-calculated parameters (e.g., GAN), which may produce sub-optimal pseudo-features for new classes, we introduce an updated sparse reconstruction method. This method generates pseudo-features through Sampling and Reconstruction (S&R). Extensive experiments on benchmark datasets, including mini-ImageNet, CUB-200, and CIFAR-100, demonstrate that our method outperforms state-of-the-art approaches under comparable experimental settings, highlighting its effectiveness in addressing the challenges of FSCIL.
Yuancheng Yang, Luyang Jin, Chao Tong 0001
ICMR4
2025 EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
abstract
Script event induction, which aims to predict the subsequent event based on the context, is a challenging task in NLP, achieving remarkable success in practical applications. However, human events are mostly recorded and presented in the form of videos rather than scripts, yet there is a lack of related research in the realm of vision. To address this problem, we introduce AVEP (Action-centric Video Event Prediction), a task that distinguishes itself from existing video prediction tasks through its incorporation of more complex logic and richer semantic information. We present a large structured dataset, which consists of about 35K annotated videos and more than 178K video clips of event, built upon existing video event datasets to support this task. The dataset offers more fine-grained annotations, where the atomic unit is represented as a multimodal event argument node, providing better structured representations of video events. Due to the complexity of event structures, traditional visual models that take patches or frames as input are not well-suited for AVEP. We propose EventFormer, a node-graph hierarchical attention based video event prediction model, which can capture both the relationships between events and their arguments and the coreferencial relationships between arguments. We conducted experiments using several SOTA video prediction models as well as LVLMs on AVEP, demonstrating both the complexity of the task and the value of the dataset. Our approach outperforms all these video prediction models. We will release the dataset and code for replicating the experiments and annotations.
Qile Su, Shoutai Zhu, Baoyu Liang, Chao Tong 0001
ACM Multimedia5
2025 SlimDL: Deploying ultra-light deep learning model on sweeping robots
Xudong Sun 0020, Zhanglin Liu, Shaoxuan Gao, Chao Tong 0001
Eng. Appl. Artif. Intell.6
2025 CosineTR: A dual-branch transformer-based network for semantic line detection
Bole Ma, Luyang Jin, Yuancheng Yang, Chao Tong 0001
Pattern Recognit.5
2024 Laplacian Projection Based Global Physical Prior Smoke Reconstruction
abstract
We present a novel framework for reconstructing fluid dynamics in real-life scenarios. Our approach leverages sparse view images and incorporates physical priors across long series of frames, resulting in reconstructed fluids with enhanced physical consistency. Unlike previous methods, we utilize a differentiable fluid simulator (DFS) and a differentiable renderer (DR) to exploit global physical priors, reducing reconstruction errors without the need for manual regularization coefficients. We introduce divergence-free Laplacian eigenfunctions (div-free LE) as velocity bases, improving computational efficiency and memory usage. By employing gradient-related strategies, we achieve better convergence and superior results. Extensive experiments demonstrate the effectiveness of our method, showcasing improved reconstruction quality and computational efficiency compared to existing approaches. We validate our approach using both synthetic and real data, highlighting its practical potential.
Shibang Xiao, Chao Tong 0001, Qifan Zhang 0003, Yunchi Cen, Frederick W. B. Li, Xiaohui Liang 0001
IEEE Trans. Vis. Comput. Graph.2
2023 G2IFu: Graph-based implicit function for single-view 3D reconstruction
Rongshan Chen, Yuancheng Yang, Chao Tong 0001
Eng. Appl. Artif. Intell.3
2023 Multi-view Pixel2Mesh++: 3D reconstruction via Pixel2Mesh with more images
Rongshan Chen, Xiang Yin 0004, Yuancheng Yang, Chao Tong 0001
Vis. Comput.4
2022 Joint Visual Perception and Linguistic Commonsense for Daily Events Causality Reasoning
abstract
Multimodal networks that juxtapose visual and linguistic modalities are currently widely adopted for solving vision-and-language tasks. They perform well in simple and intu-itive tasks, but are prone to mistakes in tasks involving latent or implicit details, due to the difficulty of capturing crucial but imperceptible visual signals in the real world. Perception errors lead to nonsensical results, but can be corrected by commonsense knowledge. To this end, we combine visual perception and linguistic commonsense to solve the challenging daily events causality reasoning task. We propose a novel Object-Aware Reasoning Network to focus on object inter-action while ignoring distracting information to refine visual perception. Further, a language branch with an independent prediction head is supervised to learn causality commonsense to help correct obvious perception errors, resulting in more plausible conclusions. Extensive experiments demonstrate that our method achieves new state-of-the-art results on Vis-Causal dataset.
Bole Ma, Chao Tong 0001
ICME2
2022 Protecting image privacy through adversarial perturbation
Baoyu Liang, Chao Tong 0001, Chao Lang, Joel J. P. C. Rodrigues, Sergei A. Kozlov
Multim. Tools Appl.2
2021 Pulmonary Nodule Classification Based on Heterogeneous Features Learning
abstract
Pulmonary cancer is one of the most dangerous cancers with a high incidence and mortality. An early accurate diagnosis and treatment of pulmonary cancer can observably increase the survival rates, where computer-aided diagnosis systems can largely improve the efficiency of radiologists. In this article, we propose a deep automated lung nodule diagnosis system based on three-dimensional convolutional neural network (3D-CNN) and support vector machine (SVM) with multiple kernel learning (MKL) algorithms. The system not only explores the computed tomography (CT) scans, but also the clinical information of patients like age, smoking history and cancer history. To extract deeper image features, a 34-layers 3D Residual Network (3D-ResNet) is employed. Heterogeneous features including the extracted image features and the clinical data are learned with MKL. The experimental results prove the effectiveness of the proposed image feature extractor and the combination of heterogeneous features in the task of lung nodule diagnosis.
Chao Tong 0001, Baoyu Liang, Mengbo Yu, Jiexuan Hu, Ali Kashif Bashir, Zhigao Zheng 0001
IEEE J. Sel. Areas Commun.1
2021 An Image Privacy Protection Algorithm Based on Adversarial Perturbation Generative Networks
abstract
Today, users of social platforms upload a large number of photos. These photos contain personal private information, including user identity information, which is easily gleaned by intelligent detection algorithms. To thwart this, in this work, we propose an intelligent algorithm to prevent deep neural network (DNN) detectors from detecting private information, especially human faces, while minimizing the impact on the visual quality of the image. More specifically, we design an image privacy protection algorithm by training and generating a corresponding adversarial sample for each image to defend DNN detectors. In addition, we propose an improved model based on the previous model by training an adversarial perturbation generative network to generate perturbation instead of training for each image. We evaluate and compare our proposed algorithm with other methods on wider face dataset and others by three indicators: Mean average precision, Averaged distortion, and Time spent. The results show that our method significantly interferes with DNN detectors while causing weak impact to the visual quality of images, and our improved model does speed up the generation of adversarial perturbations.
Chao Tong 0001, Chao Lang, Zhigao Zheng 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Robust nuclei segmentation in histopathology using ASPPU-Net and boundary refinement
Tao Wan 0001, Hongxiang Feng, Chao Tong 0001, Zengchang Qin
Neurocomputing5
2020 Pulmonary Nodule Detection Based on ISODATA-Improved Faster RCNN and 3D-CNN with Focal Loss
abstract
The early diagnosis of pulmonary cancer can significantly improve the survival rate of patients, where pulmonary nodules detection in computed tomography images plays an important role. In this article, we propose a novel pulmonary nodule detection system based on convolutional neural networks (CNN). Our system consists of two stages, pulmonary nodule candidate detection and false positive reduction. For candidate detection, we introduce Iterative Self-Organizing Data Analysis Techniques Algorithm (ISODATA) to Faster Region-based Convolutional Neural Network (Faster R-CNN) model. For false positive reduction, a three-dimensional convolutional neural network (3D-CNN) is employed to completely utilize the three-dimensional nature of CT images. In this network, Focal Loss is used to solve the class imbalance problem in this task. Experiments were conducted on LUNA16 dataset. The results show the preferable performance of the proposed system and the effectiveness of using ISODATA and Focal loss in pulmonary nodule detection is proved.
Chao Tong 0001, Baoyu Liang, Rongshan Chen, Arun Kumar Sangaiah, Zhigao Zheng 0001, Tao Wang 0037, Chenyang Yue
ACM Trans. Multim. Comput. Commun. Appl.1
2019 A deep automated skeletal bone age assessment model via region-based convolutional neural network
Baoyu Liang, Yunkai Zhai, Chao Tong 0001, Jun Li 0045, Xianying He, Qianqian Ma
Future Gener. Comput. Syst.3
2019 TimeTrustSVD: A collaborative filtering model integrating time, trust and rating information
Chao Tong 0001, Yu Lian, Jianwei Niu 0002, Joel J. P. C. Rodrigues
Future Gener. Comput. Syst.1
2018 A shilling attack detector based on convolutional neural network for collaborative recommender system in social aware network
abstract
One of the most fundamental tasks in the socially aware network (SAN) paradigm is to explore the attributes and behavior of users, which helps to design more suitable and efficient protocols. Particularly, detection of shilling attackers by mining users’ behavior is a frequently discussed topic in many social scenes like recommender systems based on collaborative filtering. As the performances of collaborative filtering are entirely based on ratings provided by users, they are vulnerable to shilling attacks which perform injection of biased profiles into rating databases to alter the systems. Current shilling attack detection methods detect spam users through artificially designed features, which are neither robust nor efficient enough. This paper illustrates a novel convolutional neural network-based method named CNN-SAD, which applies transformed network structure to exploit deep-level features from users rating profiles. Since the achieved deep-level features elaborate users rating more precisely than artificially designed features, CNN-SAD can detect shilling attacks more efficiently. According to the experimental results, the proposed method is capable of detecting the vast majority of obfuscated attacks precisely and outperforms other state-of-the-art algorithms, which contributes to applications and security in SAN.
Chao Tong 0001, Xiang Yin 0004, Jun Li 0045, Tongyu Zhu, Renli Lv, Joel J. P. C. Rodrigues
Comput. J.1
2018 Deep detection network for real-life traffic sign in vehicular networks
Xiang Long, Arun Kumar Sangaiah, Zhigao Zheng 0001, Chao Tong 0001
Comput. Networks5
2018 A novel deep learning method for aircraft landing speed prediction based on cloud-based sensor data
Chao Tong 0001, Xiang Yin 0004, Shili Wang, Zhigao Zheng 0001
Future Gener. Comput. Syst.1
2018 An efficient deep model for day-ahead electricity load forecasting with stacked denoising auto-encoders
Chao Tong 0001, Jun Li 0045, Chao Lang, Fanxin Kong, Jianwei Niu 0002, Joel J. P. C. Rodrigues
J. Parallel Distributed Comput.1
2018 A novel rating prediction method based on user relationship and natural noise
Chao Tong 0001, Yu Lian, Jianwei Niu 0002, Xiang Long
Multim. Tools Appl.1
2018 Introduction to the special issue on deep learning for biomedical and healthcare applications
Chao Tong 0001, Larry R. Medsker, Long Cheng 0005, Xiaojun Wang 0001
Neural Comput. Appl.1
2017 A Novel Classification Algorithm for New and Used Banknotes
Chao Tong 0001, Yu Lian, Zhongyu Xie, Jingyuan Feng, Fumin Zhu
Mob. Networks Appl.1
2016 A novel green algorithm for sampling complex networks
Chao Tong 0001, Yu Lian, Jianwei Niu 0002, Zhongyu Xie
J. Netw. Comput. Appl.1
2015 Weighted label propagation algorithm for overlapping community detection
abstract
Overlapping community detection algorithm research is one of hot topics in current social network analysis. In this paper, we applied the idea of weighted label propagation to overlapping community detection algorithm design, and propose a weighted label propagation algorithm (WLPA). Moreover, in order to evaluate the performance results of various overlapping community detection algorithms, we put forward a series of evaluation criteria based on error distribution curve of overlapping vertices. The experiment results show that the algorithm has a faster speed and better community detection results, and the evaluation criteria is in line with the inherent characteristics of the social network overlapping community structure.
Chao Tong 0001, Jianwei Niu 0002, Jinming Wen, Zhongyu Xie, Fu Peng
ICC1
2013 Real-time generation of personalized home video summaries on mobile devices
Jianwei Niu 0002, Da Huo 0004, Kongqiao Wang, Chao Tong 0001
Neurocomputing4
2012 Evolution of disconnected components in social networks: Patterns and a generative model
abstract
The majority of previous studies have focused on the analyses of an entire graph (network) or the giant connected component in a graph. Here we study the disconnected components (non-giant connected components) in real social networks, and reporting some interesting discoveries on how these disconnected components evolve over time. We study six diverse, real networks (citation networks, online social networks, academic collaboration networks, and others), and make the following major contributions: (a) we make empirical observations of the longevity distribution of disconnected components, and find that the curve of the distribution demonstrates a decaying trend; (b) we find that the distributions of final size of disconnected components that merge with one another or get absorbed by the giant connected component both follow power laws; (c) we find that the majority of mergings are between disconnected components and the giant connected component. The mergings that happen among disconnected components are small in scale (involve only a few components). The longevity distributions of the disconnected components in those mergings are similar, where the shortest-lived disconnected components are the most in number; and (d) we propose an empirical generative model that can produce the networks with our observed patterns.
Jianwei Niu 0002, Chao Tong 0001, Wanjiun Liao
IPCCC3
2012 Complex networks clustering algorithm based on the core influence of the nodes
abstract
This paper proposes a clustering method, group-driven process in the sociology of the clustering process simulation, whose clustering process simulates the driven process of colony formation in sociology. Clustering experiments show that the clustering accuracy and clustering speed of the algorithm in complex network are superior to the classic optimization Fast-Newman clustering algorithm.
Chao Tong 0001, Jianwei Niu 0002, Bin Dai 0009, Jinyang Fan
IPCCC1
2012 Complex networks properties analysis for mobile ad hoc networks
abstract
Recently, research on complex network theory and applications draws a lot of attention in both academy and industry. In mobile ad hoc networks (MANETs) area of research, a critical issue is to design the most effective topology for given problems. It is natural and significant to consider complex networks topology when optimising the MANET topology. Current works usually transform MANET or sensor network topologies into either small-world or scale-free. However, some fundamental problems remain unsolved. Specifically, what are the average shortest path length, degree distribution and clustering characteristics of MANETs? Do MANETs have small-world effect and scale-free property? In this work, the authors introduce complex networks theory into the context of MANET topology and study complex network properties of the MANETs to answer the above questions. The authors have theoretically analysed the degree distribution and clustering coefficient of MANETs and proposed approach to computing them. The degree distribution and clustering coefficient of MANETs are theoretically deduced from node space probability distribution on different mobility models (including but not limited to random waypoint model). Simulation results on average shortest path length, clustering coefficient and degree distribution show that in most cases MANETs do not have the small-world effect and scale-free property.
Chao Tong 0001, Jianwei Niu 0002, Guangzhi Qu, Xiang Long, Xiaopeng Gao
IET Commun.1