Tongtong Yuan

dblp:187/9122 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-8224-9891ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 7 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Explainable CNN filter pruning based on shapley value approximation in view of game theory
Tongtong Yuan, Bo Liu 0011, Yinan Tang
Expert Syst. Appl.1
2026 Adaptive metric for knowledge distillation by deep Bregman divergence
Tongtong Yuan, Bo Liu 0011, Yinan Tang
Neural Networks1
2026 Spurious Local Minima Provably Exist for Deep CNNs: Theory and Application
abstract
In this article, we prove that a general family of spurious local minima exist in the loss landscape of deep convolutional neural networks (CNNs) with strictly convex loss functions and ReLU activations. Our construction of spurious local minima is general and applies to CNNs with arbitrary architectures. We construct a local minimum $\theta $ at first, and then construct another point $\theta ^{\prime } $ in parameter space with the same empirical risk as $\theta $ . Data samples are split into some groups such that each group behaves differently under the perturbation around $\theta ^{\prime } $ to produce a lower empirical risk. We tackle the challenges caused by convolutional layers in the construction. We show that a differentiation of data samples is always possible somewhere in the feature maps, and despite network parameters being tied in each feature map, our perturbation scheme only affects the output of a single or a few neurons for a group of data samples. We then give an example of nontrivial spurious local minimum in which multiple activation patterns are explicitly constructed. Finally, based on our construction of spurious local minima, we design a deterministic optimization method to escape local minima that is applicable to CNNs, ResNets, MLPs, and transformers. Experimental results on CIFAR-10, CIFAR-100, and ImageNet-1k datasets verify our theoretical findings and show that our optimization method outperforms SGD or Adam in accuracy (by 0.27% on average) consistently on all these architectures and datasets.
Bo Liu 0011, Keyi Fu, Tongtong Yuan, Shen Geng
IEEE Trans. Neural Networks Learn. Syst.3
2025 MS-UFAD: A Large-Scale Dataset for Real-world Unified Face Attack Detection with Text Descriptions
abstract
As deepfake and adversarial attacks evolve, facial recognition systems are encountering increasingly diverse threats. Most existing face liveness detection algorithms focus on single tasks, like spoofing or deepfake attack detection. The corresponding datasets have limited coverage of attack methods, with original data mostly sourced from the internet or laboratory environments. Moreover, existing datasets lack textual annotations, particularly for attack clues, limiting algorithms’ ability to utilize semantic assistance from text. To address these issues, we propose a large-scale unified attack dataset, which includes newly collected facial videos from 5,000 individuals, along with generated videos corresponding to 52 face attack methods. The dataset contains 795k videos and 60k images across four different quality levels. Through semi-automated annotation, we provide detailed textual descriptions. This is the first face attack dataset with textual descriptions. Additionally, we propose a text-guided face attack detection method, demonstrating significant improvements in accuracy using fine-grained textual descriptions. Our dataset will be released at https://ms-ufad.github.io.
Dingheng Zeng, Zhifei Kong, Tongtong Yuan, Weihong Deng, Ying Li 0012
ICASSP9
2025 Layer-Wise Vision Injection With Disentangled Attention for Efficient LVLMs
Xuange Zhang, Dengjie Li, Bo Liu 0011, Zenghao Bao, Baisong Yang, Zhongying Liu, Tongtong Yuan
ICCV9
2025 Federated Knowledge Distillation Based on Prompt for Matching Data Distribution
Yizhang Liu, Wenze Xu, Tongtong Yuan, Weihong Deng
PRCV (18)4
2025 Rethinking cell-based neural architecture search: A theoretical perspective
Bo Liu 0011, Huiwen Zhao, Tongtong Yuan, Ting Zhang 0012, Zhaoying Liu
Neural Networks3
2025 Surveillance Video-and-Language Understanding: From Small to Large Multimodal Models
abstract
Surveillance videos play a crucial role in public security. However, current tasks related to surveillance videos primarily focus on classifying and localizing anomalous events. Despite achieving notable performance, existing methods are restricted to detecting and classifying predefined events and lack satisfactory semantic understanding. To tackle this challenge, we introduce a novel research avenue focused on Video-and-Language Understanding for surveillance (VALU), and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset, UCA (UCF-Crime Annotation), contains 23,542 sentences, with an average length of 20 words, and its annotated videos are as long as 110.7 hours. Moreover, we evaluate SOTA models on five multimodal tasks using this newly created dataset, establishing new baselines for surveillance VALU, from small to large models. Our experiments reveal that mainstream models, which perform well on previously public datasets, exhibit poor performance on surveillance video, highlighting new challenges in surveillance VALU. In addition to conducting baseline experiments to compare the performance of existing models, we also propose novel methods for multimodal anomaly detection tasks and finetune multimodal large language model models using our dataset. All the experiments highlight the necessity of constructing this multimodal dataset to advance surveillance AI. Upon the experimental results mentioned above, we conduct further in-depth analysis and discussion. The dataset and codes are provided athttps://xuange923.github.io/Surveillance-Video-Understanding.
Tongtong Yuan, Xuange Zhang, Bo Liu 0011, Zhenzhen Jiao
IEEE Trans. Circuits Syst. Video Technol.1
2024 Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges
abstract
Surveillance videos are important for public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying the predefined events with unsatisfactory semantic understanding, although they have obtained considerable performance. To address this issue, we propose a new research direction of surveillance video-and-language understanding (VALU), and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset, UCA (UCF-Crime Annotation)11The dataset is provided at https://xuange923.github.io/Surveillance-Video-Understanding., contains 23,542 sentences, with an average length of 20 words, and its annotated videos are as long as 110.7 hours. Furthermore, we benchmark SOTA models for four multimodal tasks on this newly created dataset, which serve as new baselines for surveillance VALU. Through experiments, we find that mainstream models used in previously public datasets perform poorly on surveillance video, demonstrating new challenges in surveillance VALU. We also conducted experiments on multimodal anomaly detection. These results demonstrate that our multimodal surveillance learning can improve the performance of anomaly detection. All the experiments highlight the necessity of constructing this dataset to advance surveillance AI.
Tongtong Yuan, Xuange Zhang, Bo Liu 0011, Zhenzhen Jiao
CVPR1
2024 ARPruning: An automatic channel pruning based on attention map ranking
Tongtong Yuan, Zulin Li, Bo Liu 0011, Yinan Tang
Neural Networks1
2024 Explainable Federated Medical Image Analysis Through Causal Learning and Blockchain
abstract
Federated learning (FL) enables collaborative training of machine learning models across distributed medical data sources without compromising privacy. However, applying FL to medical image analysis presents challenges like high communication overhead and data heterogeneity. This paper proposes novel FL techniques using explainable artificial intelligence (XAI) for efficient, accurate, and trustworthy analysis. A heterogeneity-aware causal learning approach selectively sparsifies model weights based on their causal contributions, significantly reducing communication requirements while retaining performance and improving interpretability. Furthermore, blockchain provides decentralized quality assessment of client datasets. The assessment scores adjust aggregation weights so higher-quality data has more influence during training, improving model generalization. Comprehensive experiments show our XAI-integrated FL framework enhances efficiency, accuracy and interpretability. The causal learning method decreases communication overhead while maintaining segmentation accuracy. The blockchain-based data valuation mitigates issues from low-quality local datasets. Our framework provides essential model explanations and trust mechanisms, making FL viable for clinical adoption in medical image analysis.
Junsheng Mu, Michel Kadoch, Tongtong Yuan, Wenzhe Lv, Qiang Liu 0030, Bohan Li 0005
IEEE J. Biomed. Health Informatics3
2024 A Survey on Performance Modeling and Prediction for Distributed DNN Training
abstract
The recent breakthroughs in large-scale DNN attract significant attention from both academia and industry toward distributed DNN training techniques. Due to the time-consuming and expensive execution process of large-scale distributed DNN training, it is crucial to model and predict the performance of distributed DNN training before its actual deployment, in order to optimize the design of distributed DNN training at low cost. This paper analyzes and emphasizes the importance of modeling and predicting the performance of distributed DNN training, categorizes and analyses the related state-of-the-art works, and discusses future challenges and opportunities for this research field. The objectives of this paper are twofold: first, to assist researchers in understanding and choosing suitable modeling and prediction tools for large-scale distributed DNN training, and second, to encourage researchers to propose more valuable research about performance modeling and prediction for distributed DNN training in the future.
Zhenhua Guo 0003, Yinan Tang, Jidong Zhai, Tongtong Yuan, Li Wang 0040, Yaqian Zhao, RenGang Li
IEEE Trans. Parallel Distributed Syst.4
2022 Bridging the Gap Between Semantic Segmentation and Instance Segmentation
abstract
Fine-grained instance segmentation is considerably more complicated and challenging than semantic segmentation. Most existing instance segmentation methods only focus on accuracy without paying much attention to inference latency, which, is critical to real-time applications, such as autonomous driving. In this paper, we aim to bridge the gap between semantic segmentation and instance segmentation by presenting a novel real-time model for instance segmentation, Sem2Ins, which effectively generates instance boundaries according to a semantic segmentation by leveraging conditional generative adversarial networks (cGANs) coupled with deep supervision and a weighted fusion layer. Specifically, supervision is imposed on each output layer, and features from different levels are fused to produce a well-generated boundary map. Sem2Ins has the following desirable features: 1) Combined with some fast semantic segmentation methods, Sem2Ins runs at a real-time speed that is fairly well-balanced against accuracy; 2) Sem2Ins works flexibly with any semantic segmentation model for instance segmentation, and if the given semantic segmentation is sufficiently good, Sem2Ins even achieves state-of-the-art in terms of accuracy; 3) deep supervision and weighted fusion can be leveraged to generate high-quality boundaries; and 4) Sem2Ins can be easily extended to panoptic segmentation. Extensive experiments performed on the Cityscapes, WildDash, KITTI and COCO benchmarks have demonstrated that 1) Sem2Ins, when combined with PSPNet and DDRNet-23-Slim, consistently outperforms the state-of-the-art real-time solution (Box2Pix) in terms of both speed and accuracy; and 2) Sem2Ins combined with DPC performs comparably to some powerful detect-and-segment approaches.
Chengxiang Yin 0001, Jian Tang 0008, Tongtong Yuan, Yanzhi Wang 0001
IEEE Trans. Multim.3
2021 Effective *-flow schedule for optical circuit switching based data center networks: A comprehensive survey
Yinan Tang, Tongtong Yuan, Bo Liu 0011, Chuangbai Xiao
Comput. Networks2
2021 Non-differentiable saddle points and sub-optimal local minima exist for deep ReLU networks
Bo Liu 0011, Zhaoying Liu, Ting Zhang 0012, Tongtong Yuan
Neural Networks4
2021 On Learning Semantic Representations for Large-Scale Abstract Sketches
abstract
In this paper, we focus on learning semantic representations for large-scale highly abstract sketches that were produced by the practical sketch-based application rather than the excessively well dawn sketches obtained by crowd-sourcing. We propose a dual-branch CNN-RNN network architecture to represent sketches, which simultaneously encodes both the static and temporal patterns of sketch strokes. Based on this architecture, we further explore learning the sketch-oriented semantic representations in two practical settings, i.e., hashing retrieval and zero-shot recognition on million-scale highly abstract sketches produced by practical online interactions. Specifically, we use our dual-branch architecture as a universal representation framework to design two sketch-specific deep models: (i) We propose a deep hashing model for sketch retrieval, where a novel hashing loss is specifically designed to further accommodate both the abstract and messy traits of sketches. (ii) We propose a deep embedding model for sketch zero-shot recognition, via collecting a large-scale edge-map dataset and proposing to extract a set of semantic vectors from edge-maps as the semantic knowledge for sketch zero-shot domain alignment. Both deep models are evaluated by comprehensive experiments on million-scale abstract sketches produced by a global online game QuickDraw and outperform state-of-the-art competitors.
Peng Xu 0005, Yongye Huang, Tongtong Yuan, Tao Xiang 0002, Timothy M. Hospedales, Yi-Zhe Song, Liang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 An Actor-Critic-Based Transfer Learning Framework for Experience-Driven Networking
abstract
Experience-driven networking has emerged as a new and highly effective approach for resource allocation in complex communication networks. Deep Reinforcement Learning (DRL) has been shown to be a useful technique for enabling experience-driven networking. In this paper, we focus on a practical and fundamental problem for experience-driven networking: when network configurations are changed, how to train a new DRL agent to effectively and quickly adapt to the new environment. We present an Actor-Critic-based Transfer learning framework for the Traffic Engineering (TE) problem using policy distillation, which we call ACT-TE. ACT-TE effectively and quickly trains a new DRL agent to solve the TE problem in a new network environment, using both old knowledge (i.e., distilled from the existing agent) and new experience (i.e., newly collected samples). We implement ACT-TE in ns-3, and compare it with commonly-used baselines using packet-level simulations on three representative network topologies: NSFNET, ARPANET and random topology. The extensive simulation results show that 1) The existing well-trained DRL agents do not work well in new network environments; 2) ACT-TE significantly outperforms both two straightforward methods (training from scratch and fine-tuning based on an existing DRL agent) and several widely-used traditional methods in terms of network utility, throughput and delay.
Dejun Yang, Jian Tang 0008, Yinan Tang, Tongtong Yuan, Yanzhi Wang 0001, Guoliang Xue
IEEE/ACM Trans. Netw.5
2019 Signal-To-Noise Ratio: A Robust Distance Metric for Deep Metric Learning
abstract
Deep metric learning, which learns discriminative features to process image clustering and retrieval tasks, has attracted extensive attention in recent years. A number of deep metric learning methods, which ensure that similar examples are mapped close to each other and dissimilar examples are mapped farther apart, have been proposed to construct effective structures for loss functions and have shown promising results. In this paper, different from the approaches on learning the loss structures, we propose a robust SNR distance metric based on Signal-to-Noise Ratio (SNR) for measuring the similarity of image pairs for deep metric learning. By exploring the properties of our SNR distance metric from the view of geometry space and statistical theory, we analyze the properties of our metric and show that it can preserve the semantic similarity between image pairs, which well justify its suitability for deep metric learning. Compared with Euclidean distance metric, our SNR distance metric can further jointly reduce the intra-class distances and enlarge the inter-class distances for learned features. Leveraging our SNR distance metric, we propose Deep SNR-based Metric Learning (DSML) to generate discriminative feature embeddings. By extensive experiments on three widely adopted benchmarks, including CARS196, CUB200-2011 and CIFAR10, our DSML has shown its superiority over other state-of-the-art methods. Additionally, we extend our SNR distance metric to deep hashing learning, and conduct experiments on two benchmarks, including CIFAR10 and NUS-WIDE, to demonstrate the effectiveness and generality of our SNR distance metric.
Tongtong Yuan, Weihong Deng, Jian Tang 0008, Yinan Tang, Binghui Chen
CVPR1
2019 Unsupervised adaptive hashing based on feature clustering
Tongtong Yuan, Weihong Deng, Jiani Hu, Zhanfu An, Yinan Tang
Neurocomputing1
2018 SketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval
abstract
We propose a deep hashing framework for sketch retrieval that, for the first time, works on a multi-million scale human sketch dataset. Leveraging on this large dataset, we explore a few sketch-specific traits that were otherwise under-studied in prior literature. Instead of following the conventional sketch recognition task, we introduce the novel problem of sketch hashing retrieval which is not only more challenging, but also offers a better testbed for large-scale sketch analysis, since: (i) more fine-grained sketch feature learning is required to accommodate the large variations in style and Abstraction, and (ii) a compact binary code needs to be learned at the same time to enable efficient retrieval. Key to our network design is the embedding of unique characteristics of human sketch, where (i) a two-branch CNN-RNN architecture is adapted to explore the temporal ordering of strokes, and (ii) a novel hashing loss is specifically designed to accommodate both the temporal and Abstract traits of sketches. By working with a 3.8M sketch dataset, we show that state-of-the-art hashing models specifically engineered for static images fail to perform well on temporal sketch data. Our network on the other hand not only offers the best retrieval performance on various code sizes, but also yields the best generalization performance under a zero-shot setting and when re-purposed for sketch recognition. Such superior performances effectively demonstrate the benefit of our sketch-specific design.
Peng Xu 0005, Yongye Huang, Tongtong Yuan, Kaiyue Pang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Zhanyu Ma, Jun Guo 0002
CVPR3
2018 Deep Transfer Network with 3D Morphable Models for Face Recognition
abstract
Data augmentation using 3D face models to synthesize faces has been demonstrated to be effective for face recognition. However, the model directly trained by using the synthesized faces together with the original real faces is not optimal. In this paper, we propose a novel approach that uses a deep transfer network (DTN) with 3D morphable models (3DMMs) for face recognition to overcome the shortage of labeled face images and the dataset bias between synthesized images and corresponding real images. We first utilize the 3DMM to synthesize faces with various poses to augment the training dataset. Then, we train a deep neural network using the synthesized face images and the original real face images. The results obtained on LFW show that the accuracy of the model utilizing synthesized data only is lower than that of the model using the original data, although the synthesized dataset contains much considerably images with more unconstrained poses. This result shows that a dataset bias exists between the synthesized faces and the real faces. We treat the synthesized faces as the source domain, and we treat the actual faces as the target domain. We use the DTN to alleviate the discrepancy between the source domain and the target domain. The DTN attempts to project source domain samples and target domain samples to a new space where they are fused together such that one cannot distinguish the domain from which a specific image is from. We optimize our DTN based on the maximum mean discrepancy (MMD) of the shared feature extraction layers and the discrimination layers. We choose AlexNet and Inception-ResNet-V1 as our benchmark models. The proposed method is also evaluated on the LFW and SLLFW databases. The experimental results show that our method can effectively address the domain discrepancy. Moreover, the dataset bias between the synthesized data and the real data is remarkably reduced, which can thus improve the performance of the convolutional neural network (CNN) model.
Zhanfu An, Weihong Deng, Tongtong Yuan, Jiani Hu
FG3
2017 Supervised hashing with extreme learning machine
abstract
Supervised hashing methods, which aim to generate semantic similarity-preserving binary codes, have been proposed to improve the performance of large-scale image retrieval. However, learning binary codes remains an NP-hard problem due to the binary constraints and complex computation. Existing hashing methods have never explored the potentiality of the label information, leading to a limited performance. To address these problems, we propose a simple supervised hashing method based on extreme learning machine (ELM). And we generate the supervised information in ELM by target code learning instead of using the traditional label code to fit the retrieval problem. With this modified label code, our method can produce high-quality binary codes and obtain high retrieval precision. Comprehensive experiments have shown our superiority to other state-of-the-art methods.
Tongtong Yuan, Weihong Deng, Jiani Hu
VCIP1