Tongtong Su

dblp:249/8204 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Zero-to-Hero: Empowering Video Appearance Transfer with Zero-Shot Initialization and Holistic Restoration
Tongtong Su, Chengyu Wang 0001, Haipeng Liao, Jun Huang 0007, Dongming Lu
AAAI1
2026 StableEKF-Transformer: Uncertainty-Aware State of Health Estimation with Dynamic Covariance Calibration and Diagonal Jacobian Parameterization
Jinqi Zhu, Tongtong Su, Di Lv, Weijia Feng, Chenyang Wang 0001
DASFAA (3)3
2026 The Power of Weighting: Multi-teacher Distillation for Communication-Efficient Federated Learning
Ruojia Zhang, Weijia Feng, Tongtong Su, Fengtao Sun, Chenyang Wang 0001, Chongke Bi
DASFAA (4)3
2026 Learning to Weigh and Distill: Gated Adaptive Knowledge Distillation for Multi-Teacher Allocation
Jiale Si, Huilin Liu, Chengmin Yan, Weijia Feng, Chenyang Wang 0001, Tongtong Su, Jinqi Zhu
INFOCOM6
2026 FedGRO: Group Relative Optimization for Resource-Efficient Federated Self-Supervised Learning in V2X
Boyue Zhang 0005, Weijia Feng, Ruojia Zhang, Rui Lan, Tongtong Su, Chenyang Wang 0001, Chongke Bi
INFOCOM5
2026 Identification of Influential Node Group in Attributed Graph through Explaining Graph Neural Network
abstract
Identification of influential groups of nodes in attributed graphs has applications in a wide range of real-world problems, for instance, collecting important proceedings in citation networks, or identifying essential genes for diagnosing disease in Protein-Protein Interaction networks. Previous approaches for influence maximization manipulated on the graph structure, despite their proliferation, neglect the node attribute information containing additional knowledge. In this work, we introduce Global Graph UNderstanding (GGUN), a perturbation-based framework leveraging the explanatory power of Graph Neural Networks. It takes into account the entire graph structure and node attributes simultaneously and fuses knowledge through GNN layers. Following the perturbation-based explanation, GGUN fills the gap between Deep Neural Network gradient-based feature importance analysis and discrete structure in the graph, which is formulated as a combinatorial optimization problem. Moreover, GGUN obtains an efficient solution by relaxing the infeasible combinatorial optimization problem with performance guaranteed. Evaluations of synthetic and real-world datasets show that GGUN outperforms baselines on both quantitative metrics and human-intelligible analysis.
Xiao Tan 0005, Tongtong Su, Yan Zhang 0100, Binghui Xu, Dian Shen, Meng Wang 0009, Beilun Wang
WWW2
2026 Teacher assistant-based knowledge distillation bridging architecture differences on heterogeneous models
Renyu Jiang, Tongtong Su, Jiale Si, Chenyang Wang 0001, Weijia Feng, Jinqi Zhu, Peiyan Yuan
Neurocomputing2
2026 Holistic prediction comparison for knowledge distillation
Tongtong Su, Chengmin Yan, Huilin Liu, Jiale Si, Xukai Wang, Jinqi Zhu, Xiguo Zhou
Neurocomputing1
2025 Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
abstract
In recent years, large text-to-video (T2V) synthesis models have garnered considerable attention for their abilities to generate videos from textual descriptions. However, achieving both high imaging quality and effective motion representation remains a significant challenge for these T2V models. Existing approaches often adapt pre-trained text-to-image (T2I) models to refine video frames, leading to issues such as flickering and artifacts due to inconsistencies across frames. In this paper, we introduce EVS, a training-free Encapsulated Video Synthesizer that composes T2I and T2V models to enhance both visual fidelity and motion smoothness of generated videos. Our approach utilizes a well-trained diffusion-based T2I model to refine low-quality video frames by treating them as out-of-distribution samples, effectively optimizing them with noising and denoising steps. Meanwhile, we employ T2V backbones to ensure consistent motion dynamics. By encapsulating the T2V temporal-only prior into the T2I generation process, EVS successfully leverages the strengths of both types of models, resulting in videos of improved imaging and motion quality. Experimental results validate the effectiveness of our approach compared to previous approaches. Our composition process also leads to a significant improvement of 1.6x-4.5x speedup in inference time.1
Tongtong Su, Chengyu Wang 0001, Jun Huang 0007, Dongming Lu
CVPR1
2025 AdaptEdit: An Adaptive Correspondence Guidance Framework for Reference-Based Video Editing
abstract
Video editing is a pivotal process for customizing video content according to user needs. However, existing text-guided methods often lead to ambiguities regarding user intentions and restrict fine-grained control for editing specific aspects in videos. To overcome these limitations, this paper introduces a novel approach named \emph{AdaptEdit}, which focuses on reference-based video editing that disentangles the editing process. It achieves this by first editing a reference image and then adaptively propagating its appearance across other frames to complete the video editing. While previous propagation methods, such as optical flow and the temporal modules of recent video generative models, struggle with object deformations and large motions, we propose an adaptive correspondence strategy that accurately transfers the appearance from the reference frame to the target frames by leveraging inter-frame semantic correspondences in the original video. By implementing a proxy-editing task to optimize hyperparameters for image token-level correspondence, our method effectively balances the need to maintain the target frame's structure while preventing leakage of irrelevant appearance. To more accurately evaluate editing beyond the semantic-level consistency provided by CLIP-style models, we introduce a new dataset, PVA, which supports pixel-level evaluation. Our method outperforms the best-performing baseline with a clear PSNR improvement of 3.6 dB.
Tongtong Su, Chengyu Wang 0001, Jun Huang 0007, Dongming Lu
IJCAI1
2025 SDPGO: Efficient Self-Distillation Training Meets Proximal Gradient Optimization
abstract
Self-knowledge distillation (SKD) enables single-model training by distilling knowledge from the model's own output, eliminating the need for a separate teacher network required in conventional distillation methods. However, current SKD methods focus mainly on replicating common features in the student model, neglecting the extraction of key features that significantly enhance student learning. Inspired by this, we devise a self-knowledge distillation framework entitled Self-Distillation training via Proximal Gradient Optimization or SDPGO, which utilizes gradient information to identify and assign greater weight to features that significantly impact classification performance, enabling the network to learn the most relevant features during training. Specifically, the proposed framework refines the gradient information into a dynamically changing weighting factor to evaluate the distillation knowledge via the dynamic weight adjustment scheme. Meanwhile, we devise the sequential iterative learning module to dynamically optimize knowledge transfer by leveraging historical predictions and real-time gradients, stabilizing training through mini-batch-based KL divergence refinement while adaptively prioritizing task-critical features for efficient self-distillation. Comprehensive experiments on image classification, object detection, and semantic segmentation demonstrate that our method consistently surpasses recent state-of-the-art knowledge distillation techniques. Code is available at: https://github.com/nanxiaotong/SDGPO.
Tongtong Su, Yun Liao, Fengbo Zheng
NeurIPS1
2024 Transmission line defect detection based on feature enhancement
Tongtong Su, Daming Liu
Multim. Tools Appl.1
2023 Self-Supervised Learning with Explorative Knowledge Distillation
abstract
Previous paradigms have combined self-supervised learning (SSL) with knowledge distillation to compress a self-supervised teacher model into a smaller student. In this work, we devise a self-supervised explorative distillation (SSED) algorithm to improve the representation quality of the lightweight models. We introduce a heterogeneous teacher to maximumly learn rich feature representation for the student, which reaches the expected goal of capturing discriminative feature information contained in network itself. SSED enforces the student to learn more diversified and perfect representations of the original class recognition task and self-supervised learning task. Extensive experiments show that SSED improves accuracy effectively on large and small models, and surpassing current top-performing SSL methods. Particularly, the linear results of our ResNet-18, trained with ResNet-50 teacher, achieves 65.5% ImageNet top-1 accuracy, which is 1.4% and 4.9% higher than OSS and DisCo. Code is available at https://github.com/nanxiaotong/SSED.
Tongtong Su, Gang Wang 0001, Xiaoguang Liu 0001
ICASSP1
2023 RADEAN: A Resource Allocation Model Based on Deep Reinforcement Learning and Generative Adversarial Networks in Edge Computing
Zhaoyang Yu 0003, Sinong Zhao, Tongtong Su, Xiaoguang Liu 0001, Gang Wang 0001, Zehua Wang 0001, Victor C. M. Leung
MobiQuitous (1)3
2023 Deep Cross-Layer Collaborative Learning Network for Online Knowledge Distillation
abstract
Recent online knowledge distillation (OKD) methods focus on capturing rich and useful intermediate information by performing multi-layer feature learning. Existing works only consider intermediate layer feature maps between the same layers and ignore valuable information across layers, which results in the lack of appropriate cross-layer supervision in detail and the process of learning. Besides, this manner provides insufficient supervision information to supervise the learning of student, since it fails to construct a qualified teacher. In this work, we propose a Deep Cross-layer Collaborative Learning network (DCCL) for OKD, which efficiently exploits fruitful knowledge of peer student models by keeping appropriate intermediate cross-layer supervision. Specifically, each student gradually integrates its own features at different layers for feature matching, so as to effectively utilize features in low and high levels for learning more composite knowledge. Moreover, we assign a collaborative knowledge learning strategy, in which a qualified teacher is established via fusing the features of last convolution layers for enhancing high-level representation. In this way, all student models continuously transfer the rich teacher’s internal representation as well as capture its dynamic growth process, and in turn assist the learning of the fusion teacher to further supervise students. In the experiments, our proposed DCCL has shown great generalization ability with various backbone models on CIFAR-100, Tiny ImageNet and ImageNet, and also demonstrated superior performance against mainstream OKD works. Our code is available here:https://github.com/nanxiaotong/DCCL.
Tongtong Su, Qiyu Liang, Zhaoyang Yu 0003, Ziyue Xu 0005, Gang Wang 0001, Xiaoguang Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 DDistill-SR: Reparameterized Dynamic Distillation Network for Lightweight Image Super-Resolution
abstract
Recent research on deep convolutional neural networks (CNNs) has provided a significant performance boost on efficient super-resolution (SR) tasks by trading off the performance and applicability. However, most existing methods focus on subtracting feature processing consumption to reduce the parameters and calculations without refining the immediate features, which leads to inadequate information in the restoration. In this paper, we propose a lightweight network termed DDistill-SR, which significantly improves the SR quality by capturing and reusing more helpful information in a static-dynamic feature distillation manner. Specifically, we propose a plug-in reparameterized dynamic unit (RDU) to promote the performance and inference cost trade-off. During the training phase, the RDU learns to linearly combine multiple reparameterizable blocks by analyzing varied input statistics to enhance layer-level representation. In the inference phase, the RDU is equally converted to simple dynamic convolutions that explicitly capture robust dynamic and static feature maps. Then, the information distillation block is constructed by several RDUs to enforce hierarchical refinement and selective fusion of spatial context information. Furthermore, we propose a dynamic distillation fusion (DDF) module to enable dynamic signals aggregation and communication between hierarchical modules to further improve performance. Empirical results show that our DDistill-SR outperforms the baselines and achieves state-of-the-art results on most super-resolution domains with much fewer parameters and less computational overhead.
Yan Wang 0086, Tongtong Su, Yusen Li, Jiuwen Cao, Gang Wang 0001, Xiaoguang Liu 0001
IEEE Trans. Multim.2
2023 STKD: Distilling Knowledge From Synchronous Teaching for Efficient Model Compression
abstract
Knowledge distillation (KD) transfers discriminative knowledge from a large and complex model (known as teacher) to a smaller and faster one (known as student). Existing advanced KD methods, limited to fixed feature extraction paradigms that capture teacher's structure knowledge to guide the training of the student, often fail to obtain comprehensive knowledge to the student. Toward this end, in this article, we propose a new approach, synchronous teaching knowledge distillation (STKD), to integrate online teaching and offline teaching for transferring rich and comprehensive knowledge to the student. In the online learning stage, a blockwise unit is designed to distill the intermediate-level knowledge and high-level knowledge, which can achieve bidirectional guidance of the teacher and student networks. Intermediate-level information interaction provides more supervisory information to the student network and is useful to enhance the quality of final predictions. In the offline learning stage, the STKD approach applies a pretrained teacher to further improve the performance and accelerate the training process by providing prior knowledge. Trained simultaneously, the student learns multilevel and comprehensive knowledge by incorporating online teaching and offline teaching, which combines the advantages of different KD strategies through our STKD method. Experimental results on the SVHN, CIFAR-10, CIFAR-100, and ImageNet ILSVRC 2012 real-world datasets show that the proposed method achieves significant performance improvements compared with the state-of-the-art methods, especially with satisfying accuracy and model size. Code for STKD is provided at https://github.com/nanxiaotong/STKD.
Tongtong Su, Zhaoyang Yu 0003, Gang Wang 0001, Xiaoguang Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 DeepSCJD: An Online Deep Learning-Based Model for Secure Collaborative Job Dispatching in Edge Computing
Zhaoyang Yu 0003, Sinong Zhao, Tongtong Su, Xiaoguang Liu 0001, Gang Wang 0001, Zehua Wang 0001, Victor C. M. Leung
ICSOC3
2021 Attention-based Feature Interaction for Efficient Online Knowledge Distillation
abstract
Existing online knowledge distillation (KD) methods solve the dependency problem of the high-capacity teacher model via mutual learning and ensemble learning. But they focus on the utilization of logits information in the last few layers and fail to construct a strong teacher model to better supervise student networks, leading to the inefficiency of KD. In this work, we propose a simple but effective online knowledge distillation algorithm, called Attentive Feature Interaction Distillation (AFID). It applies interactive teaching in which the teacher and the student can send, receive, and give feedback on an equal footing, ultimately promoting the generality of both. Specifically, we set up a Feature Interaction Module for two sub-networks to conduct low-level and mid-level feature learning. They can alternately transfer attentive features maps to exchange interesting regions and fuse the other party’s map with the features of self-extraction for information enhancement. Besides, we assign a Feature Fusion Module, in which a Peer Fused Teacher is formed to fuse the output features of two sub-networks to guide sub-networks and a Peer Ensemble Teacher is established to accomplish mutual learning between the two teachers. Integrating Feature Interaction Module and Feature Fusion Module into a unified framework takes full advantage of the interactive teaching mechanism and makes the two sub-networks capture and transfer more fine-grained features to each other. Experimental results on CIFAR-100 and ImageNet ILSVRC 2012 real datasets show that AFID achieves significant performance improvements compared with existing online KD and classical teacher-guide methods.
Tongtong Su, Qiyu Liang, Zhaoyang Yu 0003, Gang Wang 0001, Xiaoguang Liu 0001
ICDM1
2019 HDL: Hierarchical Deep Learning Model based Human Activity Recognition using Smartphone Sensors
abstract
With the development and popularization of smart-phones, human activity recognition methods based on contact perception are proposed. The smartphones which are embedded with various sensors can be used as a platform of mobile sensing for human activity recognition. In this paper, we propose an automated human activity recognition network HDL with smartphone motion sensor units. The HDL network combines DBLSTM (Deep Bidirectional Long Short-Term Memory) model and CNN (Convolutional neural network) model. The DBLSTM model is first used to model long sequence data and ultimately generate a bidirectional output vector in a abstract way. The DBLSTM model is good at dealing with serialization tasks but poor in the ability to extract features. Hence, the CNN model is then used to extract features from the abstract vector. Finally, the output layer employs a softmax function to classify human activities. We conduct experiments on the Public domain UCI dataset. The experimental results show that the proposed HDL network achieves reliable results with accuracy and F1 score as high as 97.95% and 97.27%. Compared with other networks based on the same smartphone dataset, the accuracy of HDL is higher than S-LSTM and Dropout CNN network by 2.14% and 6.97% respectively.
Tongtong Su, Huazhi Sun, Chunmei Ma, Lifen Jiang, Tongtong Xu
IJCNN1