Siyang Sun

dblp:212/6037 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 OMTS: Ordered Multipath Traffic Scheduling for Elephant Flows in Distributed AI Training Clusters
Siyang Sun, Qiang Wu 0018, Ran Wang 0004, Jie Hao 0002
WCNC1
2025 Diffusion Model-aided Resource Scheduling for Multiple GAI Training Jobs
abstract
With the prosperity of AI-Generated Content (AIGC), efficiently scheduling multiple Generative AI (GAI) distributed training jobs in a computing cluster has become crucial for pursuing higher cost-effectiveness. However, the resource-intensive nature and frequent communication demands of distributed training exacerbate resource fragmentation and network contention, resulting in low utilization and high latency. To this end, we propose an intelligent and dynamic resource scheduling method. Firstly, we propose an innovative scheduling analytical model that describes heterogeneous computing resources, communication contention, and the parameter synchronization architecture. We then formulate it as a multi-objective optimization problem. Next, we propose a Diffusion Model-based AI-generated Resource Scheduling (DARS) algorithm, to capture dynamic and high-dimensional environment and generate the optimal scheduling decisions. Finally, the policy network of deep reinforcement learning (DRL) is replaced with the proposed DARS to address the environmental uncertainty and enhance efficiency. Simulation results demonstrate that our proposed algorithm outperforms associated algorithms.
Qiang Wu 0018, Xiangbin Wang, Siyang Sun
ICCCN4
2025 Aligned Better, Listen Better for Audio-Visual Large Language Models
abstract
Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric settings. However, existing Video-LLMs and Audio-Visual Large Language Models (AV-LLMs) exhibit deficiencies in exploiting audio information, leading to weak understanding and hallucinations. To solve the issues, we delve into the model architecture and dataset. (1) From the architectural perspective, we propose a fine-grained AV-LLM, namely Dolphin. The concurrent alignment of audio and visual modalities in both temporal and spatial dimensions ensures a comprehensive and accurate understanding of videos. Specifically, we devise an audio-visual multi-scale adapter for multi-scale information aggregation, which achieves spatial alignment. For temporal alignment, we propose audio-visual interleaved merging. (2) From the dataset perspective, we curate an audio-visual caption \& instruction-tuning dataset, called AVU. It comprises 5.2 million diverse, open-ended data tuples (video, audio, question, answer) and introduces a novel data partitioning strategy. Extensive experiments show our model not only achieves remarkable performance in audio-visual understanding, but also mitigates potential hallucinations.
Shuailei Ma, Shijie Ma, Xiaoyi Bao, Chen-Wei Xie, Kecheng Zheng, Tingyu Weng, Siyang Sun
ICLR8
2025 Resilience-Driven Task-Cluster Co-Management: Proactive Mitigation of Co-Resident Threats in AI Clusters
abstract
With the prosperity of AI-generated content (AIGC), multitenant training in AI task clusters has become prevalent. To improve resource utilization, multiple tenants will coexist on the same server, while malicious tenants may exploit side-channel to pose significant co-resident eavesdropping risks. Due to the extensive attack surface, distributed training tasks are particularly vulnerable to model parameters and data leakage when subjected to the same level of security protection as inference tasks. Moreover, traditional security mechanisms, reliant on static encryption or isolation, suffer from high overhead, passive defense and poor scalability, failing to address the dynamic resilience requirements of AI clusters. However, the research on proactive resilience enhancement in AI clusters is almost blank. To fill this gap, we devise a grouping task migration mechanism (GTMM), which jointly considers adaptive server grouping and proactive task migration. Specifically, we first employ an adaptive server grouping algorithm to classify servers, offering customized protection based on tenants’ security requirements. Then, we formulate the task scheduling process as a multiobjective optimization problem for making a tradeoff between security, power consumption, and load balance. Next, we propose a deep reinforcement learning-based task migration algorithm to separate tenants that have completed co-residency for mitigating co-resident threats. Lastly, the simulation experiments demonstrate that GTMM’s security outperforms the baselines with only an affordable performance degradation.
Xiangbin Wang, Qiang Wu 0018, Ran Wang 0004, Siyang Sun
IEEE Internet Things J.5
2025 Optimal Secure NOMA Clustering and Power Allocation in Distributed Satellite-Enabled Internet of Things
abstract
The satellite-enabled Internet of Things (S-IoT) plays a crucial role by providing stable and global connectivity. However, its rapid growth brings many challenges in managing massive devices and addressing security threats. In this paper, we propose a new distributed network architecture combined with non-orthogonal multiple access (NOMA) for S-IoT. We consider a secure NOMA transmission scenario, where an appropriate legitimate device in an NOMA cluster is chosen as a jammer. To maximize the system’s total secrecy rate, we formulate an optimal secure NOMA clustering and power allocation, which is non-convex and difficult to solve directly. To solve the joint optimization problem, the original problem is transformed into two subproblems, and we propose staged algorithms to solve them efficiently. Firstly, a distributed iterative NOMA clustering algorithm is proposed to iteratively group S-IoT devices into multiple NOMA clusters. Then, a particle swarm optimization (PSO)-based optimal secure power allocation (OSPA) algorithm and a soft actor critic (SAC)-based OSPA are proposed to allocate powers for intra-cluster devices. The PSO/SAC-based OSPA algorithm can be performed among different NOMA clusters in a parallel way, which greatly improves the efficiency of power allocation. Finally, an optimal dynamic jammer strategy is proposed to dynamically select an idle device to act as a jammer, greatly improving the security performance of the systems. The simulation results demonstrate the advantages of the proposed distributed algorithms and also show that the proposed scheme greatly outperforms the state-of-the-art schemes in terms of the secrecy rate.
Bo Zhao 0022, Ruotong Zhang, Zhiquan Liu 0001, Siyang Sun, Guangliang Ren, Haolin Zhu
IEEE Internet Things J.4
2025 RPKI-Based Location-Unaware Tor Guard Relay Selection Algorithms
abstract
Tor is a well-known anonymous communication tool, used by people with various privacy and security needs. Prior works have exploited routing attacks to observe Tor traffic and deanonymize Tor users. Subsequently, location-aware relay selection algorithms have been proposed to defend against such attacks on Tor. However, location-aware relay selection algorithms are known to be vulnerable to information leakage on client locations and guard placement attacks. Can we design a new location-unaware approach to relay selection while achieving the similar goal of defending against routing attacks? Towards this end, we leverage the Resource Public Key Infrastructure (RPKI) in designing new guard relay selection algorithms. We develop a lightweight Discount Selection algorithm by only incorporating Route Origin Authorization (ROA) information, and a more secure Matching Selection algorithm by incorporating both ROA and Route Origin Validation (ROV) information. Our evaluation results show an increase in the number of ROA-ROV matched client-relay pairs using our Matching Selection algorithm, reaching 48.47% with minimal performance overhead through custom Shadow simulations and benchmarking.
Zhifan Lu, Siyang Sun
Proc. Priv. Enhancing Technol.2
2024 Relevant Intrinsic Feature Enhancement Network for Few-Shot Semantic Segmentation
abstract
For few-shot semantic segmentation, the primary task is to extract class-specific intrinsic information from limited labeled data. However, the semantic ambiguity and inter-class similarity of previous methods limit the accuracy of pixel-level foreground-background classification. To alleviate these issues, we propose the Relevant Intrinsic Feature Enhancement Network (RiFeNet). To improve the semantic consistency of foreground instances, we propose an unlabeled branch as an efficient data utilization method, which teaches the model how to extract intrinsic features robust to intra-class differences. Notably, during testing, the proposed unlabeled branch is excluded without extra unlabeled data and computation. Furthermore, we extend the inter-class variability between foreground and background by proposing a novel multi-level prototype generation and interaction module. The different-grained complementarity between global and local prototypes allows for better distinction between similar categories. The qualitative and quantitative performance of RiFeNet surpasses the state-of-the-art methods on PASCAL-5i and COCO benchmarks.
Xiaoyi Bao, Siyang Sun
AAAI3
2024 CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-Training
abstract
Learning joint and coordinated features across modalities is essential for many audio-visual tasks. Existing pre-training methods primarily focus on global information, neglecting fine-grained features and positions, leading to suboptimal performance in dense prediction tasks. To address this issue, we take a further step towards region-aware audio-visual pre-training and propose CrossMAE, which excels in Cross-modality interaction and region alignment. Specifically, we devise two masked autoencoding (MAE) pretext tasks at both pixel and embedding levels, namely Cross-Conditioned Reconstruction and Cross-Embedding Reconstruction. Taking the visual modality as an example (the same goes for audio), in Cross-Conditioned Reconstruction, the visual modality reconstructs the input image pixels conditioned on audio Attentive Tokens. As for the more challenging Cross-Embedding Reconstruction, unmasked visual tokens reconstruct complete audio features under the guidance of Learnable Queries implying positional information, which effectively enhances the interaction between modalities and exploits fine-grained semantics. Experimental results demonstrate that CrossMAE achieves state-of-the-art performance not only in classification and retrieval, but also in dense prediction tasks. Furthermore, we dive into the mechanism of modal interaction and region alignment of CrossMAE, highlighting the effectiveness of the proposed components.
Siyang Sun, Shuailei Ma, Kecheng Zheng, Xiaoyi Bao, Shijie Ma
CVPR2
2024 CoReS: Orchestrating the Dance of Reasoning and Segmentation
Xiaoyi Bao, Siyang Sun, Shuailei Ma, Kecheng Zheng, Guosheng Zhao, Xingang Wang 0003
ECCV (18)2
2024 FuseTeacher: Modality-Fused Encoders are Strong Vision Supervisors
Chen-Wei Xie, Siyang Sun, Pandeng Li, Shuailei Ma
ECCV (48)2
2023 RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training
abstract
Contrastive Language-Image Pre-training (CLIP) is attracting increasing attention for its impressive zero-shot recognition performance on different down-stream tasks. However, training CLIP is data-hungry and requires lots of image-text pairs to memorize various semantic concepts. In this paper, we propose a novel and efficient framework: Retrieval Augmented Contrastive Language-Image Pre-training (RA-CLIP) to augment embeddings by online retrieval. Specifically, we sample part of image-text data as a hold-out reference set. Given an input image, relevant image-text pairs are retrieved from the reference set to enrich the representation of input image. This process can be considered as an open-book exam: with the reference set as a cheat sheet, the proposed method doesn't need to memorize all visual concepts in the training data. It explores how to recognize visual concepts by exploiting correspondence between images and texts in the cheat sheet. The proposed RA-CLIP implements this idea and comprehensive experiments are conducted to show how RA-CLIP works. Performances on 10 image classification datasets and 2 object detection datasets show that RA-CLIP outperforms vanilla CLIP baseline by a large margin on zero-shot image classification task (+12.7%), linear probe image classification task (+6.9%) and zero-shot ROI classification task (+2.8%).
Chen-Wei Xie, Siyang Sun, Deli Zhao, Jingren Zhou 0001
CVPR2
2023 Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization
abstract
Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without any bounding-box annotations, they struggle to achieve precise localization, especially for small objects, and suffer from blurry boundaries and false positives. Moreover, the naive semi-supervised method is poor in effectively utilizing the abundance of unlabeled audio-visual pairs. In this paper, we propose a novel Semi-Supervised Learning framework for AVSL, namely Dual Mean-Teacher (DMT), comprising two teacher-student structures to circumvent the confirmation bias issue. Specifically, two teachers, pre-trained on limited labeled data, are employed to filter out noisy samples via the consensus between their predictions, and then generate high-quality pseudo-labels by intersecting their confidence maps. The optimal utilization of both labeled and unlabeled data combined with this unbiased framework enable DMT to outperform current state-of-the-art methods by a large margin, with CIoU of $\textbf{90.4\%}$ and $\textbf{48.8\%}$ on Flickr-SoundNet and VGG-Sound Source, obtaining $\textbf{8.9\%}$ and $\textbf{9.6\%}$ improvements respectively, given only $3\%$ of data positional-annotated. We also extend our framework to some existing AVSL methods and consistently boost their performance. Our code is publicly available at https://github.com/gyx-gloria/DMT.
Shijie Ma, Hu Su, Siyang Sun
NeurIPS7
2023 TMML: Text-Guided MuliModal Product Location For Alleviating Retrieval Inconsistency in E-Commerce
abstract
Image retrieval system (IRS) is commonly used in E-Commerce platforms for a wide range of applications such as price comparison and commodity recommendation. However, customers may experience inconsistent retrieval problems. Although the retrieved image contains the query object, the main product of the retrieved image is not associated with the query product. This is caused by the wrong product instance location when building the product image retrieval library. We can easily determine which product is on sale through the hint of the title, so we propose Text-Guided MuliModal Product Location (TMML) to use additional product titles to assist in locating the actual selling product instance. We design a weakly-aligned region-text data collection method to generate region-text pseudo-label by utilizing the IRS and user behavior from the E-commerce platform. To mitigate the impact of data noise, we propose a Mutual-Aware Contrastive Loss. Our results show that the proposed TMML outperforms the state-of-the-art method GLIP [11] by 3.95% in top-1 precision on our multi-objects test set, and 2.53% error located images in AliExpress has been corrected, which greatly alleviates the retrieval inconsistencies in IRS.
Youhua Tang, Siyang Sun, Baoliang Cui, Haihong Tang
SIGIR3
2022 Two stage Multi-Modal Modeling for Video Interaction Analysis in Deep Video Understanding Challenge
abstract
Interaction understanding between different entities in human-centered movie video is receiving more and more attention. Recently, a deep video understanding (DVU) task is proposed to identify interactions between different person entities on scene level task. However, limited samples of DVU dataset and multiple complex interactions make it difficult. To tackle these problems, we propose a two stage multi-modal method to predict the interaction between entities. Specifically, scene segment is first divided into several sub-scene clips, meanwhile face trajectory and person trajectory are obtained through face tracking/recognition and skeleton-based person tracing. Then, we extract and jointly train multi-modal features in the same semantic space including face emotion feature, person entity feature, visual features, text features and audio features. Finally, zero-shot transfer model and multiple classification model are proposed to predict interactions together. The experimental results show that our method performs new state-of-the-art on the DVU dataset.
Siyang Sun
ACM Multimedia1
2022 Deeply Exploit Visual and Language Information for Social Media Popularity Prediction
abstract
Social media popularity prediction task is to predict future attractiveness of new posts, which could be applied for online advertising, social recommendation, and demand prediction. Existing methods have explored multiple feature types to model the popularity prediction, including user profile, tag, space-time, category, and others. However, images and texts of social media posts, as important and primary information, are usually used by simple or insufficient processing. In this paper, we propose a method to deeply exploit visual and language information to explore the attractiveness of posts. Specifically, images are parsed from multiple perspectives including multi-modal semantic representation, perceptual image quality, and scene analysis. Different word-level and sentence-level semantic embedding are extracted from all available language texts including title, tags, concept and category. It makes social media popularity modeling more reliable with the powerful visual and language representation. Experimental results demonstrate the effectiveness of exploiting visual and language information by the proposed method, and we achieve new state-of-the-art results on the SMP Challenge at ACM Multimedia 2022.
Jianmin Wu, Dangwei Li, Chen-Wei Xie, Siyang Sun
ACM Multimedia5
2022 Deep Video Understanding with a Unified Multi-Modal Retrieval Framework
abstract
In this paper, we propose a unified multi-modal retrieval framework to tackle two typical video understanding tasks, i.e., matching movie scenes and text descriptions, and scene sentiment classification. For the task of matching movie scenes and text descriptions, it is a natural multi-modal retrieval problem, while for the task of scene sentiment classification, the proposed framework aims at finding most related sentiment tag for each movie scene, which is also a multi-modal retrieval problem. By considering these two tasks as multi-modal retrieval problems, we propose a unified multi-modal retrieval framework, which can make full use of the models pre-trained on large scale multi-modal datasets, experiments show that it is critical for the tasks which have only hundreds of training examples. To further improve the performance on movie video understanding task, we also collect a large scale video-text dataset, which contains 427,603 movie-shot and text pairs. Experimental results validate the effectiveness of this dataset.
Chen-Wei Xie, Siyang Sun, Jianmin Wu, Dangwei Li
ACM Multimedia2
2021 Fashion Focus: Multi-modal Retrieval System for Video Commodity Localization in E-commerce
abstract
Nowadays, live-stream and short video shopping in E-commerce have grown exponentially. However, the sellers are required to manually match images of the selling products to the timestamp of exhibition in the untrimmed video, resulting in a complicated process. To solve the problem, we present an innovative demonstration of multi-modal retrieval system called ``Fashion Focus'', which enables to exactly localize the product images in the online video as the focuses. Different modality contributes to the community localization, including visual content, linguistic features and interaction context are jointly investigated via presented multi-modal learning. Our system employs two procedures for analysis, including video content structuring and multi-modal retrieval, to automatically achieve accurate video-to-shop matching. Fashion Focus presents a unified framework that can orientate the consumers towards relevant product exhibitions during watching videos and help the sellers to effectively deliver the products over search and recommendation.
Yanhao Zhang 0002, Qiang Wang 0054, Cheng Da, Siyang Sun
AAAI6
2019 Multi-Loss-Aware Channel Pruning of Deep Networks
abstract
Channel pruning, which seeks to reduce the model size by removing redundant channels, is a popular solution for deep networks compression. Existing channel pruning methods usually conduct layer-wise channel selection by directly minimizing the reconstruction error of feature maps between the baseline model and the pruned one. However, they ignore the feature and semantic distributions within feature maps and real contribution of channels to the overall performance. In this paper, we propose a new channel pruning method by explicitly using both intermediate outputs of the baseline model and the classification loss of the pruned model to supervise layer-wise channel selection. Particularly, we introduce an additional loss to encode the differences in the feature and semantic distributions within feature maps between the baseline model and the pruned one. By considering the reconstruction error, the additional loss and the classification loss at the same time, our approach can significantly improve the performance of the pruned model. Comprehensive experiments on benchmark datasets demonstrate the effectiveness of the proposed method.
Yiming Hu, Siyang Sun, Jiagang Zhu, Xingang Wang 0003, Qingyi Gu
ICIP2
2019 Robust Landmark Detection and Position Measurement Based on Monocular Vision for Autonomous Aerial Refueling of UAVs
abstract
In this paper, a position measurement system, including drogue's landmark detection and position computation for autonomous aerial refueling of unmanned aerial vehicles, is proposed. A multitask parallel deep convolution neural network (MPDCNN) is designed to detect the landmarks of the drogue target. In MPDCNN, two parallel convolution networks are used, and a fusion mechanism is proposed to accomplish the effective fusion of the drogue's two salient parts' landmark detection. Considering the drogue target's geometric constraints, a position measurement method based on monocular vision is proposed. An effective fusion strategy, which fuses the measurement results of drogue's different parts, is proposed to achieve robust position measurement. The error of landmark detection with the proposed method is 3.9%, and it is obviously lower than the errors of other methods. Experimental results on the two KUKA robots platform verify the effectiveness and robustness of the proposed position measurement system for aerial refueling.
Siyang Sun, Yingjie Yin, Xingang Wang 0003, De Xu
IEEE Trans. Cybern.1
2018 A Novel Markov-Based Temporal-SoC Analysis for Characterizing PEV Charging Demand
abstract
The integration of a massive number of plug-in electric vehicles (PEVs) into current power distribution networks brings direct challenges to network planning, control, and operation. To increase the PEV penetration level with minimal negative impact, the dynamical PEV travel behaviors and charging demand need to be better understood. This paper presents a Markov-based analytical approach for modeling PEV travel behaviors and charging demand. The travel behaviors of individual PEVs are expressed mathematically through Monte Carlo simulation considering two essential factors: temporal travel purposes and state of charge (SoC). Markov model and hidden Markov model (HMM) are adopted to explicitly formulate the probabilistic correlation between multiple PEV states and SoC ranges. This modeling approach provides an efficient and generic tool for analyzing PEV travel behaviors and charging demand based on available PEV statistics. The analytical model is further adopted in the impact assessment of two PEV normal charging scheduling strategies for a range of PEV penetration levels in an IEEE 53-bus test network with field data (network parameters and realistic PEV statistics). The results demonstrate the benefit of the proposed modeling approach in network analysis considering PEV integration.
Siyang Sun, Qiang Yang 0004
IEEE Trans. Ind. Informatics1