Shang Gao 0003

dblp:28/435-3 · DBLP profile ↗
← Back
90ranked-venue papers
5as first author
73since 2021 · last 2026
0000-0002-2947-7780ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 22 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 17 since 2021Systems, architecture and hardware · 20 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 5 since 2021Security and privacy · 6 · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author
YearPublicationVenuePosition
2026 A fine-grained task scheduling strategy for resource auto-scaling over fluctuating data streams
Yinuo Fan, Dawei Sun 0001, Minghui Wu 0003, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.4
2026 A multi-domain cooperative scheduling framework for distributed stream computing systems
Dawei Sun 0001, Yueru Wang, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.4
2026 Poisoning-based Link Inference Attacks Against Federated Graph Neural Networks
abstract
Federated graph neural networks (FedGNNs) have emerged as a promising solution for handling graph data distributed across multiple owners. They enable collaborative training while preserving data decentralisation and complying with privacy and regulatory constraints. However, the inherent structural dependencies in graph data and the message-passing mechanisms of GNNs introduce both cross-client and intra-client edges in FedGNNs. Cross-client edges, in combination with federated learning (FL) protocol designs, open additional channels for information propagation and heighten the risk of privacy leakage. In FedGNNs, once edge information is compromised, adversaries can infer local neighbourhood structures and reconstruct inter-client relationships, even without direct access to raw data. Existing research on privacy inference in FL has largely overlooked edge privacy threats specific to FedGNNs. To address this gap, we propose a poisoning link inference approach with two strategies: Label Flipping Link Inference Attack (LFLIA) and Gradient Ascent Link Inference Attack (GALIA). LFLIA flips the label of a candidate node so that its perturbation propagates along structural topology during training. GALIA perturbs the candidate node’s gradient to amplify its loss. The perturbations on the candidate node can propagate to its linked neighbours by message-passing mechanism, which induces representation shifts on these linked nodes. By monitoring FedGNN outputs of a target node set before and after poisoning, an adversary can distinguish linked nodes through observable output shifts, whereas unlinked nodes exhibit little to no change. Experimental results on multiple benchmark datasets show that our poisoning-based LIA can effectively infer link existence and structure with high accuracy across diverse federated settings.
Guizhen Yang, Yanjun Zhang 0002, Leo Yu Zhang, Mengmeng Ge 0001, Shang Gao 0003
Proc. Priv. Enhancing Technol.5
2026 A Popularity-Aware Discriminative Grouping Strategy in Distributed Stream Computing Systems
abstract
Stream grouping strategy plays an important role in stateful stream computing environments. Many existing grouping strategies overlook various cost factors associated with grouping while balancing stream load. To overcome this limitation, we propose Pd-Stream, a popularity-aware discriminative grouping strategy that identifies the hot keys in dynamic real-time streams and assigns them to instances with high balance and low cost. Our solution includes: (1) A stream application model is constructed, along with a skewed data stream model and a data stream grouping model. Data stream grouping optimization problems are formalized. (2) A hot key probability estimation algorithm is designed, which estimates real-time probabilities of hot keys based on their popularity within the sampling window. (3) An instance assignment algorithm is designed using dynamic routing. This algorithm determines the minimal number of candidate instances based on the probabilities of hot keys, and selects the target instance with the lowest load through a dynamic routing table. Experimental results show that Pd-Stream provides near-optimal load balancing with low memory, achieving load imbalance as low as$10^{-5}$and replication factor as low as 1.74. It outperforms state-of-the-art works, reducing latency by 27%–46% and improving throughput by 23%–52%.
Dawei Sun 0001, Minghui Wu 0003, Shang Gao 0003, Rajkumar Buyya
IEEE Trans. Mob. Comput.4
2026 Skewness-Aware Stream Partitioning: A Key Splitting Suppression Method for Distributed Stream Processing
Dawei Sun 0001, Weilong Lv, Shang Gao 0003, Keqin Li 0001, Rajkumar Buyya
IEEE Trans. Serv. Comput.3
2025 Sketch-based Point Cloud Generation with Diffusion Model and Pre-training Enhancement
abstract
Diffusion models, known for their success in various generative tasks like image generation and super-resolution, are applied in this study for point cloud generation, a field that has not been extensively explored due to the complexity of point clouds. We propose a novel method using a diffusion model to generate high-quality 3D point clouds from 2D sketches. This method employs a self-supervised contrastive learning scheme to align sketch and point cloud modalities. Additionally, it incorporates a specific partition mixing strategy to integrate edge information during pre-training. Evaluated on two benchmark datasets, our method outperforms existing state-of-the-art approaches, showcasing the potential of diffusion models in point cloud generation and setting a new direction for future research.
Yangdong Chen, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP6
2025 Minimizing Disparities between Real and Pseudo Queries for Unsupervised Visual Grounding
abstract
Visual grounding involves the identification and localization of image regions given textual descriptions. To reduce the manual labeling effort on region-text pairs, unsupervised visual grounding aims to generate pseudo bounding box and query pairs for training grounding models. However, there exists significant disparities between real and pseudo queries in terms of object, attribute distributions, and textual formats, limiting the generalization performance of unsupervised grounding methods. To address this challenge, we propose a novel unsupervised visual grounding framework. During training, we prompt Multimodal Large Language Models to generate pseudo queries, in which the entities are beyond the object detector’s pre-defined limited categories, and are associated with richer attributes. We further devise a Modifier Tree structure to bridge the gap of textual format between real and pseudo queries. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art unsupervised approaches on public benchmark datasets, particularly when dealing with complex queries.
Changkai Ji, Jilan Xu, Yanhao Zhu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP8
2025 Human Simulacra: Benchmarking the Personification of Large Language Models
abstract
Large Language Models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted the attention of the social science community, who see the potential in leveraging LLMs to replace human participants in experiments, thereby reducing research costs and complexity. In this paper, we introduce a benchmark for LLMs personification, including a strategy for constructing virtual characters' life stories from the ground up, a Multi-Agent Cognitive Mechanism capable of simulating human cognitive processes, and a psychology-guided evaluation method to assess human simulations from both self and observational perspectives. Experimental results demonstrate that our constructed simulacra can produce personified responses that align with their target characters. We hope this work will serve as a benchmark in the field of human simulation, paving the way for future research.
Qiujie Xie, Qiming Feng, Qingqiu Li, Linyi Yang, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003, Yue Zhang 0004
ICLR9
2025 A Prediction-Driven Collaborative Scheduling Strategy for Distributed Stream Computing Systems
abstract
Multi-objective collaborative optimization is essential for improving performance in stream computing systems. However, existing approaches often neglect the interdependencies among communication overhead, load balancing, and energy consumption, and lack predictive capabilities, resulting in delayed scheduling decisions that degrade system latency and throughput. To overcome these limitations, we propose a prediction-driven collaborative framework, named Pc-Stream, which proactively identifies overloaded compute nodes and triggers task migrations in advance. This paper presents this strategy through two key components: (1) A temperature-driven neighborhood adjustment method for task topology partitioning. This method dynamically adjusts the number of migrated tasks based on a predefined temperature. Tasks with high communication volume are batchmigrated to nodes with lower utilization rates during the hightemperature phase, and migrated individually during the lowtemperature phase. (2) A sliding window mechanism that generates multiple sub-sequences for training multiple predictive models. These models enable the system to monitor load trends and proactively migrate tasks from overloaded nodes to those with sufficient resources, thereby reducing communication costs and improving load balance. Experimental results demonstrate that, under dynamic and fluctuating data stream conditions, Pc-Stream significantly enhances overall system performance: reducing average system latency by 49.9 %, and increasing average throughput by 16.9 %.
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
ICPADS4
2025 TGSAM-2: Text-Guided Medical Image Segmentation Using Segment Anything Model 2
Runtian Yuan, Ling Zhou 0002, Jilan Xu, Qingqiu Li, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
MICCAI (10)9
2025 Text-Promptable Propagation for Referring Medical Image Sequence Segmentation
abstract
Referring Medical Image Sequence Segmentation (Ref-MISS) is a novel and challenging task that aims to segment anatomical structures in medical image sequences (e.g., endoscopy, ultrasound, CT, and MRI) based on natural language descriptions. Existing 2D and 3D segmentation models struggle to explicitly track objects of interest across medical image sequences, and lack support for interactive, text-driven guidance. To address these limitations, we propose Text-Promptable Propagation (TPP), which enables the recognition of referred objects through cross-modal referring interaction, and maintains continuous tracking across the sequence via Transformer-based triple propagation, using text embeddings as queries. To support this task, we curate a large-scale benchmark, Ref-MISS-Bench, which covers 4 imaging modalities and 20 different organs and lesions. Experimental results on this benchmark demonstrate that TPP consistently outperforms state-of-the-art methods in both medical segmentation and referring video object segmentation. Code and data are available at https://github.com/yuanruntian/TPP.
Runtian Yuan, Mohan Chen 0001, Jilan Xu, Ling Zhou 0002, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ACM Multimedia9
2025 EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues
abstract
Qiming Feng, Qiujie Xie, Xiaolong Wang, Qingqiu Li, Yuejie Zhang, Rui Feng, Tao Zhang, Shang Gao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qiming Feng, Qiujie Xie, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
NAACL (Long Papers)8
2025 AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation
abstract
Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding, current MLMMs still face two major challenges: (1) insufficient region-level understanding and interaction, and (2) limited accuracy and interpretability due to single-step prediction. In this paper, we address these challenges by empowering MLMMs with anatomy-centric reasoning capabilities to enhance their interactivity and explainability. Specifically, we propose an Anatomical Ontology-Guided Reasoning (AOR) framework that accommodates both textual and optional visual prompts, centered on region-level information to enable multimodal multi-step reasoning. We also develop AOR-Instruction, a large instruction dataset for MLMMs training, under the guidance of expert physicians. Our experiments demonstrate AOR's superior performance in both Visual Question Answering (VQA) and report generation tasks. Code and data are available at: https://github.com/Liqq1/AOR.
Qingqiu Li, Zihang Cui, Seongsu Bae, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Quanli Shen, Shang Gao 0003, Junjun He
NeurIPS10
2025 Arms Race in Deep Learning: A Survey of Backdoor Defenses and Adaptive Attacks
Xiaoxing Mo, Nan Sun 0002, Leo Yu Zhang, Wei Luo 0001, Shang Gao 0003, Yong Xiang 0001
PAKDD (4)5
2025 Toward High-Availability Distributed Stream Computing Systems via Checkpoint Adaptation
abstract
ABSTRACT The importance of fault tolerance strategies for distributed streaming computing systems becomes more evident due to the increased diversity of failures. Checkpointing is considered a general and efficient method for ensuring fault tolerance. However, determining the checkpoint interval poses a challenge: shorter checkpoint intervals lead to higher overhead, while longer intervals result in extended fault recovery time. Therefore, optimizing the checkpoint interval becomes crucial for the efficient operation of streaming applications. There has been relatively limited exploration and analysis of optimal checkpoint interval settings in the context of stream computing. Many existing works considered adjusting this interval based on a single factor. This article proposes a checkpoint adaptive strategy with high availability, named Ca‐Stream. It considers multiple factors when adjusting checkpoint intervals. Specifically, it addresses the following aspects: (1) Using linear regression to predict the system's fault rate and dynamically adjusting the checkpoint interval based on these predictions. (2) Monitoring CPU time and memory consumption per task to dynamically trigger checkpoints, achieving high reliability, especially in resource‐constrained scenarios. (3) Detecting task execution times on nodes and volume of input data for tasks to identify slow tasks within the cluster. Experiments conducted on a Flink system demonstrate Ca‐Stream's benefits. It reduces checkpoint consumption time by over 38%, system recovery latency by 33%, CPU occupancy by up to 47%, and memory occupancy by 37% compared to Flink's approaches.
Dawei Sun 0001, Jia Peng, Jonathan Kua, Shang Gao 0003, Rajkumar Buyya
Concurr. Comput. Pract. Exp.5
2025 Scene-cGAN: A GAN for underwater restoration and scene depth estimation
abstract
Despite their wide scope of application, the development of underwater models for image restoration and scene depth estimation is not a straightforward task due to the limited size and quality of underwater datasets, as well as variations in water colours resulting from attenuation, absorption and scattering phenomena in the water column. To address these challenges, we present an all-in-one conditional generative adversarial network (cGAN) called Scene-cGAN. Our cGAN is a physics-based multi-domain model designed for image dewatering, restoration and depth estimation. It comprises three generators and one discriminator. To train our Scene-cGAN, we use a multi-term loss function based on uni-directional cycle-consistency and a novel dataset. This dataset is constructed from RGB-D in-air images using spectral data and concentrations of water constituents obtained from real-world water quality surveys. This approach allows us to produce imagery consistent with the radiance and veiling light corresponding to representative water types. Additionally, we compare Scene-cGAN with current state-of-the-art methods using various datasets. Results demonstrate its competitiveness in terms of colour restoration and its effectiveness in estimating the depth information for complex underwater scenes. • We present an all-in-one conditional generative adversarial network for underwater computer vision. • Our approach addresses underwater restoration and depth estimation using three generators. • We present an underwater physics-based benchmark dataset. • Experimental results show our approach is competitive with other restoration methods and provides accurate depth estimations, aligning well with the scene.
Salma P. González-Sabbagh, Antonio Robles-Kelly, Shang Gao 0003
Comput. Vis. Image Underst.3
2025 Straggler mitigation via hierarchical scheduling in elastic stream computing systems
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.3
2025 Morality-Driven Mechanism Design: Application in Hierarchical Carbon Trading Markets
Ruhan Liu, Yao Zhang 0005, Youyang Qu, Longxiang Gao, Yong Xiang 0001, Shang Gao 0003, Tom H. Luan
IEEE Internet Things J.6
2025 Ls-Stream: Lightening Stragglers in Join Operators for Skewed Data Stream Processing
abstract
Load imbalance can lead to the emergence of stragglers, i.e., join instances that significantly lag behind others in processing data streams. Currently, state-of-the-art solutions are capable of balancing the load between join instances to mitigate stragglers by managing hot keys and random partitioning. However, these solutions rely on either complicated routing strategies or resource-inefficient processing structures, making them susceptible to frequent changes in load between instances. Therefore, we present Ls-Stream, a data stream scheduler that aims to support dynamic workload assignment for join instances to lighten stragglers. This paper outlines our solution from the following aspects: (1) The models for partitioning, communication, matrix, and resource are developed, formalizing problems like imbalanced load between join instances and state migration costs. (2) Ls-Stream employs a two-level routing strategy for workload allocation by combining hash-based and key-based data partitioning, specifying the destination join instances for data tuples. (3) Ls-Stream also constructs a fine-grained model for minimizing the state migration cost. This allows us to make tradeoffs between data transfer overhead and migration benefits. (4) Experimental results demonstrate significant improvements made by Ls-Stream: reducing maximum system latency by 49.3% and increasing maximum throughput by more than 2x compared to existing state-of-the-art works.
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Keqin Li 0001, Rajkumar Buyya
IEEE Trans. Computers3
2025 An elastic reconfiguration strategy for operators in distributed stream computing systems
Dawei Sun 0001, Yinuo Fan, Chengjun Guan, Jia Rong, Shang Gao 0003, Rajkumar Buyya
J. Supercomput.5
2025 A Hierarchical Near-Source Grouping Strategy for Elastic Stream Computing Systems
abstract
Effective task scheduling in stream computing systems can reduce the latency by minimizing inter-node communication. However, this approach often requires restarting tasks to change their deployment locations, resulting in significant system overhead and making it inadequate especially in dynamically changing data stream environments. To address this issue, we propose Ns-Stream, a hierarchical data scheduler that dynamically adjusts data distribution weights between near-source and off-source tasks. Our solution includes: (1) We observe that communication overhead from off-source data processing significantly impacts system latency when tasks' resources are ample. However, as the resources become limited, the computational power required by tasks becomes the key constraint on system performance. (2) During initialization scheduling, we deploy tasks with potential communication to the same node using the graph convolutional network, thus avoiding the need for runtime task scheduling. (3) We dynamically adjust data distribution weights between near-source and off-source tasks based on their computing capabilities, prioritizing local processing of data tuples (within the same worker and node) to optimize resource utilization and reduce data transmission overhead. (4) Experimental results demonstrate significant improvements made by Ns-Stream: reducing maximum system latency by 40% and increasing maximum throughput by 55% compared to existing state-of-the-art works.
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
IEEE Trans. Serv. Comput.3
2024 3D Face Recognition with Contrastive Learning Network on Low-Quality Data
Yaping Jing, Ajmal Mian, Leo Zhang, Shang Gao 0003, Xuequan Lu
CGI (1)4
2024 A Task Dependency-Aware Scheduling Strategy for Cross-Domain Stream Computing Environments
abstract
In cross-domain stream computing, assigning highly dependent tasks to different domains causes poor performance. Existing methods ignore cross-domain and focus on load balancing and resource allocation. To address this scheduling challenge, this paper proposes a task dependency-aware scheduling strategy named Td-Stream. This strategy is discussed in the following aspects: (1) Impact analysis: Analyzing the adverse impact of communication dependencies between tasks on system performance under traditional scheduling methods in cross-domain environments. (2) Model construction: Constructing models for stream topology, task dependency, resource cost.(3) Cross-domain task allocation: Introducing a cross-domain dependent task allocation method that incorporates a resource elasticity mechanism. Experimental results demonstrate significant improvements made by Td-Stream compared to existing state-of-theart works.
Dawei Sun 0001, Yueru Wang, Shang Gao 0003, Rajkumar Buyya
HPCC4
2024 Fine-Granularity Face Sketch Synthesis
abstract
Generative Adversarial Networks (GANs) are often used in face sketch synthesis due to their powerful ability in image generation. However, most GAN based synthesis methods took the entire face as the minimum unit. Differently, we propose a novel fine-granularity face sketch synthesis framework in this paper. The core idea is to first capture local information at a fine granularity (i.e., facial component), and then generate a complete face sketch based on the fine-grained information. Specifically, we partition the face sketch into multiple components, and then train a parallel network for each component. A condition enhanced detail repair network is further designed to correct the mismatches and deformations produced during parallel generation. Extensive experiments show that our approach outperforms state-of-the-art methods from both the qualitative and quantitative perspectives.
Yangdong Chen, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICASSP7
2024 ControlCap: Controllable Captioning via No-Fuss Lexicon
abstract
Controllable captioning has received much attention in recent years. Although substantial progress has been made, existing methods still face challenges such as high training costs, intricate control signals and limited control capabilities. To address these issues, we propose a straightforward and unified framework called ControlCap. It uses a no-fuss lexicon as control signal and controls the style and content of visual descriptions through Soft Guidance (a global guide to the caption distribution) and Hard Force (integrating signals without additional training). Extensive experiments, both quantitative and qualitative, have been conducted on three benchmark captioning tasks. Results demonstrate the control ability of ControlCap: it can produce controlled captions that are coherent and diverse while keeping the core content intact.
Qiujie Xie, Qiming Feng, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP6
2024 Exploring Object-Centered External Knowledge for Fine-Grained Video Paragraph Captioning
abstract
Video paragraph captioning task aims to generate a detailed, fluent and relevant paragraph for a given video. Prior studies often focus on isolating visual objects (potential main components in a sentence) from the overall video content. They rarely explore the latent semantic relations between objects and high-level video concepts, resulting in dull or even incorrect descriptions. To create fine-grained and contextually relevant paragraph captions, we propose a novel framework that constructs a concept graph from a commonsense knowledge base and infers richer semantic meaning from the visual objects. Moreover, we employ a Vision-Guided Concept Selection Network that incorporates an under-sentence supervision mechanism to align the external knowledge with the visual information. Through extensive experiments on ActivityNet captions and YouCook2, the effectiveness of our method is demonstrated compared to state-of-the-art methods.
Guorui Yu, Yimin Hu, Yiqian Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP7
2024 Temporal Feature Aggregation for Efficient 2D Video Grounding
abstract
Video grounding aims to locate the target video moment in an untrimmed video based on a text query. Most existing methods employ 3D CNNs as the video feature extractor, incurring substantial computational costs. Only a few methods use 2D backbones for video feature extraction, and they suffer from diminished accuracy due to the inherent lack of temporal information within 2D features. To address this problem, we propose a novel 2D video grounding method called TFA that improves accuracy while minimizing computational costs. Our approach involves a query-guided temporal feature aggregation module designed to explicitly capture temporal information. We disentangle time intervals of input video frames and prediction spans to reduce computational overhead. Additionally, we introduce deformable attention into the multi-modal encoder for further enhancement. Extensive experiments on two public datasets demonstrate that our method outperforms previous 2D video grounding methods and achieves competitive results with most 3D methods at significantly reduced costs.
Mohan Chen 0001, Yiren Zhang, Jueqi Wei, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICME7
2024 Memory-Augmented Transformer for Efficient End-to-End Video Grounding
abstract
Video grounding aims to localize a specific segment corresponding to a text query in an untrimmed video. Due to the tremendous computational cost required to process the video frames, the de facto paradigm of video grounding is to extract video features using pretrained video encoders. The parameters of the video encoders are fixed during training, which limits the performance of the localization model. To solve this problem, we propose a Memory-Augmented Transformer (MAT) model. Specifically, each video is split into non-overlapping clips, and our MAT processes videos in a clip-by-clip manner while caching video features into FIFO cached memory queues. By enabling early return, our MAT outperforms previous methods with only less than 60% frames seen. Extensive experimental results on three public benchmark datasets demonstrate that our MAT can achieve competitive performance while being much more efficient than currently prevailing two-stage methods. Code is available at https://github.com/xuyw1997/MAT.
Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICME6
2024 Point Cloud Normal Estimation via Representation Learning on Height Maps
abstract
Point Cloud Normal Estimation via Representation Learning on Height Maps
Dasith de Silva Edirimuni, Ye Zhu 0002, Shang Gao 0003, Zhiyong Wang 0001, Antonio Robles-Kelly, Xuequan Lu
MMAsia4
2024 Robust Backdoor Detection for Deep Learning via Topological Evolution Dynamics
abstract
A backdoor attack in deep learning inserts a hidden backdoor in the model to trigger malicious behavior upon specific input patterns. Existing detection approaches assume a metric space (for either the original inputs or their latent representations) in which normal samples and malicious samples are separable. We show that this assumption has a severe limitation by introducing a novel SSDT (Source-Specific and Dynamic-Triggers) backdoor, which obscures the difference between normal samples and malicious samples.To overcome this limitation, we move beyond looking for a perfect metric space that would work for different deep-learning models, and instead resort to more robust topological constructs. We propose TED (Topological Evolution Dynamics) as a model-agnostic basis for robust backdoor detection. The main idea of TED is to view a deep-learning model as a dynamical system that evolves inputs to outputs. In such a dynamical system, a benign input follows a natural evolution trajectory similar to other benign inputs. In contrast, a malicious sample displays a distinct trajectory, since it starts close to benign samples but eventually shifts towards the neighborhood of attacker-specified target samples to activate the backdoor.Extensive evaluations are conducted on vision and natural language datasets across different network architectures. The results demonstrate that TED not only achieves a high detection rate, but also significantly outperforms existing state-of-the-art detection approaches, particularly in addressing the sophisticated SSDT attack. The code to reproduce the results is made public on GitHub.
Xiaoxing Mo, Yechao Zhang, Leo Yu Zhang, Wei Luo 0001, Nan Sun 0002, Shengshan Hu, Shang Gao 0003, Yang Xiang 0001
SP7
2024 Lc-Stream: An elastic scheduling strategy with latency constraints in geo-distributed stream computing environments
abstract
Summary An effective scheduling strategy is critical for achieving better performance in real‐time stream processing systems. How to quickly and efficiently process real‐time data stream is always challenging, especially when clusters are collaborating in a Geo‐Distributed computing environment. To address these challenges, we propose an elastic scheduling strategy with Latency Constraints in Geo‐Distributed stream computing environments called Lc‐Stream. This article discusses our work from the following aspects: (1) An optimized data stream redirection method that is proposed based on queuing network algorithm, along with a computing resource model, a latency constrained scheduling model and a communication energy consumption model. (2) An updated node selection method based on the inter‐layer task correlation, to reduce the communication latency between groups at the executor granularity. (3) A network cluster distribution for Geo‐Distributed computing environment to ensure energy saving under low transmission latency. Experimental results show that compared to R‐Storm, Lc‐Stream reduces total latency by over 19% and increases throughput by over 37% in typical cross‐domain multi‐task topologies. Compared to Ts‐Stream, Lc‐Stream also reduces total latency by over 15% and increases throughput by over 21%. At the same time, it helps to balance the load among the systems and avoid overuse of compute nodes.
Dawei Sun 0001, Yueru Wang, Jialiang Sui, Shang Gao 0003, Jia Rong, Rajkumar Buyya
Concurr. Comput. Pract. Exp.4
2024 Orchestrating scheduling, grouping and parallelism to enhance the performance of distributed stream computing system
Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
Expert Syst. Appl.3
2024 An adaptive load balancing strategy for stateful join operator in skewed data stream environments
Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.3
2024 DGD-cGAN: A dual generator for image dewatering and restoration
Salma P. González-Sabbagh, Antonio Robles-Kelly, Shang Gao 0003
Pattern Recognit.3
2024 Elastic Scaling of Stateful Operators Over Fluctuating Data Streams
abstract
Elastic scaling of parallel operators has emerged as a powerful approach to reduce response time in stream applications with fluctuating inputs. Many state-of-the-art works focus on stateless operators and change the operator parallelism from one aspect. They often lack efficient management of operator states and overlook the costs associated with resource over-provisioning. To overcome these limitations, we introduce Es-Stream for elastic scaling of stateful operators over fluctuating data streams, which includes: 1) We observe that under-provisioning of operator parallelism leads to data pile-up, resulting in longer system latency, while over-provisioning of operator parallelism causes idle instances and additional resource consumption. 2) The Es-Stream system scales in two dimensions: the parallelism of operators and the number of resources. It dynamically adjusts operators to an optimal parallelism while scaling the resources used by the stream application. 3) When the parallelism of stateful operators changes, upstream operators backup downstream operators’ state and cache the emitted data tuples at dynamic time intervals, ensuring the operator parallelism is adjusted in a low-overhead way. 4) Experimental results demonstrate that Es-Stream provides promising performance improvements, reducing the maximum system latency by 3x and saving the maximum state recovery time by 2x, compared to existing state-of-the-art works.
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Keqin Li 0001, Rajkumar Buyya
IEEE Trans. Serv. Comput.3
2023 Enhanced Knowledge Injection for Radiology Report Generation
abstract
Automatic generation of radiology reports holds crucial clinical value, as it can alleviate substantial workload on radiologists and remind less experienced ones of potential anomalies. Despite the remarkable performance of various image captioning methods in the natural image field, generating accurate reports for medical images still faces challenges, i.e., disparities in visual and textual data, and lack of accurate domain knowledge. To address these issues, we propose an enhanced knowledge injection framework, which utilizes two branches to extract different types of knowledge. The Weighted Concept Knowledge (WCK) branch is responsible for introducing clinical medical concepts weighted by TF-IDF scores. The Multimodal Retrieval Knowledge (MRK) branch extracts triplets from similar reports, emphasizing crucial clinical information related to entity positions and existence. By integrating this finer-grained and well-structured knowledge with the current image, we are able to leverage the multi-source knowledge gain to ultimately facilitate more accurate report generation. Extensive experiments have been conducted on two public benchmarks, demonstrating that our method achieves superior performance over other state-of-the-art methods. Ablation studies further validate the effectiveness of two extracted knowledge sources.
Qingqiu Li, Jilan Xu, Runtian Yuan, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
BIBM8
2023 Semi-MedSeq: Semi-supervised Semantic Segmentation for Medical Image Sequences
abstract
In clinical practice, medical imaging techniques include 2D video-based examinations that capture sequential scans, and 3D volumetric imaging that forms a comprehensive 3D representation from a stack of 2D slices. The medical image sequences produced by the above techniques provide valuable spatio-temporal characteristics for analysis and segmentation, but the annotation of image sequences is extremely time-consuming and labor-intensive. To exploit the coherence and address the scarcity of labeled data, we propose a novel semi-supervised semantic segmentation framework for medical image sequences, which consists of a conditional network and a denoising network. Specifically, we embed a Sequential Feature Reconstruction module into both networks. This module reconstructs the target frame from contiguous frames and captures their shared visual features. Guided by the context-enhancing information from the conditioning network, the denoising network suppresses background noise via a Diffusion-based Noise Elimination module. Extensive experiments are conducted on 2D and 3D tasks, including cardiac segmentation, polyp segmentation, placenta vessel segmentation and abdomen multi-organ segmentation. The results show our method is superior to existing semi-supervised methods and exhibits advantages over fully-supervised medical image segmentation methods with only 1/2 labeled data, validating its effectiveness and generalization ability.
Runtian Yuan, Jilan Xu, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM8
2023 Motion-Aware Video Paragraph Captioning via Exploring Object-Centered Internal Knowledge
abstract
Video paragraph captioning task aims at generating a fine-grained, coherent and relevant paragraph for a video. Different from the images where objects are static, the temporal states of objects are changing in videos. The dynamic information could be contributed to understanding the whole video content. Existing works rarely put focus on modeling the dynamic changing state of the objects in the videos, causing the activities occurred in videos are poorly or wrongly depicted in paragraphs. To address this problem, we propose a novel Object State Tracking Network, which can capture the temporal state change of objects. However, due to the similarity of the consecutive frames in the videos, the information of the video is redundant and noisy. We further propose a semantic alignment mechanism, and enable the sentence information to refine the visual information. Extensive experiments on ActivityNet Captions demonstrate the effectiveness of our method.
Yimin Hu, Guorui Yu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICASSP7
2023 Video Captioning via Relation-Aware Graph Learning
abstract
Recent neural models for video captioning usually employed an encoder-decoder framework. However, most approaches either neglected the spatial and temporal interactions between objects in a video or implicitly modelled the interactions, resulting in less desired performance. In this paper, we propose a novel relation-aware graph learning framework. It explicitly models both spatial and temporal relations for objects. In particular, a relation-aware graph is designed to depict the spatial relations between different objects in a scene. Parallelly, a temporal graph network is designed to perform relational reasoning for the same objects in adjacent frames. Features of both types of relations are learned and fused for the follow-up language decoder. Experiments on two bench-mark datasets show the effectiveness of our framework. It achieves state-of-the-art performance with CIDEr scores on MSVD and MSR-VTT.
Heming Jing, Qiujie Xie, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
ICASSP7
2023 Conditional Video-Text Reconstruction Network with Cauchy Mask for Weakly Supervised Temporal Sentence Grounding
abstract
Temporal sentence grounding aims to detect the target segment most related to a given query in an untrimmed video. To alleviate the expensive annotation cost for temporal labels, researchers paid more attention to weakly supervised setting. Prior studies neglected the utilization of video representation reconstruction, which led to an unbalanced alignment learning. Moreover, they used different strategies to generate proposals which ignored the temporal structure in a query. In this paper, we propose a novel Conditional Video-Text Reconstruction Network (CVTRN). It supports conditional reconstruction of video and text representation. Specifically, video and text features are fused to compute semantic alignment, which is the condition of reconstruction. A new mask strategy for mask conditioned sentence reconstruction is also devised. This strategy focuses more on boundary regions than the widely used Gaussian mask in previous methods. Experimental results on two public benchmark datasets show that our CVTRN outperforms the state-of-the-art methods.
Jueqi Wei, Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
ICME6
2023 SPTNET: Span-based Prompt Tuning for Video Grounding
abstract
When a Pre-trained Language Model (PLM) is adopted in video grounding task, it usually acts as a text encoder without having its knowledge fully utilized. Also, there exists an inconsistency problem between the pre-training and downstream objectives. To solve the issues, we propose a new paradigm, named Span-based Prompt Tuning (SPTNet). It can convert the video grounding task into a cloze form. Specifically, a query is first changed into a form with mask token by a template, then the video and the query embeddings are integrated through a cross-modal transformer. The start and end points of the query matching time span are predicted with the embedding of the mask token. Experimental results on two public benchmarks ActivityNet Captions and Charades-STA show that our SPTNet achieves surpassing performance compared with state-of-the-art methods.
Yiren Zhang, Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003
ICME6
2023 A Frequency-aware Grouping Strategy for Stateful Operators in Distributed Stream Processing Systems
abstract
Current optimization for stateful flow processing computation tends to focus on load balancing without considering the utilization of downstream instance resources. To address this issue, we propose a data stream grouping method called Fa-Stream, specifically designed for stateful operators and incorporating field values frequency-awareness. Fa-Stream is implemented in three main aspects: (1) A data stream grouping model is built using Count-Min Sketch and Gated Recurrent Unit (GRU) to predict and analyze the frequency of field values. It selectively chooses high-frequency field values, and the communication distance model and instance resource constraint model are designed to adjust the weights of downstream instances for high-frequency field values. (2) A cyclic access routing table is generated, and weights are dynamically adjusted by a rebalancing scheme to avoid load skewness. Consistent hash grouping is implemented for low-frequency field values, and dual mapping is used to prevent large-scale migration caused by scaling. To validate the effectiveness of Fa-Stream, comparative experiments between Partial Key Grouping (PKG) and Fa-Stream are conducted using the Storm platform. Results demonstrate that Fa-Stream improves tuple throughput by 12.3%, reduces system delay by 14.2%, and increases load balancing degree by 42.8%. Furthermore, fa-Stream exhibits efficiency and stability across different data skews and tuple input rates.
Dawei Sun 0001, Weilong Lv, Shang Gao 0003, Jia Rong
ICPADS4
2023 A Latency Guaranteed Scheduling Strategy under Performance Constraints in Big Data Stream Computing Environments
abstract
Efficient utilization of computing resources in a stream computing environment is crucial for system performance. Existing scheduling strategies can hardly guarantee latency under performance constraints, let alone accounting for the communication cost incurred by scheduling itself. To address these issues, we propose Lg-Stream, a latency guaranteed scheduling strategy under performance constraints. This paper discusses the Lg-Stream strategy from the following aspects: (1) We model the topology as a queuing network to evaluate the system latency; (2) For scenarios with limited resources and latency constraints, we ensure that each executor-to-component allocation optimizes the system's processing latency to the maximum, consequently altering the components’ parallelism; (3) We place executors that communicate with each other on the same node as much as possible. Experimental results demonstrated that in comparison to existing state-of-the-art scheduling strategies, it reduces the average system latency up to 30%.
Dawei Sun 0001, Chengjun Guan, Yinuo Fan, Jia Rung, Shang Gao 0003
ICPADS5
2023 A Survey on Out-of-Distribution Evaluation of Neural NLP Models
abstract
Adversarial robustness, domain generalization and dataset biases are three active lines of research contributing to out-of-distribution (OOD) evaluation on neural NLP models. However, a comprehensive, integrated discussion of the three research lines is still lacking in the literature. This survey will 1) compare the three lines of research under a unifying definition; 2) summarize their data-generating processes and evaluation protocols for each line of research; and 3) emphasize the challenges and opportunities for future work.
Ming Liu 0028, Shang Gao 0003, Wray L. Buntine
IJCAI3
2023 Bio-Inspired Dual-Network Model to Tackle Statistical Heterogeneity in Federated Learning
abstract
The problem of statistical heterogeneity in Federated Learning has been a major challenge, with existing solutions making unrealistic assumptions about the availability of shared datasets and high bandwidth between clients and the server. Solving this problem is crucial for the success of Federated Learning in real-world scenarios. In this work, we propose a biologically inspired dual-network model FedDual, which mimics how the human brain learns and memorizes the information. The model consists of a neocortical and a hippocampal network similar to those in the human brain. The hippocampal network is comprised by an image classification model, while the neocortical network is a variational auto-encoder responsible for long-term and re-callable memory. In this manner, FedDual uses the neocortical network to generate pseudo-patterns (synthetic data) on the server (global model). This allows for the hippocampal network to be trained with these pseudo-patterns. The dual-network architecture allows devices to share information via the weight updates of the neocortical network to the server without sending the actual data. We compare FedDual against alternatives elsewhere in the literature when applied to widely available datasets. FedDual not only achieves a margin of accuracy improvement over the alternatives, but also converges faster, requiring less communication rounds.
Adnan Ahmad, Vinh Loi Chau, Antonio Robles-Kelly, Shang Gao 0003, Longxiang Gao, Lianhua Chi, Wei Luo 0001
IJCNN4
2023 Backdoor Attack on Deep Neural Networks in Perception Domain
abstract
As deep neural networks (DNNs) are widely deployed in various applications, the security of pretrained DNNs is crucial since backdoors can be introduced through poisoned training. A backdoored DNN model works properly when benign inputs are provided, but it produces targeted misclassification on the inputs with an intended pattern known as a trojan trigger. Current technologies for trigger generation mainly focus on the physical and model domains. In this work, we investigate trojan triggers from the perception domain, especially the physical process of collecting light rays when they pass through the lens and hit the optical sensors. A new type of backdoor attack, Lens Flare attack, is introduced. It concentrates on the perception domain and is more physically plausible and stealthy. Experiments show that the DNNs with Lens Flare backdoor can achieve accuracy comparable to their original counterpart on benign input while misclassifying the input with high certainty if the Lens Flare trigger is present. It is also demonstrated that the Lens Flare backdoor is resistant to state-of-the-art backdoor defenses.
Xiaoxing Mo, Leo Yu Zhang, Nan Sun 0002, Wei Luo 0001, Shang Gao 0003
IJCNN5
2023 CAMG: Context-Aware Moment Graph Network for Multimodal Temporal Activity Localization via Language
Yuelin Hu, Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
NLPCC (1)7
2023 Towards uniform point distribution in feature-preserving point cloud filtering
abstract
While a popular representation of 3D data, point clouds may contain noise and need filtering before use. Existing point cloud filtering methods either cannot preserve sharp features or result in uneven point distributions in the filtered output. To address this problem, this paper introduces a point cloud filtering method that considers both point distribution and feature preservation during filtering. The key idea is to incorporate a repulsion term with a data term in energy minimization. The repulsion term is responsible for the point distribution, while the data term aims to approximate the noisy surfaces while preserving geometric features. This method is capable of handling models with fine-scale features and sharp features. Extensive experiments show that our method quickly yields good results with relatively uniform point distribution.
Shuaijun Chen, Jinxi Wang, Wei Pan 0010, Shang Gao 0003, Meili Wang 0001, Xuequan Lu
Comput. Vis. Media4
2023 3D face recognition: A comprehensive survey in 2022
abstract
In the past ten years, research on face recognition has shifted to using 3D facial surfaces, as 3D geometric information provides more discriminative features. This comprehensive survey reviews 3D face recognition techniques developed in the past decade, both conventional methods and deep learning methods. These methods are evaluated with detailed descriptions of selected representative works. Their advantages and disadvantages are summarized in terms of accuracy, complexity, and robustness to facial variations (expression, pose, occlusion, etc.). A review of 3D face databases is also provided, and a discussion of future research challenges and directions of the topic.
Yaping Jing, Xuequan Lu, Shang Gao 0003
Comput. Vis. Media3
2023 Enhanced graph neural network for session-based recommendation
Zhenzhen Sheng, Tao Zhang 0022, Yuejie Zhang, Shang Gao 0003
Expert Syst. Appl.4
2023 Designing a Secure Blockchain-Based Supply Chain Management Framework
abstract
Supply chain management (SCM) faces a critical security issue because of the asymmetry of information delivered to various parties in the ecosystem and the lack of corresponding supervision. In response, we propose the use of blockchain technology to address the SCM security issues and put forward a blockchain-based SCM framework. We apply design science paradigm to guide the blockchain-based SCM framework development and implementation of a proof-of-concept prototype. We use Hyperledger Fabric and Composer to develop the prototype artifact. Performance evaluation results issued from Hyperledger Caliper prove the superiority and robustness of the proposed blockchain-based framework in terms of security and efficiency requirements, and performance metrics including throughput and latency. Also, the evaluation results show that the IT artifact is stable, and the high stability can reduce the risks of system vulnerabilities and breakdown.
Jiongbin Liu, William Yeoh 0002, Longxiang Gao, Shang Gao 0003, Ojelanki K. Ngwenyama
J. Comput. Inf. Syst.4
2023 Graph classification via discriminative edge feature learning
abstract
Spectral graph convolutional neural networks (GCNNs) have been producing encouraging results in graph classification tasks. However, most spectral GCNNs utilize fixed graphs when aggregating node features, while omitting edge feature learning and failing to get an optimal graph structure. Moreover, many existing graph datasets do not provide initialized edge features, further restraining the ability of learning edge features via spectral GCNNs. In this paper, we try to address this issue by designing an edge feature scheme and an add-on layer between every two stacked graph convolution layers in spectral GCNN. Both are lightweight while effective in filling the gap between edge feature learning and performance enhancement of graph classification. The edge feature scheme makes edge features adapt to node representations at different spectral graph convolution layers. The add-on layers help adjust the edge features to an optimal graph structure. To test the effectiveness of our method, we take Euclidean positions as initial node features and extract graphs with semantic information from point cloud objects. The node features of our extracted graphs are more scalable for edge feature learning than most existing graph datasets (in one-hot encoded label format). Three new graph datasets are constructed based on ModelNet40, ModelNet10 and ShapeNet Part datasets. Experimental results show that our method outperforms state-of-the-art graph classification methods on the new datasets. Our code and the constructed graph datasets will be released to the community.
Xuequan Lu, Shang Gao 0003, Antonio Robles-Kelly, Yuejie Zhang
Pattern Recognit.3
2023 Cyber Information Retrieval Through Pragmatics Understanding and Visualization
abstract
The amount of cybersecurity-related information is extraordinarily increasing, given the fast-growing number of cybersecurity attacks and the significant influence brought by them. How to efficiently obtain and precisely understand the relevant knowledge in the sea of information on cybersecurity becomes a challenge. In this article, we propose an innovative cybersecurity retrieval scheme that supports automatic indexing and searching of cybersecurity information based on semantic contents and hidden metadata. The proposed scheme leverages a customized neural model that incorporates new linguistic features and word embedding by identifying the entities related to cybersecurity incidents from the text. We implement a novel cybersecurity search engine to demonstrate effective, understandable and pragmatic cybersecurity information retrieval based on the proposed schema. Comprehensive performance evaluation over real-world datasets has been conducted to validate the new algorithms and techniques developed for cybersecurity information retrieval. The new engine makes it possible to conduct augmented search, cybersecurity analytics, and visualization, with the ultimate goal of providing direct and efficient results to help people obtain and truly understand cybersecurity information.
Nan Sun 0002, Jun Zhang 0010, Shang Gao 0003, Leo Yu Zhang, Seyit Ahmet Çamtepe, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.3
2022 Single-Modality Endoscopic Polyp Segmentation via Random Color Reversal Synthesis and Two-Branched Learning
abstract
Endoscopic polyp segmentation plays a fundamental role in the diagnosis and treatment of colorectal cancer. However, polyp segmentation often suffers from limited accuracy due to its large variations in appearance, blurry boundary and severe imbalanced illumination. In this paper, we propose a novel Translation Assisted Segmentation Network (TASNet) for polyp segmentation of single-modality endoscopic images. It consists of two branches, i.e. an image-to-image translation branch and an image segmentation branch. These two branches communicate via a shared encoder. For the image-to-image translation branch, a Color Reversal Strategy is established to treat the original image as source image and synthesize target images. Moreover, we introduce a Random Color Reversal Synthesis module for progressive segmentation. Extensive experiments show that our framework achieves superior performance than state-of-the-art methods on five widely-used endoscopic image datasets.
Mingzhu Chen, Jilan Xu, Runtian Yuan, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM8
2022 MedSeq: Semantic Segmentation for Medical Image Sequences
abstract
Medical image segmentation plays a critical role in computer-aided diagnosis, while the diversity and complexity of medical images make it difficult to segment precisely. In practice, medical images of specific modalities (e.g. Magnetic Resonance Imaging, Colonoscopy and Ultrasonography) are collected as sequences independently for every patient. However, 1) there exists few works exploiting sequence information among successive frames, neglecting inter-frame relationships that are useful to locate target objects; 2) the performance of medical image segmentation is limited to the low contrast or blurry boundary of medical images, and intra-frame dependencies are not fully explored. Thus in this paper, we propose MedSeq for segmenting objects of interest in medical image sequences. Following the “locate-then-refine” paradigm, we locate target regions by modeling cross-frame relationships and then perform refinement on coarse masks. More specifically, we design a Cross-frame Attention module to learn correlations among frames, taking advantages of their similar appearances. For refinement, we propose a novel Boundary-aware Transformer to improve the segmentation of boundary patches. Extensive experiments are conducted on benchmark datasets of Cardiac Segmentation and Video Polyp Segmentation. Our method achieves superior performance over the state-of-the-art methods.
Runtian Yuan, Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003
BIBM8
2022 Client Selection Based on Diversity Scaling for Federated Learning on Non-IID Data
Yuechao Ren, Atul Sajjanhar, Shang Gao 0003, Seng W. Loke
BROADNETS3
2022 Semi-supervised Continual Learning with Meta Self-training
abstract
Continual learning (CL) aims to enhance sequential learning by alleviating the forgetting of previously acquired knowledge. Recent advances in CL lack consideration of the real-world scenarios, where labeled data are scarce and unlabeled data are abundant. To narrow this gap, we focus on semi-supervised continual learning (SSCL). We exploit unlabeled data under limited supervision in the CL setting and demonstrate the feasibility of semi-supervised learning in CL. In this work, we propose a novel method, namely Meta-SSCL, which combines meta-learning with pseudo-labeling and data augmentations to learn a sequence of semi-supervised tasks without catastrophic forgetting. Extensive experiments on CL benchmark text classification datasets show that our method achieves promising results in SSCL.
Stella Ho, Ming Liu 0028, Lan Du 0002, Longxiang Gao, Shang Gao 0003
CIKM6
2022 CREAM: Weakly Supervised Object Localization via Class RE-Activation Mapping
abstract
Weakly Supervised Object Localization (WSOL) aims to localize objects with image-level supervision. Existing works mainly rely on Class Activation Mapping (CAM) de-rived from a classification model. However, CAM-based methods usually focus on the most discriminative parts of an object (i.e., incomplete localization problem). In this paper, we empirically prove that this problem is associated with the mixup of the activation values between less discrimi-native foreground regions and the background. To address it, we propose Class RE-Activation Mapping (CREAM), a novel clustering-based approach to boost the activation values of the integral object regions. To this end, we in-troduce class-specific foreground and background context embeddings as cluster centroids. A CAM-guided momen-tum preservation strategy is developed to learn the context embeddings during training. At the inference stage, the re-activation mapping is formulated as a parameter es-timation problem under Gaussian Mixture Model, which can be solved by deriving an unsupervised Expectation- Maximization based soft-clustering algorithm. By simply integrating CREAM into various WSOL approaches, our method significantly improves their performance. CREAM achieves the state-of-the-art performance on CUB, ILSVRC and OpenImages benchmark datasets. Code will be avail-able at https://github.com/lazzcharles/CREAM.
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
CVPR8
2022 Semantic-Driven Saliency-Context Separation for Video Captioning
abstract
Video captioning aims at generating a natural language de-scription for a given video clip including not only salient sce-narios but also contextual scenarios. The former reveal the highlight of a video and are usually the focus of most existing captioning methods. The latter, however, are not well ex-plored and even ignored easily, though they may provide cer-tain detailed and latent information that can help with a better understanding of the video. To effectively exploit the infor-mation contained in both, a novel video captioning network is proposed. It has two key modules: Cross-Modality Selection (CMS) and Saliency-Context Adaptive Decoder (SCAD). Specifically, CMS mainly focuses on utilizing the semantic information to distinguish saliency and context. Meanwhile, SCAD adaptively identifies both the saliency and context to generate more detailed and precise captions. Experiments on two benchmark datasets, i.e., MSVD and MSR-VTT, demon-strate the effectiveness of our model through the comparison with state-of-the-art methods.
Heming Jing, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICME7
2022 STDNet: Spatio-Temporal Decomposed Network for Video Grounding
abstract
Previous methods for video grounding treated either the query or the video as a whole, while neglecting their respective semantics in the orthogonal space and time dimensions. Since spatial semantics appears frequently in a video, temporal semantics is more discriminative and deserves more attention. Based on such considerations, we propose a novel Spatio-Temporal Decomposed Network (STDNet) which decomposes the query and the video into their spatial and temporal semantics, respectively. Specifically, spatial and temporal words are selected from the query, and the video is split into two pathways. Spatial cross-modal attention is computed first and serves as prior knowledge for temporal attention. A new localization strategy is also devised which regresses the segment's start conditioned on the end and essentially breaks the independence assumption made in previous methods. Experimental results on three public benchmark datasets show that our STDNet outperforms the state-of-the-art methods.
Yuanwu Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
ICME7
2022 A Generic Enhancer for Backdoor Attacks on Deep Neural Networks
Bilal Hussain Abbasi, Leo Yu Zhang, Shang Gao 0003, Antonio Robles-Kelly, Robin Doss
ICONIP (7)4
2022 TCCNet: Temporally Consistent Context-Free Network for Semi-supervised Video Polyp Segmentation
abstract
Automatic video polyp segmentation (VPS) is highly valued for the early diagnosis of colorectal cancer. However, existing methods are limited in three respects: 1) most of them work on static images, while ignoring the temporal information in consecutive video frames; 2) all of them are fully supervised and easily overfit in presence of limited annotations; 3) the context of polyp (i.e., lumen, specularity and mucosa tissue) varies in an endoscopic clip, which may affect the predictions of adjacent frames. To resolve these challenges, we propose a novel Temporally Consistent Context-Free Network (TCCNet) for semi-supervised VPS. It contains a segmentation branch and a propagation branch with a co-training scheme to supervise the predictions of unlabeled image. To maintain the temporal consistency of predictions, we design a Sequence-Corrected Reverse Attention module and a Propagation-Corrected Reverse Attention module. A Context-Free Loss is also proposed to mitigate the impact of varying contexts. Extensive experiments show that even trained under 1/15 label ratio, TCCNet is comparable to the state-of-the-art fully supervised methods for VPS. Also, TCCNet surpasses existing semi-supervised methods for natural image and other medical image segmentation tasks.
Jilan Xu, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Xuequan Lu, Shang Gao 0003
IJCAI8
2022 A Differential Privacy Mechanism for Deceiving Cyber Attacks in IoT Networks
Guizhen Yang, Mengmeng Ge 0001, Shang Gao 0003, Xuequan Lu, Leo Yu Zhang, Robin Doss
NSS3
2022 A novel privacy protection scheme for location-based services using collaborative caching
Nisha Nisha, Iynkaran Natgunanathan, Shang Gao 0003, Yong Xiang 0001
Comput. Networks3
2022 An energy efficient and runtime-aware framework for distributed stream computing systems
Dawei Sun 0001, Yijing Cui, Minghui Wu 0003, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.4
2022 A multi-level collaborative framework for elastic stream computing systems
Dawei Sun 0001, Shang Gao 0003, Xunyun Liu, Rajkumar Buyya
Future Gener. Comput. Syst.2
2022 MQDS: An energy saving scheduling strategy with diverse QoS constraints towards reconfigurable cloud storage systems
Xindong You, Dawei Sun 0001, Xueqiang Lv, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.4
2022 A state lossless scheduling strategy in distributed stream computing systems
Minghui Wu 0003, Dawei Sun 0001, Yijing Cui, Shang Gao 0003, Xunyun Liu, Rajkumar Buyya
J. Netw. Comput. Appl.4
2021 A Machine Learning-Based Elastic Strategy for Operator Parallelism in a Big Data Stream Computing System
Wei Li 0228, Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
BROADNETS3
2021 End-to-End Dynamic Pipelining Tuning Strategy for Small Files Transfer
Shimin Wu, Dawei Sun 0001, Shang Gao 0003, Guangyan Zhang
BROADNETS3
2021 Automated Security Assessment for the Internet of Things
abstract
Internet of Things (IoT) based applications face an increasing number of potential security risks, which need to be systematically assessed and addressed. Expert-based manual assessment of IoT security is a predominant approach, which is usually inefficient. To address this problem, we propose an automated security assessment framework for IoT networks. Our framework first leverages machine learning and natural language processing to analyze vulnerability descriptions for predicting vulnerability metrics. The predicted metrics are then input into a two-layered graphical security model, which consists of an attack graph at the upper layer to present the network connectivity and an attack tree for each node in the network at the bottom layer to depict the vulnerability information. This security model automatically assesses the security of the IoT network by capturing potential attack paths. We evaluate the viability of our approach using a proof-of-concept smart building system model which contains a variety of real-world IoT devices and poten-tial vulnerabilities. Our evaluation of the proposed framework demonstrates its effectiveness in terms of automatically predicting the vulnerability metrics of new vulnerabilities with more than 90% accuracy, on average, and identifying the most vulnerable attack paths within an IoT network. The produced assessment results can serve as a guideline for cybersecurity professionals to take further actions and mitigate risks in a timely manner.
Xuanyu Duan, Mengmeng Ge 0001, Triet Huynh Minh Le, Faheem Ullah, Shang Gao 0003, Xuequan Lu, Muhammad Ali Babar 0001
PRDC5
2021 Lr-Stream: Using latency and resource aware scheduling to improve latency and throughput for streaming applications
Dawei Sun 0001, Hanyu He, Hongbin Yan, Shang Gao 0003, Xunyun Liu, Xinqi Zheng
Future Gener. Comput. Syst.4
2021 Deep neural-based vulnerability discovery demystified: data, model and performance
Guanjun Lin, Leo Yu Zhang, Shang Gao 0003, Yonghang Tai, Jun Zhang 0010
Neural Comput. Appl.4
2020 Label Generation Network based on Self-selected Historical Information for Multiple Disease Classification on Chest Radiography
abstract
Deep learning has made significant break through's in image classification, but accurate diagnosis on chest radiography remains challenging due to a variety of potential diseases contained in one scan. Complex relations among diseases have significant clinical meanings, but are always ignored in most of previous work. Thus in this paper, we propose a novel Label Generation Network (LGN) which treats the label sequence as the caption of a radiology image and utilizes RNN to generate the disease labels according to the semantic relations and co-occurrence dependency among them. However, the sequential generation process of RNN makes it hard to capture the complex topological relations among diseases. To mitigate this problem, a Historical Information Module (HIM) is especially introduced to LGN, in which all the generated labels are fully considered when generating a new label. Moreover, a specific self-attention mechanism is applied in HIM to learn the topological disease relations and utilize them to select useful historical information which can provide positive guidance to the prediction of new label. Very positive results have been obtained in our experiments on the benchmark dataset of Chest X-ray14, which significantly outperform the state-of-the-art methods.
Yuelin Hu, Yuejie Zhang, Tao Zhang 0022, Shang Gao 0003, Weiguo Fan
BIBM4
2020 Sentiment Analysis of Film Reviews Based on Deep Learning Model Collaborated with Content Credibility Filtering
Xindong You, Xueqiang Lv, Shangqian Zhang, Dawei Sun 0001, Shang Gao 0003
CollaborateCom (1)5
2020 Video Captioning With Temporal And Region Graph Convolution Network
abstract
Video captioning aims to generate a natural language description for a given video clip that includes not only spatial information but also temporal information. To better exploit such spatial-temporal information attached to videos, we propose a novel video captioning framework with Temporal Graph Network (TGN) and Region Graph Network (RGN). TGN mainly focuses on utilizing the sequential information of frames that most of existing methods ignore. RGN is designed to explore the relationships among salient objects. Different from previous work, we introduce Graph Convolution Network (GCN) to encode frames with their sequential information and build a region graph for utilizing object information. We also particularly adopt a stack GRU decoder with a coarse-to-fine structure for caption generation. Very promising experimental results on two benchmark datasets (MSVD and MSR-VTT) show the effectiveness of our model.
Xinlong Xiao, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003, Weiguo Fan
ICME5
2020 Data Analytics of Crowdsourced Resources for Cybersecurity Intelligence
Nan Sun 0002, Jun Zhang 0010, Shang Gao 0003, Leo Yu Zhang, Seyit Ahmet Çamtepe, Yang Xiang 0001
NSS3
2020 Dynamic redirection of real-time data streams for elastic stream computing
Dawei Sun 0001, Shang Gao 0003, Xunyun Liu, Xindong You, Rajkumar Buyya
Future Gener. Comput. Syst.2
2019 Fog Computing Based Traffic and Car Parking Intelligent System
Walaa Alajali, Shang Gao 0003, Abdulrahman D. Alhusaynat
ICA3PP (2)2
2019 State and runtime-aware scheduling in elastic stream computing systems
Dawei Sun 0001, Shang Gao 0003, Xunyun Liu, Fengyun Li, Xinqi Zheng, Rajkumar Buyya
Future Gener. Comput. Syst.2
2018 Rethinking elastic online scheduling of big data streaming applications over high-velocity continuous data streams
Dawei Sun 0001, Hongbin Yan, Shang Gao 0003, Xunyun Liu, Rajkumar Buyya
J. Supercomput.3
2017 Shortest Path Discovery in Consideration of Obstacle in Mobile Social Network Environments
Dawei Sun 0001, Wentian Qu, Shang Gao 0003, Li Liu 0026
CollaborateCom3
2017 Performance Analysis of Storm in a Real-World Big Data Stream Computing Environment
Hongbin Yan, Dawei Sun 0001, Shang Gao 0003, Zhangbing Zhou
CollaborateCom3
2016 Supporting Adaptive Tour with High Level Petri Nets
abstract
One of the issues for tour planning applications is to adaptively provide personalized advices for different types of tourists and tour activities. This paper proposes a high level Petri Nets based approach to providing some level of adaptation by implementing adaptive navigation in a tour node space. The new model supports dynamic reordering or removal of tour nodes along a tour path; it supports multiple travel modes and incorporates multimodality within its tour planning logic to derive adaptive tour. Examples are given to demonstrate how to realize adaptive interfaces and personalization. Future directions are also discussed at the end of this paper.
Shang Gao 0003, Junyu Niu, Dawei Sun 0001
KES1
2016 A Strategy to Improve Accuracy of Multi-dimensional Feature Forecasting in Big Data Stream Computing Environments
Dawei Sun 0001, Shang Gao 0003, Fengyun Li
WISE (1)3
2012 Modeling a Dynamic Data Replication Strategy to Increase System Availability in Cloud Computing Environments
Dawei Sun 0001, Guiran Chang, Shang Gao 0003, Lizhong Jin, Xingwei Wang 0001
J. Comput. Sci. Technol.3
2007 Enhancing Web-Based Adaptive Learning with Colored Timed Petri Net
Shang Gao 0003, Robert Dew
KSEM1
2005 Supporting Adaptive Learning in Hypertext Environment: A High Level Timed Petri Net Based Approach
abstract
One problem for hypertext-based learning application is to control learning paths for different learning activities. This paper first introduced related concepts of hypertext learning state space and Petri net, then proposed a high level timed Petri Net based approach to provide some kinds of adaptation for learning activities. Examples were given while explaining ways to realizing adaptive instructions. Possible future directions were also discussed at the end of this paper.
Shang Gao 0003, Zili Zhang 0001, Igor T. Hawryszkiewycz
ICALT1
2005 Supporting Adaptive Learning with High Level Timed Petri Nets
Shang Gao 0003, Zili Zhang 0001, Jason Wells, Igor T. Hawryszkiewycz
KES (3)1
2005 Supporting Awareness in Asynchronous Collaborative Environments
Shang Gao 0003, DongBai Xue, Igor T. Hawryszkiewycz
WEBIST1