Shaokun Wang

dblp:249/5296 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-8945-1200ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 GOAL: Geometrically Optimal Alignment for Continual Generalized Category Discovery
abstract
Continual Generalized Category Discovery (C-GCD) requires identifying novel classes from unlabeled data while retaining knowledge of known classes over time. Existing methods typically update classifier weights dynamically, resulting in forgetting and inconsistent feature alignment. We propose GOAL, a unified framework that introduces a fixed Equiangular Tight Frame (ETF) classifier to impose a consistent geometric structure throughout learning. GOAL conducts supervised alignment for labeled samples and confidence-guided alignment for novel samples, enabling stable integration of new classes without disrupting old ones. Experiments on four benchmarks show that GOAL outperforms prior methods, reducing forgetting by 16.1% and boosting novel class discovery by 3.2%, establishing a strong solution for long-horizon continual discovery.
Jizhou Han, Chenhao Ding, Songlin Dong, Yuhang He 0001, Shaokun Wang, Yihong Gong
AAAI5
2026 StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
abstract
Continual Text-to-Video Retrieval (CTVR) is a challenging multimodal continual learning setting, where models must incrementally learn new semantic categories while maintaining accurate text-video alignment for previously learned ones, thus making it particularly prone to catastrophic forgetting. A key challenge in CTVR is feature drift, which manifests in two forms: intra-modal feature drift caused by continual learning within each modality, and non-cooperative feature drift across modalities that leads to modality misalignment. To mitigate these issues, we propose StructAlign, a structured cross-modal alignment method for CTVR. First, StructAlign introduces a simplex Equiangular Tight Frame (ETF) geometry as a unified geometric prior to mitigate modality misalignment. Building upon this geometric prior, we design a cross-modal ETF alignment loss that aligns text and video features with category-level ETF prototypes, encouraging the learned representations to form an approximate simplex ETF geometry. In addition, to suppress intra-modal feature drift, we design a Cross-modal Relation Preserving loss, which leverages complementary modalities to preserve cross-modal similarity relations, providing stable relational supervision for feature updates. By jointly addressing non-cooperative feature drift across modalities and intra-modal feature drift, StructAlign effectively alleviates catastrophic forgetting in CTVR. Extensive experiments on benchmark datasets demonstrate that our method shows competitive advantages over state-of-the-art continual retrieval approaches.
Shaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu, Yupeng Hu 0003, Liqiang Nie
SIGIR1
2025 Dynamic Integration of Task-Specific Adapters for Class Incremental Learning
abstract
Non-exemplar Class Incremental Learning (NECIL) enables models to continuously acquire new classes without retraining from scratch and storing old task exemplars, addressing privacy and storage issues. However, the absence of data from earlier tasks exacerbates the challenge of catastrophic forgetting in NECIL. In this paper, we propose a novel framework called Dynamic Integration of task-specific Adapters (DIA), which comprises two key components: Task-Specific Adapter Integration (TSAI) and Patch-Level Model Alignment. TSAI boosts compositionality through a patch-level adapter integration strategy, aggregating richer task-specific information while maintaining low computation costs. Patch-Level Model Alignment maintains feature consistency and accurate decision boundaries via two specialized mechanisms: Patch-Level Distillation Loss (PDL) and Patch-Level Feature Reconstruction (PFR). Specifically, on the one hand, the PDL preserves feature-level consistency between successive models by implementing a distillation loss based on the contributions of patch tokens to new class learning. On the other hand, the PFR promotes classifier alignment by reconstructing old class features from previous tasks that adapt to new task knowledge, thereby preserving well-calibrated decision boundaries. Comprehensive experiments validate the effectiveness of our DIA, revealing significant improvements on NECIL benchmark datasets while maintaining an optimal balance between computational complexity and accuracy.
Jiashuo Li, Shaokun Wang, Yuhang He 0001, Xing Wei 0001, Yihong Gong
CVPR2
2025 CIA: Class- and Instance-aware Adaptation for Vision-Language Models
abstract
Few-shot parameter-efficient tuning methods demonstrate promising potential for Vision-Language (V-L) models in downstream tasks. However, existing approaches primarily focus on class-level alignment between image and text features, overlooking crucial instance-specific semantic information. This limitation leads to suboptimal performance on challenging tasks and restricted generalization capability to unseen data. To address these issues, we propose Class- and Instance-aware Adaptation (CIA), a novel framework that simultaneously optimizes both class-level and instance-level alignments. Specifically, CIA introduces a novel instance encoder that leverages cross-modal self-attention to generate instance-specific text features, accompanied by a carefully designed regularization mechanism to maintain consistency between class-level and instance-level representations. Extensive experiments across 15 benchmark datasets demonstrate that CIA significantly improves the downstream adaptation of V-L models.
Lin Peng 0003, Cong Wan, Shaokun Wang, Xiang Song 0005, Yuhang He 0001, Yihong Gong
ACM Multimedia3
2025 Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) focuses on classifying known categories while simultaneously discovering novel categories from unlabeled data. However, previous GCD methods face challenges due to inconsistent optimization objectives and category confusion. This leads to feature overlap and ultimately hinders performance on novel categories. To address these issues, we propose the Neural Collapse-inspired Generalized Category Discovery (NC-GCD) framework. By pre-assigning and fixing Equiangular Tight Frame (ETF) prototypes, our method ensures an optimal geometric structure and a consistent optimization objective for both known and novel categories. We introduce a Consistent ETF Alignment Loss that unifies supervised and unsupervised ETF alignment and enhances category separability. Additionally, a Semantic Consistency Matcher (SCM) is designed to maintain stable and consistent label assignments across clustering iterations. Our method significantly enhancing novel category accuracy and demonstrating its effectiveness.
Jizhou Han, Shaokun Wang, Yuhang He 0001, Chenhao Ding, Xinyuan Gao, Songlin Dong, Yihong Gong
NeurIPS2
2024 Non-exemplar Domain Incremental Learning via Cross-Domain Concept Integration
Yuhang He 0001, Songlin Dong, Xinyuan Gao, Shaokun Wang, Yihong Gong
ECCV (49)5
2024 Enhancing Pre-trained ViTs for Downstream Task Adaptation: A Locality-Aware Prompt Learning Method
abstract
Vision Transformers (ViTs) excel in extracting global information from image patches. However, their inherent limitation lies in effectively extracting information within local regions, hindering their applicability and performance. Particularly, fully supervised pre-trained ViTs, such as Vanilla ViT and CLIP, face the challenge of locality vanishing when adapting to downstream tasks. To address this, we introduce a novel LOcality-aware pRompt lEarning (LORE) method, aiming to improve the adaptation of pre-trained ViTs to downstream tasks. LORE integrates a data-driven Black Box module (i.e., a pre-trained ViT encoder) with a knowledge-driven White Box module. The White Box module is a locality-aware prompt learning mechanism to compensate for ViTs' deficiency in incorporating local information. More specifically, it begins with the design of a Locality Interaction Network (LIN), which treats an image as a neighbor graph and employs graph convolution operations to enhance local relationships among image patches. Subsequently, a Knowledge-Locality Attention (KLA) mechanism is proposed to capture critical local regions from images, learning Knowledge-Locality (K-L) prototypes utilizing relevant semantic knowledge. Afterwards, K-L prototypes guide the training of a Prompt Generator (PG) to generate locality-aware prompts for images. The locality-aware prompts, aggregating crucial local information, serve as additional input for our Black Box module. Combining pre-trained ViTs with our locality-aware prompt learning mechanism, our Black-White Box model enables the capture of both global and local information, facilitating effective downstream task adaptation. Experimental evaluations across four downstream tasks demonstrate the effectiveness and superiority of our LORE.
Shaokun Wang, Yuhang He 0001, Yihong Gong
ACM Multimedia1
2024 Global self-sustaining and local inheritance for source-free unsupervised domain adaptation
Lin Peng 0003, Yuhang He 0001, Shaokun Wang, Xiang Song 0005, Songlin Dong, Xing Wei 0001, Yihong Gong
Pattern Recognit.3
2023 Non-Exemplar Class-Incremental Learning via Adaptive Old Class Reconstruction
abstract
In the Class-Incremental Learning (CIL) task, rehearsal-based approaches have received a lot of attention recently. However, storing old class samples is often infeasible in application scenarios where device memory is insufficient or data privacy is important. Therefore, it is necessary to rethink Non-Exemplar Class-Incremental Learning (NECIL). In this paper, we propose a novel NECIL method named POLO with an adaPtive Old cLass recOnstruction mechanism, in which a density-based prototype reinforcement method (DBR), a topology-correction prototype adaptation method (TPA), and an adaptive prototype augmentation method (APA) are designed to reconstruct pseudo features of old classes in new incremental sessions. Specifically, the DBR focuses on the low-density features to maintain the model's discriminative ability for old classes. Afterward, the TPA is designed to adapt old class prototypes to new feature spaces in the incremental learning process. Finally, the APA is developed to further adapt pseudo feature spaces of old classes to new feature spaces. Experimental evaluations on four benchmark datasets demonstrate the effectiveness of our proposed method over the state-of-the-art NECIL methods.
Shaokun Wang, Weiwei Shi 0003, Yuhang He 0001, Yihong Gong
ACM Multimedia1
2023 Deep Temporal State Perception Toward Artificial Cyber-Physical Systems
abstract
Cyber–physical systems (CPS), as the cornerstone of smart city, has been attracting great interest from academia and industry. It aims to monitor/control physical components via communication and computation while ensuring effectiveness, intelligence, and security. The related research has pointed that the state perception on physical device is the prerequisite for boosting overall CPS performance. Toward this end, we present an effective deep temporal perception network to achieve classification-based state detection. Namely, we first design a multifeature encoding network for multiview time-series representation. Concretely, on the one hand, we utilize two piecewise aggregate representation strategies to obtain the key temporal trends; on the other hand, we adopt a temporal symbolic representation strategy to capture the necessary contextual semantic correlations. Thereafter, we develop a comprehensive representation enhancement module to improve feature comprehension capability and thus boosting the overall performance and interpretability. Corresponding comparison experiments, ablation studies, and data visualization analyses on benchmark data sets have verified the effectiveness of our model.
Shaokun Wang, Fan Liu 0008, Hongyun Fan, Yupeng Hu 0003, Shijun Liu
IEEE Internet Things J.2
2023 Semantic Knowledge Guided Class-Incremental Learning
abstract
Driven by practical needs, research on Class-Incremental Learning (CIL) has received more and more attentions in recent years. A technical challenge to be conquered by CIL methods is the catastrophic forgetting problem, where the model’s performance improves rapidly on new classes while deteriorates drastically on old ones. The main causes behind catastrophic forgetting include network drifts, inter-class confusions, etc. In this paper, we propose a novel CIL method that solves the catastrophic forgetting problem from two aspects. First, to solve the inter-class confusion problem, we propose a novel Semantic knOwledge gUided ciL framework (SOUL) that consists of a CNN feature extractor and a Bi-GCN (Graph Convolutional Network) classifier. In each CIL session, we use the semantic knowledge extracted from the class labels to build two inter-class relation graphs among all the encountered old and new classes. Using these two relation graphs, we develop a Bi-GCN classifier to fuse two kinds of semantic relations in a balanced way, and then to transfer the inter-class relations from semantic modality to image classification weights. The entire SOUL framework is trained end-to-end by the standard BP algorithm, which optimizes the Bi-GCN classifier and the CNN feature extractor jointly. Second, to prevent the network drift, we develop the local topology preserving strategy that divides the global topological structure of the learned feature space into a set of local topological relations, and maintains these local relations at CIL session. Experimental evaluations demonstrate the state-of-the-art performance accuracies on benchmark image classification datasets.
Shaokun Wang, Weiwei Shi 0003, Songlin Dong, Xinyuan Gao, Xiang Song 0005, Yihong Gong
IEEE Trans. Circuits Syst. Video Technol.1
2023 Micro-Influencer Recommendation by Multi-Perspective Account Representation Learning
abstract
Influencer marketing is emerging as a new marketing method, changing the marketing strategies of brands profoundly. In order to help brands find suitable micro-influencers as marketing partners, the micro-influencer recommendation is regarded as an indispensable part of influencer marketing. However, previous works only focus on modeling theindividual imageof brands/micro-influencers, which is insufficient to represent the characteristics of brands/micro-influencers over the marketing scenarios. In this case, we propose a micro-influencer ranking joint learning framework which models brands/micro-influencers from the perspective ofindividual image,target audiences, andcooperation preferences. Specifically, to model accounts’individual image, we extract topics information and images semantic information from historical content information, and fuse them to learn the account content representation. We introducetarget audiencesas a new kind of marketing role in the micro-influencer recommendation, in which audiences information of brand/micro-influencer is leveraged to learn the multi-modal account audiences representation. Afterward, we build the attribute co-occurrence graph network to minecooperation preferencesfrom social media interaction information. Based on account attributes, thecooperation preferencesbetween brands and micro-influencers are refined to attributes’ co-occurrence information. The attribute node embeddings learned in the attribute co-occurrence graph network are further utilized to construct the account attribute representation. Finally, the global ranking function is designed to generate ranking scores for all brand-micro-influencer pairs from the three perspectives jointly. The extensive experiments on a publicly available dataset demonstrate the effectiveness of our proposed model over the state-of-the-art methods.
Shaokun Wang, Tian Gan 0002, Yu-An Liu 0028, Jianlong Wu, Liqiang Nie
IEEE Trans. Multim.1
2022 Discover Micro-Influencers for Brands via Better Understanding
abstract
With the rapid development of the influencer marketing industry in recent years, the cooperation between brands and micro-influencers on marketing has achieved much attention. As a key sub-task of influencer marketing, micro-influencer recommendation is gaining momentum. However, in influencer marketing campaigns, it is not enough to only consider marketing effectiveness. Towards this end, we propose a concept-based micro-influencer ranking framework, to address the problems of marketing effectiveness and self-development needs for the task of micro-influencer recommendation. Marketing effectiveness is improved by concept-based social media account representation and a micro-influencer ranking function. We conduct social media account representation from the perspective of historical activities and marketing direction. And two adaptive learned metrics, endorsement effect score and micro-influencer influence score, are defined to learn the micro-influencer ranking function. To meet self-development needs, we design a bi-directional concept attention mechanism to focus on brands’ and micro-influencers’ marketing direction over social media concepts. Interpretable concept-based parameters are utilized to help brands and micro-influencers make marketing decisions. Extensive experiments conducted on a real-world dataset demonstrate the advantage of our proposed method compared with the state-of-the-art methods.
Shaokun Wang, Tian Gan 0002, Yu-An Liu 0028, Jianlong Wu, Liqiang Nie
IEEE Trans. Multim.1
2021 Multi-resolution Representation for Streaming Time Series Retrieval
abstract
Streaming time series retrieval (TSR) has been widely concerned in academia and industry. Considering the large volume, high dimensionality and continuous accumulation features of time series, there is limited capability to perform in-depth similarity searching directly on the raw time series data. Therefore, time series representation, which can provide the dimension reduction-based approximate results for the raw data, should be utilized in the first step for streaming TSR. However, the existing representation-based TSR methods mainly have two limitations: on the one hand, the representation efficiency of the current methods is too slow to adapt for real-time streaming time series representation; on the other hand, the retrieval efficiency of them is also not ideal, and thus fails to recognize the specific given sequence patterns on the streaming data effectively. In this paper, we present an efficient retrieval method on streaming time series. Concretely, our method can incrementally represent the features of streaming data to automatically prune the corresponding dissimilar sequences and retain the most similar candidates for efficient one-pass searching. Extensive experiments on real world datasets have been conducted to demonstrate the superiority of our method.
Yongqi Li 0001, Fubin Yao, Shaokun Wang, Peng Zhan
Int. J. Pattern Recognit. Artif. Intell.4
2019 Seeking Micro-influencers for Brand Promotion
abstract
What made you want to wear the clothes you are wearing? Where is the place you want to visit for your next-coming holiday? Why do you like the music you frequently listen to? If you are like most people, you probably made these decisions as a result of watching influencers on social media. Furthermore, influencer marketing is an opportunity for brands to take advantage of social media using a well-defined and well-designed social media marketing strategy. However, choosing the right influencers is not an easy task. With more people gaining an increasing number of followers in social media, finding the right influencer for an E-commerce company becomes paramount. In fact, most marketers cite it as a top challenge for their brands. To address the aforementioned issues, we proposed a data-driven micro-influencer ranking scheme to solve the essential question of finding out the right micro-influencer. Specifically, we represented brands and influencers by fusing their historical posts' visual and textual information. A novel k-buckets sampling strategy with a modified listwise learning to rank model were proposed to learn a brand-micro-influncer scoring function. In addition, we developed a new Instagram brand micro-influencer dataset, consisting of 360 brands and 3,748 micro-influencers, which can benefit future researchers in this area. The extensive evaluations demonstrate the advantage of our proposed method compared with the state-of-the-art methods.
Tian Gan 0002, Shaokun Wang, Meng Liu 0006, Xuemeng Song, Yiyang Yao, Liqiang Nie
ACM Multimedia2