Suojuan Zhang

dblp:208/9761 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0003-2193-2288ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Roll Call Vote Prediction With Graph Reconstruction and Attention-Based Pruning
Jiayue Chen, Zikai Yin, Ziwei Zhao 0002, Fake Lin, Zhi Zheng 0008, Suojuan Zhang, Tong Xu 0001, Yang Wang 0001, Enhong Chen
IEEE Trans. Comput. Soc. Syst.6
2025 PhysFFTFormer: A Frequency Domain-based Vision Transformer for Efficient Remote Physiological Measurement
abstract
Remote Photoplethysmography (rPPG) is a non-contact technique for extracting physiological signals from facial videos. Recently, Transformer-based architectures have exhibited remarkable performance in rPPG estimation, owing to their excellent long-range spatial-temporal modeling capacities. However, challenges persist in applying Transformer-based rPPG methods, where quadratic computational costs and inadequate feature modeling diversity remain formidable. To address these challenges, we leverage the power of the Frequency Domain-based Vision Transformer and propose an end-to-end model PhysFFTFormer. Specifically, by integrating customized Frequency-Domain Spatiotemporal Self-Attention and Frequency-Domain Discriminative Feed-Forward modules, PhysFFTFormer efficiently captures spatial-temporal dependencies with reduced computational complexity. To extract rich and diverse feature representations, we design a dual-pathway architecture to utilize both raw and differential video frames. Furthermore, the Frequency-Domain Spatiotemporal Cross-Attention module is introduced to enhance information exchange and enable feature complementation between the two pathways. Extensive experiments on multiple benchmark datasets demonstrate PhysFFTFormer’s state-of-the-art performance, robustness, and potential for real-world non-contact health monitoring.
Sirui Zhao, Tong Xu 0001, Yu Sun 0021, Hao Wang 0076, Suojuan Zhang, Enhong Chen
ICME6
2024 A Unified Framework for Adaptive Representation Enhancement and Inversed Learning in Cross-Domain Recommendation
Luankang Zhang, Hao Wang 0076, Suojuan Zhang, Mingjia Yin, Yongqiang Han, Defu Lian, Enhong Chen
DASFAA (3)3
2024 When Box Meets Graph Neural Network in Tag-aware Recommendation
abstract
Last year has witnessed the re-flourishment of tag-aware recommender systems supported by the LLM-enriched tags. Unfortunately, though large efforts have been made, current solutions may fail to describe the diversity and uncertainty inherent in user preferences with only tag-driven profiles. Recently, with the development of geometry-based techniques, e.g., box embeddings, the diversity of user preferences now could be fully modeled as the range within a box in high dimension space. However, defect still exists as these approaches are incapable of capturing high-order neighbor signals, i.e., semantic-rich multi-hop relations within the user-tag-item tripartite graph, which severely limits the effectiveness of user modeling. To deal with this challenge, in this paper, we propose a novel framework, called BoxGNN, to perform message aggregation via combinations of logical operations, thereby incorporating high-order signals. Specifically, we first embed users, items, and tags as hyper-boxes rather than simple points in the representation space, and define two logical operations, i.e., union and intersection, to facilitate the subsequent process. Next, we perform the message aggregation mechanism via the combination of logical operations, to obtain the corresponding high-order box representations. Finally, we adopt a volume-based learning objective with Gumbel smoothing techniques to refine the representation of boxes. Extensive experiments on two publicly available datasets and one LLM-enhanced e-commerce dataset have validated the superiority of BoxGNN compared with various state-of-the-art baselines. The code is released online: https://github.com/critical88/BoxGNN.
Fake Lin, Ziwei Zhao 0002, Xi Zhu 0004, Shitian Shen, Xueying Li 0004, Tong Xu 0001, Suojuan Zhang, Enhong Chen
KDD8
2024 Dataset Regeneration for Sequential Recommendation
abstract
The sequential recommender (SR) system is a crucial component of modern recommender systems, as it aims to capture the evolving preferences of users. Significant efforts have been made to enhance the capabilities of SR systems. These methods typically follow the model-centric paradigm, which involves developing effective models based on fixed datasets. However, this approach often overlooks potential quality issues and flaws inherent in the data. Driven by the potential of data-centric AI, we propose a novel data-centric paradigm for developing an ideal training dataset using a model-agnostic dataset regeneration framework called DR4SR. This framework enables the regeneration of a dataset with exceptional cross-architecture generalizability. Additionally, we introduce the DR4SR+ framework, which incorporates a model-aware dataset personalizer to tailor the regenerated dataset specifically for a target model. To demonstrate the effectiveness of the data-centric paradigm, we integrate our framework with various model-centric methods and observe significant performance improvements across four widely adopted datasets. Furthermore, we conduct in-depth analyses to explore the potential of the data-centric paradigm and provide valuable insights. The code can be found at https://github.com/USTC-StarTeam/DR4SR.
Mingjia Yin, Hao Wang 0076, Wei Guo 0006, Yong Liu 0020, Suojuan Zhang, Sirui Zhao, Defu Lian, Enhong Chen
KDD5
2024 Bridging Gaps in Content and Knowledge for Multimodal Entity Linking
abstract
Multimodal Entity Linking (MEL) aims to address the ambiguity in multimodal mentions and associate them with Multimodal Knowledge Graphs (MMKGs). Existing works primarily focus on designing multimodal interaction and fusion mechanisms to enhance the performance of MEL. However, these methods still overlook two crucial gaps within the MEL task. One is the content discrepancy between mentions and entities, manifested as uneven information density. The other is the knowledge gap, indicating insufficient knowledge extraction and reasoning during the linking process. To bridge these gaps, we propose a novel framework FissFuse, as well as a plug-and-play knowledge-aware re-ranking method KAR. Specifically, FissFuse collaborates with the Fission and Fusion branches, establishing dynamic features for each mention-entity pair and adaptively learning multimodal interactions to alleviate content discrepancy. Meanwhile, KAR is endowed with carefully crafted instruction for intricate knowledge reasoning, serving as re-ranking agents empowered by Large Language Models (LLMs). Extensive experiments on two well-constructed MEL datasets demonstrate outstanding performance of FissFuse compared with various baselines. Comprehensive evaluations and ablation experiments validate the effectiveness and generality of KAR.
Pengfei Luo, Tong Xu 0001, Che Liu 0001, Suojuan Zhang, Linli Xu 0002, Minglei Li 0001, Enhong Chen
ACM Multimedia4
2024 MLC-DKT: A multi-layer context-aware deep knowledge tracing model
Suojuan Zhang, Jie Pu, Shuanghong Shen, Enhong Chen
Knowl. Based Syst.1
2023 A generalized multi-skill aggregation method for cognitive diagnosis
Suojuan Zhang, Xiaohan Yu 0004, Enhong Chen, Fei Wang 0063, Zhenya Huang
World Wide Web (WWW)1
2019 PROMETHEE for prioritized criteria
Xiuli Qi, Xiaohan Yu 0004, Lei Wang 0012, Xianglin Liao, Suojuan Zhang
Soft Comput.5
2018 ELECTRE methods in prioritized MCDM environment
Xiaohan Yu 0004, Suojuan Zhang, Xianglin Liao, Xiuli Qi
Inf. Sci.2