VLDB 2026 Research / reviewers in the wild / expert
Jin Su
dblp:70/10295
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Protein 3D Graph Structure Learning for Robust Structure-Based Protein Property PredictionabstractProtein structure-based property prediction has emerged as a promising approach for various biological tasks, such as protein function prediction and sub-cellular location estimation. The existing methods highly rely on experimental protein structure data and fail in scenarios where these data are unavailable. Predicted protein structures from AI tools (e.g., AlphaFold2) were utilized as alternatives. However, we observed that current practices, which simply employ accurately predicted structures during inference, suffer from notable degradation in prediction accuracy. While similar phenomena have been extensively studied in general fields (e.g., Computer Vision) as model robustness, their impact on protein property prediction remains unexplored. In this paper, we first investigate the reason behind the performance decrease when utilizing predicted structures, attributing it to the structure embedding bias from the perspective of structure representation learning. To study this problem, we identify a Protein 3D Graph Structure Learning Problem for Robust Protein Property Prediction (PGSL-RP3), collect benchmark datasets, and present a protein Structure embedding Alignment Optimization framework (SAO) to mitigate the problem of structure embedding bias between the predicted and experimental protein structures. Extensive experiments have shown that our framework is model-agnostic and effective in improving the property prediction of both predicted structures and experimental structures. Yufei Huang 0002, Siyuan Li 0002, Lirong Wu, Jin Su, Odin Zhang, Zhangyang Gao, Jiangbin Zheng 0002, Stan Z. Li |
AAAI | 4 |
| 2024 | Analysis and Evaluation of Manned/Unmanned Collaborative Emergency Rescue Equipment SystemabstractUnmanned equipment is more and more widely used in complex dangerous scenarios such as emergency rescue. The cooperative operation of manned/unmanned hybrid formation will be the main method at present. This paper proposes a topology analysis and evaluation method of manned/unmanned cooperative equipment systems and constructs a topology model from the logic layer to the physical layer. According to the concepts of degree and betweenness centrality, the key nodes analysis and invulnerability evaluation method of manned/unmanned equipment system are proposed. Finally, the method's effectiveness is verified by simulation, which provides system and data support for dynamic reconstruction and real-time decision-making of manned/unmanned equipment systems. Jin Su, Chunming Li, Yuanqing Xia |
CSCWD | 1 |
| 2024 | SaProt: Protein Language Modeling with Structure-aware VocabularyabstractLarge-scale protein language models (PLMs), such as the ESM family, have achieved remarkable performance in various downstream tasks related to protein structure and function by undergoing unsupervised training on residue sequences. They have become essential tools for researchers and practitioners in biology. However, a limitation of vanilla PLMs is their lack of explicit consideration for protein structure information, which suggests the potential for further improvement. Motivated by this, we introduce the concept of a ``structure-aware vocabulary" that integrates residue tokens with structure tokens. The structure tokens are derived by encoding the 3D structure of proteins using Foldseek. We then propose SaProt, a large-scale general-purpose PLM trained on an extensive dataset comprising approximately 40 million protein sequences and structures. Through extensive evaluation, our SaProt model surpasses well-established and renowned baselines across 10 significant downstream tasks, demonstrating its exceptional capacity and broad applicability. We have made the code, pre-trained model, and all relevant materials available at https://github.com/westlake-repl/SaProt. Jin Su, Chenchen Han, Junjie Shan, Xibin Zhou, Fajie Yuan |
ICLR | 1 |
| 2023 | Traffic accident location study based on AD-DBSCAN Algorithm with Adaptive ParametersabstractAiming at the shortcomings of the traditional Density-Based Spatial Clustering of Applications with Noise -DBSCAN algorithm such as insignificant clustering effect and the choice of parameter combinations. This paper proposes an AD-DBSCAN algorithm with adaptive parameters, which makes the algorithm more difficult in the selection of the parameters. By establishing a DBSCAN algorithm model to adapt to finding the optimal distance threshold and the minimum number of neighbor points, the clustering is more accurate, and the noise point identified in the data is more accurate. Through the observation of the calculation model of the Calinski-Harabasz index, the evaluation index of the clustering algorithm, the selection of the optimal best distance threshold and the minimum number of neighborhood points, the accuracy of noise point recognition is improved by 5 times in the clustering algorithm, and the Calinski-Harabasz index improved by about 39.84%. The applicability of the algorithm in clustering the locations of urban road traffic accidents is verified. Xijun Zhang, Jin Su, Xianli Zhang |
CSCWD | 2 |
| 2022 | Exploring evolution-aware & -free protein language models as protein function predictorsabstractLarge-scale Protein Language Models (PLMs) have improved performance in protein prediction tasks, ranging from 3D structure prediction to various function predictions. In particular, AlphaFold, a ground-breaking AI system, could potentially reshape structural biology. However, the utility of the PLM module in AlphaFold, Evoformer, has not been explored beyond structure prediction. In this paper, we investigate the representation ability of three popular PLMs: ESM-1b (single sequence), MSA-Transformer (multiple sequence alignment), and Evoformer (structural), with a special focus on Evoformer. Specifically, we aim to answer the following key questions: (1) Does the Evoformer trained as part of AlphaFold produce representations amenable to predicting protein function? (2) If yes, can Evoformer replace ESM-1b and MSA-Transformer? (3) How much do these PLMs rely on evolution-related protein data? In this regard, are they complementary to each other? We compare these models by empirical study along with new insights and conclusions. All code and datasets for reproducibility are available at https://github.com/elttaes/Revisiting-PLMs . Mingyang Hu, Fajie Yuan, Kevin Yang, Fusong Ju, Jin Su, Qiuyang Ding |
NeurIPS | 5 |
| 2022 | Multi-class fuzzy support matrix machine for classification in roller bearing fault diagnosis
Haiyang Pan, Jinde Zheng, Jin Su, Jinyu Tong |
Adv. Eng. Informatics | 4 |
| 2021 | A Bacterial Foraging Optimization Algorithm for User Interface Layout Design in Complex Human-computer Interaction SystemabstractThe layout design of user interface (UI) is a key issue in making a better human-computer interaction system for complex system such as spacecraft or vehicles, which directly affects the information level, tactical performance and operational efficiency. Firstly, this study builds a model for the layout design of UI which concerns the importance and use frequency of the layout components. Then, an improved bacterial foraging optimization algorithm is proposed to solve the layout optimization problems, which include the process of chemotaxis, replicates and migration. The preliminary results showed the feasibility of the proposed method. Nianfu Jin, Jin Su, Fengying Pang |
CSCWD | 4 |
| 2021 | Research on Tasks Scheduling Model based on Awareness-Control Resource of the Vehicle CrewabstractIn this paper, four kinds of task scheduling models are established, which are limited by the awareness-control resource of the vehicle crew. Based on the models, the tasks are scheduled from the perspective of the required time and resources, which achieves the reasonable allocation of the vehicle crew awareness-control resources and the requirements of the crew task load. The models are solved by a self-adaptive particle swarm optimization algorithm with dynamically changing inertia weight (DCWPSO). The proposed method, on the one hand, optimizes the task load distribution of the vehicle crew in general while keeping the vehicle crew with high work efficiency in the whole task cycle. On the other hand, the completion contribution of the vehicle crew in the task chain is increased by the time arrangement of key tasks, thus improving the overall task service effectiveness. ZhanHua Yang, Jin Su, Xiwen Shang |
CSCWD | 2 |
| 2021 | Multi-granularity Textual Adversarial Attack with Behavior CloningabstractRecently, the textual adversarial attack models become increasingly popular due to their successful in estimating the robustness of NLP models.However, existing works have obvious deficiencies.(1) They usually consider only a single granularity of modification strategies (e.g.word-level or sentence-level), which is insufficient to explore the holistic textual space for generation; (2) They need to query victim models hundreds of times to make a successful attack, which is highly inefficient in practice.To address such problems, in this paper we propose MAYA, a Multi-grAnularitY Attack model to effectively generate high-quality adversarial samples with fewer queries to victim models.Furthermore, we propose a reinforcement-learning based method to train a multi-granularity attack agent through behavior cloning with the expert knowledge from our MAYA algorithm to further reduce the query times.Additionally, we also adapt the agent to attack blackbox models that only output labels without confidence scores.We conduct comprehensive experiments to evaluate our attack models by attacking BiLSTM, BERT and RoBERTa in two different black-box attack settings and three benchmark datasets.Experimental results show that our models achieve overall better attacking performance and produce more fluent and grammatical adversarial samples compared to baseline models.Besides, our adversarial attack agent significantly reduces the query times in both attack settings. Yangyi Chen, Jin Su, Wei Wei 0002 |
EMNLP (1) | 2 |