VLDB 2026 Research / reviewers in the wild / expert
Shang Shang
dblp:73/9634
· DBLP profile ↗
10ranked-venue papers
6as first author
5since 2021 · last 2025
0009-0001-7596-3018ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 61% Language models and text generation · 30% Reinforcement learning · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Distributed computing theory · 56% Coding theory · 44% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein design |
1.7 | 2 | 2025 | AffinityFlow: Guided Flows for Antibody Affinity Maturation · ICML 2025 Data Distillation for extrapolative protein design through exact preference optimization · ICLR 2025 |
Machine learning › Generative modeling
flow matching |
0.9 | 1 | 2025 | AffinityFlow: Guided Flows for Antibody Affinity Maturation · ICML 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | Data Distillation for extrapolative protein design through exact preference optimization · ICLR 2025 |
Machine learning › Generative modeling › protein design
protein structure generation |
0.9 | 1 | 2025 | AffinityFlow: Guided Flows for Antibody Affinity Maturation · ICML 2025 |
Bioinformatics and computational biology › protein design
antibody affinity maturation |
0.9 | 1 | 2025 | AffinityFlow: Guided Flows for Antibody Affinity Maturation · ICML 2025 |
Machine learning › Reinforcement learning
preference learning |
0.3 | 1 | 2025 | Data Distillation for extrapolative protein design through exact preference optimization · ICLR 2025 |
Bioinformatics and computational biology
protein structure prediction |
0.3 | 1 | 2025 | AffinityFlow: Guided Flows for Antibody Affinity Maturation · ICML 2025 |
Distributed computing theory
consensus |
0.2 | 1 | 2013 | An upper bound on the convergence time for quantized consensus · INFOCOM 2013 |
Coding theory
upper bounds |
0.2 | 1 | 2013 | An upper bound on the convergence time for quantized consensus · INFOCOM 2013 |
Distributed computing theory
distributed algorithms |
0.0 | 1 | 2013 | An upper bound on the convergence time for quantized consensus · INFOCOM 2013 |
Methods — techniques the papers use, named apart from their topics
progressive search · 1.7preference optimization · 1.7inverse folding · 1.7flow matching · 1.7data distillation · 1.7co-teaching · 1.7alternating optimization · 1.7random walk · 0.2markov chain coupling · 0.2electric network theory · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Data Distillation for extrapolative protein design through exact preference optimizationabstractThe goal of protein design typically involves increasing fitness (extrapolating) beyond what is seen during training (e.g., towards higher stability, stronger binding affinity, etc.). State-of-the-art methods assume that one can safely steer proteins towards such extrapolated regions by learning from pairs alone. We hypothesize that noisy training pairs are not sufficiently informative to capture the fitness gradient and that models learned from pairs specifically may fail to capture three-way relations important for search, e.g., how two alternatives fair relative to a seed. Building on the success of preference alignment models in large language models, we introduce a progressive search method for extrapolative protein design by directly distilling into the model relevant triplet relations. We evaluated our model's performance in designing AAV and GFP proteins and demonstrated that the proposed framework significantly improves effectiveness in extrapolation tasks. Mostafa Karimi, Sharmi Banerjee, Tommi S. Jaakkola, Bella Dubrov, Shang Shang, Ron Benson |
ICLR | 5 |
| 2025 | AffinityFlow: Guided Flows for Antibody Affinity MaturationabstractAntibodies are widely used as therapeutics, but their development requires costly affinity maturation, involving iterative mutations to enhance binding affinity. This paper explores a sequence-only scenario for affinity maturation, using solely antibody and antigen sequences. Recently AlphaFlow wraps AlphaFold within flow matching to generate diverse protein structures, enabling a sequence-conditioned generative model of structure. Building on this, we propose an alternating optimization framework that (1) fixes the sequence to guide structure generation toward high binding affinity using a structure-based predictor, then (2) applies inverse folding to create sequence mutations, refined by a sequence-based predictor. A key challenge is the lack of labeled data for training both predictors. To address this, we develop a co-teaching module that incorporates valuable information from noisy biophysical energies into predictor refinement. The sequence-based predictor selects consensus samples to teach the structure-based predictor, and vice versa. Our method, AffinityFlow, achieves state-of-the-art performance in proof-of-concept affinity maturation experiments. Karla-Luise Herpoldt, Chenchao Zhao, Marcus D. Collins, Shang Shang, Ron Benson |
ICML | 6 |
| 2025 | Can LLMs deeply detect complex malicious queries? A framework for jailbreaking via obfuscating intentabstractAbstract This paper delves into a possible security flaw in large language models (LLMs), particularly in their capacity to identify malicious intent within intricate or ambiguous inquiries. We have discovered that LLMs might overlook the malicious nature of highly veiled requests, even without alterations to the malevolent text in those queries, thus exposing a significant weakness in their content analysis systems. To be specific, we pinpoint and scrutinize two aspects of this vulnerability: (i) LLMs’ diminished capability to perceive maliciousness when parsing extremely obscured queries, and (ii) LLMs’ inability to discern malicious intent in queries that have been intentionally altered to increase their ambiguity by modifying the malevolent content itself. To illustrate and tackle this problem, we propose a theoretical framework and analytical strategy, and introduce a novel black-box jailbreak attack technique called IntentObfuscator. This technique exploits the identified vulnerability by concealing the genuine intentions behind user prompts, thereby compelling LLMs to inadvertently produce restricted content and circumvent their inherent content safety protocols. We elaborate on two specific applications within this framework: ”Obscure Intention” and ”Create Ambiguity,” which skillfully manipulate the complexity and ambiguity of queries to effectively dodge the detection of malicious intent. We empirically confirm the efficacy of the IntentObfuscator approach across various models, including ChatGPT-3.5, ChatGPT-4, Qwen, and Baichuan, achieving an average jailbreak success rate of 69.21%. Remarkably, our tests on ChatGPT-3.5, boasting 100 million weekly active users, yielded an impressive success rate of 83.65%. Additionally, we verify our approach across a range of sensitive content categories, including graphic violence, racism, sexism, political sensitivity, cybersecurity threats, and criminal techniques, further highlighting the considerable impact of our findings on refining ”Red Team” tactics against LLM content security frameworks. Shang Shang, Xinqiang Zhao, Zhongjiang Yao, Yepeng Yao, Liya Su, Zijing Fan, Zhengwei Jiang |
Comput. J. | 1 |
| 2025 | Advanced code slicing with pre-trained model fine-tuned for open-source component malware detectionabstractAbstract Open Source Software (OSS) is an essential part of modern software development, with platforms such as PyPI for Python, NPM for JavaScript, and RubyGems for Ruby facilitating code sharing and reuse. However, these repositories also pose significant security risks due to potential software supply chain attacks, where payloads are injected into components, propagating threats to downstream users and critical infrastructure. Existing automatic malicious component detection tools, particularly for PyPI, struggle to distinguish between subtle differences in malicious and benign behaviors, leading to high false positive rates. To address these issues, we systematically compare and explore these subtle differences, offering a more refined and accurate detection method, Open-Source Component Code Slices BERT (OCS-BERT). OCS-BERT leverages taint-based program slicing to isolate sensitive behavior segments and fine-tunes pre-trained model to capture subtle semantic differences across programming languages. This system excels in detecting malicious Python components and exhibits encouraging cross-language transferability to JavaScript's NPM and Ruby's RubyGems. Additionally, OCS-BERT successfully detected 107 malicious components from a total of 25,759 newly-uploaded PyPI components, taking two weeks to complete the process. This achievement demonstrates the effectiveness of our method, which serves as a potent enhancement to the current repertoire of software supply chain detection methodologies. Yongshan Wang, Siyuan Pang, Zijing Fan, Shang Shang, Yepeng Yao, Zhengwei Jiang, Baoxu Liu |
Comput. J. | 4 |
| 2024 | IntentObfuscator: A Jailbreaking Method via Confusing LLM with Prompts
Shang Shang, Zhongjiang Yao, Yepeng Yao, Liya Su, Zijing Fan, Zhengwei Jiang |
ESORICS (4) | 1 |
| 2014 | The application of differential privacy for rank aggregation: Privacy and accuracy
Shang Shang, Tiance Wang, Paul W. Cuff, Sanjeev R. Kulkarni |
FUSION | 1 |
| 2013 | An upper bound on the convergence time for quantized consensusabstractWe analyze a class of distributed quantized consensus algorithms for arbitrary networks. In the initial setting, each node in the network has an integer value. Nodes exchange their current estimate of the mean value in the network, and then update their estimate by communicating with their neighbors in a limited capacity channel in an asynchronous clock setting. Eventually, all nodes reach consensus with quantized precision. We start the analysis with a special case of a distributed binary voting algorithm, then proceed to the expected convergence time for the general quantized consensus algorithm proposed by Kashyap et al. We use the theory of electric networks, random walks, and couplings of Markov chains to derive an O(N3log N) upper bound for the expected convergence time on an arbitrary graph of size N, improving on the state of art bound of O(N4log N) for binary consensus and O(N5) for quantized consensus algorithms. Our result is not dependent on the graph topology. Simulations are performed to validate the analysis. Shang Shang, Paul W. Cuff, Pan Hui 0001, Sanjeev R. Kulkarni |
INFOCOM | 1 |
| 2013 | Improving augmented reality using recommender systemsabstractWith the rapid development of smart devices and wireless communication, especially with the pre-launch of Google Glass, augmented reality (AR) has received enormous attention recently. AR adds virtual objects into a user's real-world environment enabling live interaction in three dimensions. Limited by the small display of AR devices, content selection is one of the key issues to improve user experience. In this paper, we present an aggregated random walk algorithm incorporating personal preferences, location information, and temporal information in a layered graph. By adaptively changing the graph edge weight and computing the rank score, the proposed AR recommender system predicts users' preferences and provides the most relevant recommendations with aggregated information. Zhuo Zhang 0009, Shang Shang, Sanjeev R. Kulkarni, Pan Hui 0001 |
RecSys | 2 |
| 2012 | An upper bound on the convergence time for distributed binary consensus
Shang Shang, Paul W. Cuff, Sanjeev R. Kulkarni, Pan Hui 0001 |
FUSION | 1 |
| 2011 | Wisdom of the Crowd: Incorporating Social Influence in Recommendation ModelsabstractRecommendation systems have received considerable attention recently. However, most research has been focused on improving the performance of collaborative filtering (CF) techniques. Social networks, indispensably, provide us extra information on people's preferences, and should be considered and deployed to improve the quality of recommendations. In this paper, we propose two recommendation models, for individuals and for groups respectively, based on social contagion and social influence network theory. In the recommendation model for individuals, we improve the result of collaborative filtering prediction with social contagion outcome, which simulates the result of information cascade in the decision-making process. In the recommendation model for groups, we apply social influence network theory to take interpersonal influence into account to form a settled pattern of disagreement, and then aggregate opinions of group members. By introducing the concept of susceptibility and interpersonal influence, the settled rating results are flexible, and inclined to members whose ratings are "essential". Shang Shang, Pan Hui 0001, Sanjeev R. Kulkarni, Paul W. Cuff |
ICPADS | 1 |