VLDB 2026 Research / reviewers in the wild / expert
Jianfeng Si
dblp:46/11145
· DBLP profile ↗
7ranked-venue papers
3as first author
2since 2021 · last 2026
0009-0004-6297-2726ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 46% Language models and text generation · 18% Generative modeling · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 77% Data mining · 23% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
2.0 | 2 | 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment · ACL (1) 2026 Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Machine learning › Trustworthy machine learning › AI safety
content safety |
1.0 | 1 | 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Machine learning › Generative modeling › diffusion model
controllable generation |
1.0 | 1 | 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model |
1.0 | 1 | 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › self-play
self-play reinforcement learning |
1.0 | 1 | 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model safety |
0.3 | 1 | 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment · ACL (1) 2026 |
Machine learning › Trustworthy machine learning › safety evaluation
red teaming |
0.3 | 1 | 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Natural language and speech › Information extraction and text analysis › sentiment analysis
fine-grained opinion mining |
0.2 | 1 | 2015 | Extracting Verb Expressions Implying Negative Opinions · AAAI 2015 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.2 | 1 | 2015 | Extracting Verb Expressions Implying Negative Opinions · AAAI 2015 |
Computational finance and economics › financial market prediction
stock prediction |
0.2 | 1 | 2014 | Exploiting Social Relations and Sentiment for Stock Prediction · EMNLP 2014 |
Web and social media mining
social network analysis |
0.2 | 1 | 2014 | Exploiting Social Relations and Sentiment for Stock Prediction · EMNLP 2014 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field |
0.1 | 1 | 2015 | Extracting Verb Expressions Implying Negative Opinions · AAAI 2015 |
Data mining › text mining
topic model |
0.1 | 1 | 2014 | Exploiting Social Relations and Sentiment for Stock Prediction · EMNLP 2014 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.0self-play · 1.0reinforcement learning · 1.0direct preference optimization · 1.0co-training · 1.0regression · 0.4labeled topic model · 0.4markov network · 0.2linguistic feature modeling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-TrainingabstractCurrent methods for content safety in Large Language Models (LLMs), such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often rely on multi-stage training pipelines and lack fine-grained, post-deployment controllability. To address these limitations, we propose a unified co-training framework that efficiently integrates multiple safety behaviors: positive (lawful/prosocial), negative (unfiltered/risk-prone) and rejective (refusal-oriented/conservative) within a single SFT stage. Notably, each behavior is dynamically activated via a simple system-level instruction, or magic token, enabling stealthy and efficient behavioral switching at inference time. This flexibility supports diverse deployment scenarios, such as positive for safe user interaction, negative for internal red-teaming, and rejective for context-aware refusals triggered by upstream moderation signals. This co-training strategy induces a distinct Safety Alignment Margin in the output space, characterized by well-separated response distributions corresponding to each safety mode. The existence of this margin provides empirical evidence for the model's safety robustness and enables unprecedented fine-grained control. Experiments show that our method matches the safety alignment quality of SFT+DPO, with our 8B model notably surpassing DeepSeek-R1 (671B) in safety performance, while significantly reducing both training complexity and deployment costs. This work presents a scalable, efficient, and highly controllable solution for LLM content safety. Jianfeng Si, Lin Sun 0010, Zhewen Tan, Xiangzheng Zhang |
AAAI | 1 |
| 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety AlignmentabstractZhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu, Kaiqi Guan, Huiyan Jin, Jiawen Tao, Xiaokun Yuan, Xiangzheng Zhang, Duohe Ma, Tong Yang, Lin Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu, Kaiqi Guan, Huiyan Jin, Jiawen Tao, Xiaokun Yuan, Xiangzheng Zhang, Duohe Ma, Tong Yang 0003, Lin Sun 0010 |
ACL (1) | 3 |
| 2015 | Extracting Verb Expressions Implying Negative OpinionsabstractIdentifying aspect-based opinions has been studied extensively in recent years. However, existing work primarily focused on adjective, adverb, and noun expressions. Clearly, verb expressions can imply opinions too. We found that in many domains verb expressions can be even more important to applications because they often describe major issues of products or services. These issues enable brands and businesses to directly improve their products or services. To the best of our knowledge, this problem has not received much attention in the literature. In this paper, we make an attempt to solve this problem. Our proposed method first extracts verb expressions from reviews and then employs Markov Networks to model rich linguistic features and long distance relationships to identify negative issue expressions. Since our training data is obtained from titles of reviews whose labels are automatically inferred from review ratings, our approach is applicable to any domain without manual involvement. Experimental results using real-life review datasets show that our approach outperforms strong baselines. Huayi Li, Arjun Mukherjee, Jianfeng Si, Bing Liu 0001 |
AAAI | 3 |
| 2015 | Review Authorship Attribution in a Similarity Space
Tieyun Qian, Bing Liu 0001, Qing Li 0001, Jianfeng Si |
J. Comput. Sci. Technol. | 4 |
| 2014 | Exploiting Social Relations and Sentiment for Stock PredictionabstractIn this paper we first exploit cash-tags ("$" followed by stocks' ticker symbols) in Twitter to build a stock network, where nodes are stocks connected by edges when two stocks co-occur frequently in tweets.We then employ a labeled topic model to jointly model both the tweets and the network structure to assign each node and each edge a topic respectively.This Semantic Stock Network (SSN) summarizes discussion topics about stocks and stock relations.We further show that social sentiment about stock (node) topics and stock relationship (edge) topics are predictive of each stock's market.For prediction, we propose to regress the topic-sentiment time-series and the stock's price time series.Experimental results demonstrate that topic sentiments from close neighbors are able to help improve the prediction of a stock markedly. Jianfeng Si, Arjun Mukherjee, Bing Liu 0001, Sinno Jialin Pan, Qing Li 0001, Huayi Li |
EMNLP | 1 |
| 2014 | Users' interest grouping from online reviews based on topic frequency and order
Jianfeng Si, Qing Li 0001, Tieyun Qian, Xiaotie Deng |
World Wide Web | 1 |
| 2012 | Leveraging Network Structure for Incremental Document Clustering
Tieyun Qian, Jianfeng Si, Qing Li 0001 |
APWeb | 2 |