VLDB 2026 Research / reviewers in the wild / expert
Shisong Tang
dblp:319/0177
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-4550-3950ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Heterogeneous Multi-treatment Uplift Modeling for Trade-off Optimization in Short-Video RecommendationabstractThe rapid proliferation of short videos on social media platforms presents unique challenges and opportunities for recommendation systems. Users exhibit diverse preferences, and the responses resulting from different strategies often conflict with one another, potentially exhibiting inverse correlations between metrics such as watch time and video view counts. Existing uplift models face limitations in handling the heterogeneous multi-treatment scenarios of short-video recommendations, often failing to effectively capture both the synergistic and individual causal effects of different strategies. Furthermore, traditional fixed-weight approaches for balancing these responses lack personalization and can result in biased decision-making. To address these issues, we propose a novel Heterogeneous Multi-treatment Uplift Modeling (HMUM) framework for trade-off optimization in short-video recommendations. HMUM comprises an Offline Hybrid Uplift Modeling (HUM) module, which captures the synergistic and individual effects of multiple strategies, and an Online Dynamic Decision-Making (DDM) module, which estimates the weights of different user responses in real-time for personalized decision-making. Evaluated on two public datasets, an industrial dataset, and online A/B experiments on the Kuaishou platform, our model demonstrated superior offline performance and significant improvements in key metrics. It is now fully deployed on the platform, benefiting hundreds of millions of users. Chenhao Zhai, Chang Meng, Shuchang Liu 0001, Shisong Tang, Xiaoqiang Feng, Xiu Li 0001 |
KDD (1) | 6 |
| 2026 | Mining Citywide Dengue Spread Patterns in Singapore Through Hotspot Dynamics from Open Web Data
Gaoxi Xiao, Stefan Ma, Hechang Chen, Shisong Tang, Flora D. Salim |
WWW | 5 |
| 2026 | Deep learning models for digital medical imaging: a survey
Xueju Wang, Yujia Cong, Lele Cong, Xianling Cong, Shisong Tang, Hechang Chen |
Appl. Intell. | 6 |
| 2025 | SlideCoder: Layout-aware RAG-enhanced Hierarchical Slide Generation from DesignabstractWenxin Tang, Jingyu Xiao, Wenxuan Jiang, Xi Xiao, Yuhang Wang, Xuxin Tang, Qing Li, Yuehe Ma, Junliang Liu, Shisong Tang, Michael R. Lyu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Wenxin Tang, Jingyu Xiao, Wenxuan Jiang, Xi Xiao 0001, Yuhang Wang 0036, Xuxin Tang, Qing Li 0006, Yuehe Ma, Shisong Tang, Michael R. Lyu |
EMNLP | 10 |
| 2025 | Calibrating Video Watch-time Predictions with Credible Prototype AlignmentabstractAccurately predicting user watch-time is crucial for enhancing user stickiness and retention in video recommendation systems. Existing watch-time prediction approaches typically involve transformations of watch-time labels for prediction and subsequent reversal, ignoring both the natural distribution properties of label and the instance representation confusion that results in inaccurate predictions. In this paper, we propose ProWTP, a two-stage method combining prototype learning and optimal transport for watch-time regression prediction, suitable for any deep recommendation model. Specifically, we observe that the watch-ratio (the ratio of watch-time to video duration) within the same duration bucket exhibits a multimodal distribution. To facilitate incorporation into models, we use a hierarchical vector quantised variational autoencoder (HVQ-VAE) to convert the continuous label distribution into a high-dimensional discrete distribution, serving as credible prototypes for calibrations. Based on this, ProWTP views the alignment between prototypes and instance representations as a Semi-relaxed Unbalanced Optimal Transport (SUOT) problem, where the marginal constraints of prototypes are relaxed. And the corresponding optimization problem is reformulated as a weighted Lasso problem for solution. Moreover, ProWTP introduces the assignment and compactness losses to encourage instances to cluster closely around their respective prototypes, thereby enhancing the prototype-level distinguishability. Finally, we conducted extensive offline experiments on two industrial datasets, demonstrating our consistent superiority in real-world application. Chao Cui, Shisong Tang, Fan Li 0017, Jiechao Gao, Hechang Chen |
ICML | 2 |
| 2025 | VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
Tianke Zhang, Chang Meng, Xiaobei Wang, Jinpeng Wang 0002, Yifan Zhang 0004, Shisong Tang, Changyi Liu, Haojie Ding, Kaiyu Jiang, Kaiyu Tang, Hai-Tao Zheng 0002, Fan Yang 0094, Tingting Gao, Di Zhang 0026, Kun Gai |
KDD (2) | 7 |
| 2025 | Aligning and Balancing ID and Multimodal Representations for RecommendationabstractLarge-scale recommendation systems mainly rely on sparse ID features, struggling with data sparsity. It's important to use multimodal information to assist ID learning for better performance. However, there exists two challenges: (1) distribution discrepancy between multimodal and ID makes direct integration prone to user-item mismatch; (2) slower convergence of multimodal representations compared to ID, causing optimization imbalance under a unified objective, which limits the potential of multimodal representations. In this paper, we comprehensively investigate the two problems and proposes a framework named AB-Rec to align and balance ID and multimodal representations learning for recommendation. We design three alignment tasks to fine-tune a pre-trained multimodal large language model (MLLM), which is then utilized to generate a unified multimodal representation for each item. AB-Rec aligns the distributions of ID and multimodal representations by minimizing the in-batch Wasserstein distance, and maximizes the distance between the two types of representations for the same item to avoid representation collapse. To solve the optimization imbalance, we propose a gradient modulation method that adaptively controls the optimization process by monitoring the contribution differences between ID and multimodal representations. Finally, we conduct extensive offline experiments on four datasets and an A/B test on an online video platform, demonstrating the effectiveness and scalability of our proposed method. Binrui Wu, Shisong Tang, Fan Li 0017, Chang Meng, Jingyu Xiao, Jiechao Gao |
KDD (2) | 2 |
| 2025 | Leveraging Label Distributions as Anchors to Enhance Video RecommendationabstractIn video recommendation systems, accurately predicting watch time is crucial for enhancing user engagement and retention. Traditional methods typically apply label transformations or mitigate duration bias to improve performance but overlook that erroneous instance representations are the primary cause of significant prediction errors. Moreover, these approaches predominantly rely on point perdition, limiting their robustness. To address these challenges, we propose LDA, a novel prediction paradigm that optimizes instance representations by explicitly leveraging label distributions as anchors within the model, enabling more accurate and robust predictions. Our analysis reveals that watch ratio across different duration groups exhibit distinct multi-peak distributions, reflecting the strong aggregation of user behavior. Based on this finding, we employ Vector Quantized Variational Auto-encoder (VQ-VAE) to convert the continuous watch ratio distribution into representative anchors that capture these multi-peak characteristics within each duration group. Subsequently, we project both instance representations and anchors into a common space and utilize Optimal Transport (OT) to generate pseudo-labels aligned with the anchor distribution, allowing instances to obtain structured coordinates within this space during training. Finally, we derive optimized instance representations for watch time prediction by aggregating anchor vectors through weighted integration. Extensive offline experiments on two datasets and large-scale online A/B testing on a short-video platform with over 300 million DAUs demonstrate the consistent superiority of LDA in watch time prediction. Chao Cui, Shisong Tang, Fan Li 0017, Huafeng Cao, Jiechao Gao, Hechang Chen |
KDD (2) | 3 |
| 2025 | Contrastive Prototype Framework for Calibrating Video RecommendationabstractOnline video recommendation systems often build binary labels based on play complete rate (i.e., the ratio of watch time to video duration), such as complete play and effective play, using them as implicit feedback for Click-Through Rate (CTR) prediction tasks to gauge user interest. Existing works tend to improve prediction accuracy by designing complex models, overlooking that a key cause of inaccurate predictions is the disorganization of instance representation space. To address this issue, we explore a novel approach using prototype learning to calibrate the instance representation space of deep recommendation models and propose a model-agnostic Contrastive Prototype Framework (CPF). Firstly, CPF partitions the instance space into different subspaces based on duration, then generates positive and negative prototype pairs for each subspace from pre-trained recommendation model. Subsequently, we map the instance representations to the prototype space and calibrate them by reducing the distance to the corresponding prototypes. Ultimately, the prediction is derived from the linear combination of the estimated values associated with each prototype. To prevent disorganization in the prototype space during training, we design contrastive and orthogonality losses to constrain the learning of prototypes. Additionally, we show that how CPF effectively addresses the duration bias from the perspective of causal intervention. Offline experiments on two datasets demonstrate that CPF improves recommendation accuracy over several baseline models in predicting five widely used implicit feedback labels. We have also deployed CPF on a short video platform, validating its effectiveness in real-world scenarios. Fan Li 0032, Jiazhen Huang, Shisong Tang, Huafeng Cao, Haochen Sui, Xiaoyu Kang |
ACM Multimedia | 3 |
| 2025 | Tackling Continual Offline RL through Selective Weights Activation on Aligned SpacesabstractContinual offline reinforcement learning (CORL) has shown impressive ability in diffusion-based continual learning systems by modeling the joint distributions of trajectories. However, most research only focuses on limited continual task settings where the tasks have the same observation and action space, which deviates from the realistic demands of training agents in various environments. In view of this, we propose Vector-Quantized Continual Diffuser, named VQ-CD, to break the barrier of different spaces between various tasks. Specifically, our method contains two complementary sections, where the quantization spaces alignment provides a unified basis for the selective weights activation. In the quantized spaces alignment, we leverage vector quantization to align the different state and action spaces of various tasks, facilitating continual training in the same space. Then, we propose to leverage a unified diffusion model attached by the inverse dynamic model to master all tasks by selectively activating different weights according to the task-related sparse masks. Finally, we conduct extensive experiments on 15 continual learning (CL) tasks, including conventional CL task settings (identical state and action spaces) and general CL task settings (various state and action spaces). Compared with 17 baselines, our method reaches the SOTA performance. Jifeng Hu, Sili Huang, Li Shen 0008, Zhejian Yang, Shengchao Hu, Shisong Tang, Hechang Chen, Lichao Sun 0001, Yi Chang 0001, Dacheng Tao |
NeurIPS | 6 |
| 2024 | Smart Data-Driven Proactive Push to Edge Network for User-Generated VideosabstractTo reduce costs and improve performance, video Content Delivery Networks (CDNs) have started to incorporate lightweight edge nodes, e.g., WiFi access points. Because of this, it is necessary for CDNs to intelligently select which video files should be placed at their core data centers vs. these edge nodes. This is more complex than traditional CDN management, as lightweight edge nodes are much more numerous and unstable than data centers. With this in mind, we present SDPush —- a system for managing content placement in edge CDNs. SDPush tackles two problems. First, it is necessary for SDPush to select which files to proactive push. To address this, we build a file popularity prediction model that effectively identifies video files that will receive many views. Second, SDPush should determine how many replicas of each file to push. To address this, we design a model to predict the benefits of pushing particular files (regarding traffic savings) and then formulate the replica decision problem as a lightweight problem, which is solvable within seconds, even for platforms that accommodate millions of daily active users. Through a trace-driven evaluation and a live deployment on a real video platform, we validate SDPush’s effectiveness, offloading peak-period traffic by 12.1% to 23.9% from the data center to edge nodes, thereby reducing the CDN costs. Xiaoteng Ma, Qing Li 0006, Junkun Peng, Gareth Tyson, Ziwen Ye, Shisong Tang, Shengbin Meng, Gabriel-Miro Muntean |
INFOCOM | 6 |
| 2024 | Contextual Distillation Model for Diversified RecommendationabstractThe diversity of recommendation is equally crucial as accuracy in improving user experience. Existing studies, e.g., Determinantal Point Process (DPP) and Maximal Marginal Relevance (MMR), employ a greedy paradigm to iteratively select items that optimize both accuracy and diversity. However, prior methods typically exhibit quadratic complexity, limiting their applications to the re-ranking stage and are not applicable to other recommendation stages with a larger pool of candidate items, such as the pre-ranking and ranking stages. In this paper, we propose Contextual Distillation Model (CDM), an efficient recommendation model that addresses diversification, suitable for the deployment in all stages of industrial recommendation pipelines. Specifically, CDM utilizes the candidate items in the same user request as context to enhance the diversification of the results. We propose a contrastive context encoder that employs attention mechanisms to model both positive and negative contexts. For the training of CDM, we compare each target item with its context embedding and utilize the knowledge distillation framework to learn the win probability of each target item under the MMR algorithm, where the teacher is derived from MMR outputs. During inference, ranking is performed through a linear combination of the recommendation and student model scores, ensuring both diversity and efficiency. We perform offline evaluations on two industrial datasets and conduct online A/B test of CDM on the short-video platform KuaiShou. The considerable enhancements observed in both recommendation quality and diversity, as shown by metrics, provide strong superiority for the effectiveness of CDM. Fan Li 0017, Xu Si, Shisong Tang, Dingmin Wang, Kunyan Han, Guorui Zhou, Yang Song 0008, Hechang Chen |
KDD | 3 |
| 2023 | Counterfactual Video Recommendation for Duration DebiasingabstractDuration bias widely exists in video recommendations, where models tend to recommend short videos for the higher ratio of finish playing and thus possibly fail to capture users' true interests. In this paper, we eliminate the duration bias from both data and model. First, based on the extensive data analysis, we observe that play completion rate of videos with the same duration presents a bimodal distribution. Hence, we propose to perform threshold division to construct binary labels as training labels for alleviating the drawback of finish playing labels overly biased towards short videos. Algorithmically, we resort to causal inference, which enables us to inspect causal relationships of video recommendations with a causal graph. We identify that duration has two kinds of effect on prediction: direct and indirect. Duration bias lies in the direct effect, while the indirect effect benefits prediction. To this end, we design a model-agnostic Counterfactual Video Recommendation for Duration Debiasing (CVRDD) framework, which incorporates multi-task learning to estimate different causal effect during training. In the inference phase, we perform counterfactual inference to remove the direct effect of duration for unbiased prediction. We conduct experiments on two industrial datasets, and in addition to achieving highly promising results on traditional top-k recommendation metrics, CVRDD also improves the user watch time. Shisong Tang, Qing Li 0006, Dingmin Wang, Ci Gao, Wentao Xiao, Dan Zhao 0003, Yong Jiang 0001, Aoyang Zhang |
KDD | 1 |
| 2022 | Knowledge-based Temporal Fusion Network for Interpretable Online Video Popularity PredictionabstractPredicting the popularity of online videos has many real-world applications, such as recommendation, precise advertising, and edge caching strategies. Despite many efforts have been dedicated to the online video popularity prediction, there still exist several challenges: (1) The meta-data from online videos is usually sparse and noisy, which makes it difficult to learn a stable and robust representation. (2) The influence of content features and temporal features in different life cycles of online videos is dynamically changing, so it is necessary to build a model that can capture the dynamics. (3) Besides, there is a great need to interpret the predictive behavior of the model to assist administrators of video platforms in the subsequent decision-making. Shisong Tang, Qing Li 0006, Xiaoteng Ma, Ci Gao, Dingmin Wang, Yong Jiang 0001, Aoyang Zhang, Hechang Chen |
WWW | 1 |