Shuanglong Li

dblp:98/1148 · DBLP profile ↗
← Back
11ranked-venue papers in the field
0as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 6Data Mining & Knowledge Discovery · 3Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2026 A Multi-Stage Structural Captioning Framework for Enhancing Chinese Image-to-Video Generation in Baidu
abstract
The rapid advancement of video generation technology, particularly in the domain of image-to-video generation, is significantly transforming both personal and industrial applications. The performance of video generation models is heavily influenced by the quality of data, specifically video content and captions. Among these, the quality of video captions directly impacts the model's ability to follow user instructions. Currently, video captions are typically generated using vision-language model (VLM)-based methods. However, issues such as hallucinations, incomplete descriptions, and inaccuracies remain prevalent. To address these challenges, we propose a novel optimization framework for a multi-expert, structured video captioning model, specifically designed to enhance the performance of Chinese image-to-video tasks in Baidu's business scenarios. First, we construct a structured video caption dataset by combining a large language model (LLM) with human annotations, and introduce a hallucination-aware adaptive GRPO algorithm. A two-stage fine-tuning alignment strategy is employed to improve the accuracy of fine-grained content descriptions and mitigate hallucinations. Second, we design independent video sub-expert models to characterize subject motion intensity and camera movement, thereby expanding the dimensionality of video captions and addressing the issue of inaccurate sub-dimension descriptions in a single VLM. Finally, we leverage multi-level structured video captions to guide the training of the video generation model, resulting in improved quality and consistency in video generation. Extensive experiments demonstrate that our method consistently outperforms the baseline. Specifically, it achieves a 2.49% improvement in F1-score on event-level cross-validation and reduces hallucination occurrences by 30.9%. Moreover, our framework significantly enhances the quality of Chinese video generation, yielding a 2.9% increase in the VBench total score and a 9.2% improvement in the proportion of high-quality generated videos.
Zhipeng Jin, Xiawei Li, Wen Tao, Yi Yang 0031, Cong Han 0002, Shuanglong Li, Zhongmin Cai
KDD (1)8
2025 Large Vison-Language Foundation Model in Baidu AIGC Image Advertising
abstract
Recent advances in generative artificial intelligence have revolutionized information retrieval and content generation, opening up new opportunities for the e-commerce industry. Alignment learning between small models and parallel corpora cannot meet current needs. The success of ChatGPT demonstrates that large models need to first establish a fundamental understanding, and then utilize high-quality corpora for generation. Having a large model foundation is indispensable. In this paper, we establish a fundamental 10B multimodal model foundation for multimodal generation tasks and propose a scene-based alignment learning approach called conditional sample supervised fine-tuning for downstream generation tasks. Meanwhile, diffusion models are known to be vulnerable to outliers in training data. To address this, we utilize an alternative diffusion loss function that preserves the high quality of generated data like the original squared L2 loss while being robust to outliers.In practical test sets, the multimodal foundation fully demonstrates its alignment and comprehension abilities for graphic and textual content. Additionally, conditional fine-tuning and the design of the loss function significantly enhance the quality of generated content. The quality rate of images has increased by 34.3 percentage points, and prompt control has improved by 19.8 percentage points. The application of our framework in Baidu Search Ads has led to significant revenue growth. For instance, ads with generated image creatives have achieved a 29% higher click-through rate (CTR), resulting in a daily consumption of 3 million yuan.
Zhipeng Jin, Wen Tao, Yi Yang 0031, Cong Han 0002, Shuanglong Li
KDD (1)6
2025 Retrieval-Augmented Image Captioning and Generation with Entity Concepts Enhancement for Baidu Multimodal Advertising
abstract
Recent advancements in generative artificial intelligence are driving a significant transformation in information retrieval and content generation, creating substantial opportunities for online advertising. Text-to-image generation technology has become increasingly prevalent in advertising content production, demonstrating promising performance improvements in terms of semantic relevance and visual appeal. However, existing models often suffer from inadequate representation of entity concepts, such as prominent product brands and recognizable landmarks. This inherent limitation subsequently leads to notable deficiencies in brand tonality, industry-specific relevance, and market adaptability of the generated advertising content. To address this challenge, we propose a multimodal ad content generation framework specifically engineered for online advertising system, particularly focused on resolving the deficiency in entity concepts. Our framework is comprised of two phases: first, an image captioning module with entity-aware learning based on multimodal large language model, leveraging retrieval-augmented techniques to incorporate entity concepts into image descriptions; second, a text-to-image diffusion model refined on image-text pairs enriched with entity concepts to facilitate entity-grounded image generation. Extensive experiments validate the effectiveness of our framework, demonstrating superior performance in both image captioning and image generation compared to existing methods, particularly in the accuracy of depiction of relevant entities in advertising images. Moreover, the deployment of the framework in the system primary traffic of Baidu Search Ads, has brought significant enhancements to advertisement revenue for both advertisers and the platform.
Kang Zhao 0002, Zhipeng Jin, Wen Tao, Yi Yang 0031, Cong Han 0002, Shuanglong Li, Zhongmin Cai
SIGIR7
2024 Multi-Stage Refined Visual Captioning for Baidu Ad Creatives Generation
abstract
High-quality multimodal training data is of critical importance for improving of multimodal model performance. However, the utilization of web-crawled vision-caption pairs is hindered by the presence of noise and irrelevance, as well as a lack of Chinese data. Large Language Models (LLM) and Large Multimodal Models (LMM) has demonstrated promising performance in cross-modal understanding and generation. In light of this, we propose a Chinese visual captioning pipeline for the synthesis of high-quality data. Our pipeline is comprised of two phases: the initial training of an encoder for visual understanding; and the subsequent fine-tuning of a captioning model in a two-stage iterative human-in-the-loop process, where the captioning model incorporates the pre-trained vision encoder and LLM by a visual cross-attention querying transformer. Extensive experiments have been conducted to validate our framework, including both quantitative and qualitative evaluation of captions generated from images and videos. The synthesis pipeline has been integrated into the ad image creative generation process in Baidu Search Ads, resulting in enhanced capabilities in prompt following.
Yi Yang 0031, Kang Zhao 0002, Zhipeng Jin, Wen Tao, Shuanglong Li
CIKM7
2024 A Multi-Node Multi-GPU Distributed GNN Training Framework for Large-Scale Online Advertising
abstract
Graph Neural Networks (GNNs) have become critical in various domains such as online advertising but face scalability challenges due to the growing size of graph data, leading to the needs for advanced distributed GPU computation strategies across multiple nodes. This paper presents PGLBox-Cluster, a robust distributed graph learning framework constructed atop the PaddlePaddle platform, implemented to efficiently process graphs comprising billions of nodes and edges. Through strategic partitioning of the model, node attributes, and graph data and leveraging industrial-grade RPC and NCCL for communication, PGLBox-Cluster facilitates effective distributed computation. The extensive experimental results confirm that PGLBox-Cluster achieves a 1.94x to 2.93x speedup over the single-node configuration, significantly elevating graph neural network scalability and efficiency by handling datasets exceeding 3 billion nodes and 120 billion edges with its novel asynchronous communication and graph partitioning techniques. The repository is released at This Link.
Xuewu Jiao, Xinsheng Luo, Jiang Bian 0003, Junchao Yang 0001, Mingqing Hu, Weipeng Lu, Shikun Feng, Danlei Feng, Haoyi Xiong, Shuanglong Li
CIKM13
2024 Scaling Vison-Language Foundation Model to 12 Billion Parameters in Baidu Dynamic Image Advertising
abstract
Dynamic image advertising is an add-on service in search advertising that matches visuals to search ads in real-time. However, the image matching system encompasses various sub-tasks with different objectives, increasing the complexity of achieving global optimization. Besides, prevalent long-tailed data poses a challenge to the multimodal representation learning in dynamic image advertising. Recently, vision-language pre-trained models have achieved remarkable performance across a variety of multimodal tasks, and implemented as the foundational representation model in electronic business scenarios. In this paper, to improve multimodal content understanding in Dynamic Image adVERtising, we present a viSion-language rEpresentation model (referred to as DIVERSE) that learns on cross-view and cross-token contrastive loss. Moreover, with large-scale curated advertising image-text data and extensive efficient training techniques, we scale DIVERSE to 12 billion parameters, which is the biggest Chinese multimodal representation model in industrial practices. Experiment results demonstrate the distinct advantages of DIVERSE12B in business datasets, with competitive performance on public benchmarks. Further evaluation in downstream applications including ad text-image retrieval, text-image relevance modeling, and image content moderation, shows that it outperforms previous separately-trained models across offline and online metrics. Moreover, DIVERSE12B has been implemented on the system primary traffic of Baidu Search Ads, bringing considerable increase to both user experience, and revenue for advertisers and search engine.
Kang Zhao 0002, Zhipeng Jin, Yi Yang 0031, Wen Tao, Xiaodong Chen 0006, Cong Han 0002, Shuanglong Li
CIKM8
2024 Enhancing Baidu Multimodal Advertisement with Chinese Text-to-Image Generation via Bilingual Alignment and Caption Synthesis
abstract
Recent advances in generative artificial intelligence have revolutionized information retrieval and content generation, opening up new opportunities for the e-commerce industry. In particular, text-to-image generation models offer a novel approach to guiding the image generation process using natural language input, which is inspiring for multimodal search advertising. Traditional multimodal search ads require advertisers to prepare ad creatives, such as ad images, which is time-consuming and requires uniform image specifications and content quality inspection. To this end, we propose a streamlined generation framework for search ad image creatives. First, we prepare a Chinese image caption model with high-quality image-caption pairs to bootstrap training data refinement. With curated high-quality images and synthesized descriptive captions, we then train a Chinese text-to-image generation model, the largest to date, using SDXL and a 10-billion multimodal text encoder. Specifically, we introduce a two-stage bilingual multimodal representation alignment process to seamlessly integrate the text encoder with the generation model. Extensive experiments validate the effectiveness of our framework, including assessments of image captioning and image generation. The implementation of our framework in Baidu Search Ads shows significant revenue increase, For example, beauty industry ads with generated image creatives achieve a 29% higher click-through rate (CTR).
Kang Zhao 0002, Zhipeng Jin, Yi Yang 0031, Wen Tao, Cong Han 0002, Shuanglong Li
SIGIR7
2023 Large Scale Multiplex Latent Graphs for Recommendation in Baidu
abstract
In commercial systems, the interaction between user behavior and advertising content forms a complex graph network that exhibits multi-scenario and multi-modal characteristics. Multi-scenario refers to the fact that ads are typically displayed in different scenarios on the same platform, such as Baidu App’s news feed and short video scenes. Across different scenarios, users’ interests have both commonalities and differences. Multi-modal refers to the fact that the influence of advertising content modality on users varies, and users’ sensitivity to different content modalities may also differ. We believe that the fundamental relationship between user behavior and advertising content may be multi-layered, influenced by the types of presenting content scenarios and the unique features of each modality. For example, in the news feed scenario, the importance of video titles may be different from that in the video stream scenario, where the visual content of the video may be more prominent. We propose a novel heterogeneous graph neural network method that integrates multi-scenario domains and multi-modal content features, which we call the Domain-Aware Multi-Modal graph model (DAMM-GCN). Our model achieves significant improvements in multiple scenarios compared with the traditional flat graph models such as Metapath2vec [1], GraphSAGE [2] and MMGCN [3]. Furthermore, our model has been deployed in the Baidu advertising system, obtaining 2.08% improvement on CPM (Cost Per Mille).
Zhipeng Jin, Yi Yang 0031, Xuewu Jiao, Shuanglong Li
IEEE Big Data6
2023 PGLBox: Multi-GPU Graph Learning Framework for Web-Scale Recommendation
abstract
While having been used widely for large-scale recommendation and online advertising, the Graph Neural Network (GNN) has demonstrated its representation learning capacity to extract embeddings of nodes and edges through passing, transforming, and aggregating information over the graph. In this work, we propose PGLBox1 - a multi-GPU graph learning framework based on PaddlePaddle [24], incorporating with optimized storage, computation, and communication strategies, to train deep GNNs based on web-scale graphs for the recommendation. Specifically, PGLBox adopts a hierarchical storage system with three layers to facilitate I/O, where graphs and embeddings are stored in the HBMs and SSDs, respectively, with MEMs as the cache. To fully utilize multi-GPUs and I/O bandwidth, PGLBox proposes an asynchronous pipeline with three stages - it first samples the subgraphs from the input graph, then pulls & updates embeddings and trains GNNs on the subgraph with parameters updating queued at the end of the pipeline. Thanks to the capacity of PGLBox in handling web-scale graphs, it becomes feasible to unify the view of GNN-based recommendation tasks for multiple advertising verticals and fuse all these graphs into a unified yet huge one. We evaluate PGLBox using a bucket of realistic GNN training tasks for the recommendation, and compare the performance of PGLBox on top of a multi-GPU server (Tesla A100×8) and the legacy training system based on a 40-node MPI cluster at Baidu. The overall comparisons show that PGLBox could save up to 55% monetary cost for training GNN models, and achieve up to 14× training speedup with the same accuracy as the legacy trainer. The open-source implementation of PGLBox is available at https://github.com/PaddlePaddle/PGL/tree/main/apps/PGLBox.
Xuewu Jiao, Weibin Li 0004, Xinxuan Wu, Jiang Bian 0003, Siming Dai, Xinsheng Luo, Mingqing Hu, Zhengjie Huang, Danlei Feng, Junchao Yang 0001, Shikun Feng, Haoyi Xiong, Dianhai Yu, Shuanglong Li, Jingzhou He, Yanjun Ma
KDD16
2023 Enhancing Dynamic Image Advertising with Vision-Language Pre-training
abstract
In the multimedia era, image becomes an effective medium in search advertising. Dynamic Image Advertising (DIA), a system that matches queries with appropriate ad images and generates multimodal ads, is introduced to improve user experience and ad revenue. The core of DIA is a query-image matching module performing ad image retrieval and relevance modeling. Current query-image matching suffers from data scarcity and inconsistency, and insufficient cross-modal fusion. Also, the retrieval and relevance models are separately trained, affecting overall performance. In this paper, we propose a vision-language framework for query-image matching. It consists of two parts. First, we design a base model combining different encoders and tasks, and train it on large-scale image-text pairs to learn general multimodal representation. Then, we fine-tune the base model on advertising business data, unifying relevance modeling and retrieval through multi-objective learning. Our framework has been implemented in Baidu search advertising system "Phoneix Nest". Online evaluation shows that it improves cost per mille (CPM) and click-through rate (CTR) by 1.04% and 1.865% on the system main traffic.
Zhoufutu Wen, Zhipeng Jin, Yi Yang 0031, Xiaodong Chen 0006, Shuanglong Li
SIGIR7
2022 Enhanced Video BERT for Fast Video Advertisement Retrieval
abstract
Recently, video BERT based on cross-modal attention has achieved excellent performance in many cross-modal tasks in academia. Nevertheless, the expensive computation cost of cross-modal attention makes video BERT impractical for large-scale search in industrial applications. Inspired by the success of the tree-based deep model (TDM) in the recommendation system, we present a enhanced video BERT (EVB). It provides a practical solution to deploy the heavy video BERT for the large-scale query-to-video search. The proposed EVB overcomes the limitation of TDM relying on global features, and makes the tree structure based on a global feature compatible with version BERT using a set of local features. What’s more, we proposes a similarity-based dynamic construction to integrate the optimization of model efficiency. The proposed EVB has been deployed in our video advertising platform and brings a considerable boost in CVR and CTR for advertisers.
Yi Yang 0031, Zhipeng Jin, Xuewu Jiao, Shuanglong Li, Ping Li 0001
IEEE Big Data7