EDBT 2026 Demo / reviewers in the wild / expert
Wen Tao
dblp:118/4552
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Multi-Stage Structural Captioning Framework for Enhancing Chinese Image-to-Video Generation in BaiduabstractThe rapid advancement of video generation technology, particularly in the domain of image-to-video generation, is significantly transforming both personal and industrial applications. The performance of video generation models is heavily influenced by the quality of data, specifically video content and captions. Among these, the quality of video captions directly impacts the model's ability to follow user instructions. Currently, video captions are typically generated using vision-language model (VLM)-based methods. However, issues such as hallucinations, incomplete descriptions, and inaccuracies remain prevalent. To address these challenges, we propose a novel optimization framework for a multi-expert, structured video captioning model, specifically designed to enhance the performance of Chinese image-to-video tasks in Baidu's business scenarios. First, we construct a structured video caption dataset by combining a large language model (LLM) with human annotations, and introduce a hallucination-aware adaptive GRPO algorithm. A two-stage fine-tuning alignment strategy is employed to improve the accuracy of fine-grained content descriptions and mitigate hallucinations. Second, we design independent video sub-expert models to characterize subject motion intensity and camera movement, thereby expanding the dimensionality of video captions and addressing the issue of inaccurate sub-dimension descriptions in a single VLM. Finally, we leverage multi-level structured video captions to guide the training of the video generation model, resulting in improved quality and consistency in video generation. Extensive experiments demonstrate that our method consistently outperforms the baseline. Specifically, it achieves a 2.49% improvement in F1-score on event-level cross-validation and reduces hallucination occurrences by 30.9%. Moreover, our framework significantly enhances the quality of Chinese video generation, yielding a 2.9% increase in the VBench total score and a 9.2% improvement in the proportion of high-quality generated videos. Zhipeng Jin, Xiawei Li, Wen Tao, Yi Yang 0031, Cong Han 0002, Shuanglong Li, Zhongmin Cai |
KDD (1) | 4 |
| 2025 | How to Make Large Language Models Generate 100% Valid Molecules?abstractMolecule generation is key to drug discovery and materials science, enabling the design of novel compounds with specific properties.Large language models (LLMs) can learn to perform a wide range of tasks from just a few examples.However, generating valid molecules using representations like SMILES is challenging for LLMs in few-shot settings.In this work, we explore how LLMs can generate 100% valid molecules.We evaluate whether LLMs can use SELFIES, a representation where every string corresponds to a valid molecule, for valid molecule generation but find that LLMs perform worse with SELFIES than with SMILES.We then examine LLMs' ability to correct invalid SMILES and find their capacity limited.Finally, we introduce SmiSelf, a cross-chemical language framework for invalid SMILES correction.SmiSelf converts invalid SMILES to SELFIES using grammatical rules, leveraging SELFIES' mechanisms to correct the invalid SMILES.Experiments show that SmiSelf ensures 100% validity while preserving molecular characteristics and maintaining or even enhancing performance on other metrics.SmiSelf helps expand LLMs' practical applications in biomedicine and is compatible with all SMILES-based generative models.Code is available at https: //github.com/wentao228/SmiSelf. Wen Tao, Jing Tang 0004, Alvin Chan, Bryan Hooi, Baolong Bi, Nanyun Peng 0001, Yuansheng Liu, Yiwei Wang 0001 |
EMNLP | 1 |
| 2025 | UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text SynthesisabstractText-to-image generation has greatly advanced content creation, yet accurately rendering visual text remains a key challenge due to blurred glyphs, semantic drift, and limited style control. Existing methods often rely on pre-rendered glyph images as conditions, but these struggle to retain original font styles and color cues, necessitating complex multi-branch designs that increase model overhead and reduce flexibility. To address these issues, we propose a segmentation-guided framework that uses pixel-level visual text masks -- rich in glyph shape, color, and spatial detail -- as unified conditional inputs. Our method introduces two core components: (1) a fine-tuned bilingual segmentation model for precise text mask extraction, and (2) a streamlined diffusion model augmented with adaptive glyph conditioning and a region-specific loss to preserve textual fidelity in both content and style. Our approach achieves state-of-the-art performance on the AnyText benchmark, significantly surpassing prior methods in both Chinese and English settings. To enable more rigorous evaluation, we also introduce two new benchmarks: GlyphMM-benchmark for testing layout and glyph consistency in complex typesetting, and MiniText-benchmark for assessing generation quality in small-scale text regions. Experimental results show that our model outperforms existing methods by a large margin in both scenarios, particularly excelling at small text rendering and complex layout preservation, validating its strong generalization and deployment readiness. Yuanrui Wang, Cong Han 0002, Zhipeng Jin, Xiawei Li, SiNan Du, Wen Tao, Shuanglong Li, Yi Yang 0031, Chun Yuan 0003 |
ICCV | 7 |
| 2025 | Towards Synergistic Path-based Explanations for Knowledge Graph Completion: Exploration and EvaluationabstractKnowledge graph completion (KGC) aims to alleviate the inherent incompleteness of knowledge graphs (KGs), a crucial task for numerous applications such as recommendation systems and drug repurposing. The success of knowledge graph embedding (KGE) models provokes the question about the explainability: ``\textit{Which the patterns of the input KG are most determinant to the prediction}?'' Particularly, path-based explainers prevail in existing methods because of their strong capability for human understanding. In this paper, based on the observation that a fact is usually determined by the synergy of multiple reasoning chains, we propose a novel explainable framework, dubbed KGExplainer, to explore synergistic pathways. KGExplainer is a model-agnostic approach that employs a perturbation-based greedy search algorithm to identify the most crucial synergistic paths as explanations within the local structure of target predictions. To evaluate the quality of these explanations, KGExplainer distills an evaluator from the target KGE model, allowing for the examination of their fidelity. We experimentally demonstrate that the distilled evaluator has comparable predictive performance to the target KGE. Experimental results on benchmark datasets demonstrate the effectiveness of KGExplainer, achieving a human evaluation accuracy of 83.3\% and showing promising improvements in explainability. Code is available at \url{https://github.com/xiaomingaaa/KGExplainer} Tengfei Ma 0002, Xiang Song 0003, Wen Tao, Mufei Li, Jiani Zhang 0003, Xiaoqin Pan, Yijun Wang 0002, Bosheng Song, Xiangxiang Zeng |
ICLR | 3 |
| 2025 | Large Vison-Language Foundation Model in Baidu AIGC Image AdvertisingabstractRecent advances in generative artificial intelligence have revolutionized information retrieval and content generation, opening up new opportunities for the e-commerce industry. Alignment learning between small models and parallel corpora cannot meet current needs. The success of ChatGPT demonstrates that large models need to first establish a fundamental understanding, and then utilize high-quality corpora for generation. Having a large model foundation is indispensable. In this paper, we establish a fundamental 10B multimodal model foundation for multimodal generation tasks and propose a scene-based alignment learning approach called conditional sample supervised fine-tuning for downstream generation tasks. Meanwhile, diffusion models are known to be vulnerable to outliers in training data. To address this, we utilize an alternative diffusion loss function that preserves the high quality of generated data like the original squared L2 loss while being robust to outliers.In practical test sets, the multimodal foundation fully demonstrates its alignment and comprehension abilities for graphic and textual content. Additionally, conditional fine-tuning and the design of the loss function significantly enhance the quality of generated content. The quality rate of images has increased by 34.3 percentage points, and prompt control has improved by 19.8 percentage points. The application of our framework in Baidu Search Ads has led to significant revenue growth. For instance, ads with generated image creatives have achieved a 29% higher click-through rate (CTR), resulting in a daily consumption of 3 million yuan. Zhipeng Jin, Wen Tao, Yi Yang 0031, Cong Han 0002, Shuanglong Li |
KDD (1) | 2 |
| 2025 | Retrieval-Augmented Image Captioning and Generation with Entity Concepts Enhancement for Baidu Multimodal AdvertisingabstractRecent advancements in generative artificial intelligence are driving a significant transformation in information retrieval and content generation, creating substantial opportunities for online advertising. Text-to-image generation technology has become increasingly prevalent in advertising content production, demonstrating promising performance improvements in terms of semantic relevance and visual appeal. However, existing models often suffer from inadequate representation of entity concepts, such as prominent product brands and recognizable landmarks. This inherent limitation subsequently leads to notable deficiencies in brand tonality, industry-specific relevance, and market adaptability of the generated advertising content. To address this challenge, we propose a multimodal ad content generation framework specifically engineered for online advertising system, particularly focused on resolving the deficiency in entity concepts. Our framework is comprised of two phases: first, an image captioning module with entity-aware learning based on multimodal large language model, leveraging retrieval-augmented techniques to incorporate entity concepts into image descriptions; second, a text-to-image diffusion model refined on image-text pairs enriched with entity concepts to facilitate entity-grounded image generation. Extensive experiments validate the effectiveness of our framework, demonstrating superior performance in both image captioning and image generation compared to existing methods, particularly in the accuracy of depiction of relevant entities in advertising images. Moreover, the deployment of the framework in the system primary traffic of Baidu Search Ads, has brought significant enhancements to advertisement revenue for both advertisers and the platform. Kang Zhao 0002, Zhipeng Jin, Wen Tao, Yi Yang 0031, Cong Han 0002, Shuanglong Li, Zhongmin Cai |
SIGIR | 4 |
| 2024 | Multi-Stage Refined Visual Captioning for Baidu Ad Creatives GenerationabstractHigh-quality multimodal training data is of critical importance for improving of multimodal model performance. However, the utilization of web-crawled vision-caption pairs is hindered by the presence of noise and irrelevance, as well as a lack of Chinese data. Large Language Models (LLM) and Large Multimodal Models (LMM) has demonstrated promising performance in cross-modal understanding and generation. In light of this, we propose a Chinese visual captioning pipeline for the synthesis of high-quality data. Our pipeline is comprised of two phases: the initial training of an encoder for visual understanding; and the subsequent fine-tuning of a captioning model in a two-stage iterative human-in-the-loop process, where the captioning model incorporates the pre-trained vision encoder and LLM by a visual cross-attention querying transformer. Extensive experiments have been conducted to validate our framework, including both quantitative and qualitative evaluation of captions generated from images and videos. The synthesis pipeline has been integrated into the ad image creative generation process in Baidu Search Ads, resulting in enhanced capabilities in prompt following. Yi Yang 0031, Kang Zhao 0002, Zhipeng Jin, Wen Tao, Shuanglong Li |
CIKM | 5 |
| 2024 | Scaling Vison-Language Foundation Model to 12 Billion Parameters in Baidu Dynamic Image AdvertisingabstractDynamic image advertising is an add-on service in search advertising that matches visuals to search ads in real-time. However, the image matching system encompasses various sub-tasks with different objectives, increasing the complexity of achieving global optimization. Besides, prevalent long-tailed data poses a challenge to the multimodal representation learning in dynamic image advertising. Recently, vision-language pre-trained models have achieved remarkable performance across a variety of multimodal tasks, and implemented as the foundational representation model in electronic business scenarios. In this paper, to improve multimodal content understanding in Dynamic Image adVERtising, we present a viSion-language rEpresentation model (referred to as DIVERSE) that learns on cross-view and cross-token contrastive loss. Moreover, with large-scale curated advertising image-text data and extensive efficient training techniques, we scale DIVERSE to 12 billion parameters, which is the biggest Chinese multimodal representation model in industrial practices. Experiment results demonstrate the distinct advantages of DIVERSE12B in business datasets, with competitive performance on public benchmarks. Further evaluation in downstream applications including ad text-image retrieval, text-image relevance modeling, and image content moderation, shows that it outperforms previous separately-trained models across offline and online metrics. Moreover, DIVERSE12B has been implemented on the system primary traffic of Baidu Search Ads, bringing considerable increase to both user experience, and revenue for advertisers and search engine. Kang Zhao 0002, Zhipeng Jin, Yi Yang 0031, Wen Tao, Xiaodong Chen 0006, Cong Han 0002, Shuanglong Li |
CIKM | 5 |
| 2024 | Enhancing Baidu Multimodal Advertisement with Chinese Text-to-Image Generation via Bilingual Alignment and Caption SynthesisabstractRecent advances in generative artificial intelligence have revolutionized information retrieval and content generation, opening up new opportunities for the e-commerce industry. In particular, text-to-image generation models offer a novel approach to guiding the image generation process using natural language input, which is inspiring for multimodal search advertising. Traditional multimodal search ads require advertisers to prepare ad creatives, such as ad images, which is time-consuming and requires uniform image specifications and content quality inspection. To this end, we propose a streamlined generation framework for search ad image creatives. First, we prepare a Chinese image caption model with high-quality image-caption pairs to bootstrap training data refinement. With curated high-quality images and synthesized descriptive captions, we then train a Chinese text-to-image generation model, the largest to date, using SDXL and a 10-billion multimodal text encoder. Specifically, we introduce a two-stage bilingual multimodal representation alignment process to seamlessly integrate the text encoder with the generation model. Extensive experiments validate the effectiveness of our framework, including assessments of image captioning and image generation. The implementation of our framework in Baidu Search Ads shows significant revenue increase, For example, beauty industry ads with generated image creatives achieve a 29% higher click-through rate (CTR). Kang Zhao 0002, Zhipeng Jin, Yi Yang 0031, Wen Tao, Cong Han 0002, Shuanglong Li |
SIGIR | 5 |
| 2024 | Learning to Denoise Biomedical Knowledge Graph for Robust Molecular Interaction PredictionabstractMolecular interaction prediction plays a crucial role in forecasting unknown interactions between molecules, such as drug-target interaction (DTI) and drug-drug interaction (DDI), which are essential in the field of drug discovery and therapeutics. Although previous prediction methods have yielded promising results by leveraging the rich semantics and topological structure of biomedical knowledge graphs (KGs), they have primarily focused on enhancing predictive performance without addressing the presence of inevitable noise and inconsistent semantics. This limitation has hindered the advancement of KG-based prediction methods. To address this limitation, we propose BioKDN (BiomedicalKnowledge GraphDenoisingNetwork) for robust molecular interaction prediction. BioKDN refines the reliable structure of local subgraphs by denoising noisy links in a learnable manner, providing a general module for extracting task-relevant interactions. To enhance the reliability of the refined structure, BioKDN maintains consistent and robust semantics by smoothing relations around the target interaction. By maximizing the mutual information between reliable structure and smoothed relations, BioKDN emphasizes informative semantics to enable precise predictions. Experimental results on real-world datasets show that BioKDN surpasses state-of-the-art models in DTI and DDI prediction tasks, confirming the effectiveness and robustness of BioKDN in denoising unreliable interactions within contaminated KGs. Tengfei Ma 0002, Yujie Chen 0002, Wen Tao, Dashun Zheng, Xuan Lin, Patrick Pang 0001, Yijun Wang 0002, Longyue Wang, Bosheng Song, Xiangxiang Zeng, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Prediction of multi-relational drug-gene interaction via Dynamic hyperGraph Contrastive LearningabstractDrug-gene interaction prediction occupies a crucial position in various areas of drug discovery, such as drug repurposing, lead discovery and off-target detection. Previous studies show good performance, but they are limited to exploring the binding interactions and ignoring the other interaction relationships. Graph neural networks have emerged as promising approaches owing to their powerful capability of modeling correlations under drug-gene bipartite graphs. Despite the widespread adoption of graph neural network-based methods, many of them experience performance degradation in situations where high-quality and sufficient training data are unavailable. Unfortunately, in practical drug discovery scenarios, interaction data are often sparse and noisy, which may lead to unsatisfactory results. To undertake the above challenges, we propose a novel Dynamic hyperGraph Contrastive Learning (DGCL) framework that exploits local and global relationships between drugs and genes. Specifically, graph convolutions are adopted to extract explicit local relations among drugs and genes. Meanwhile, the cooperation of dynamic hypergraph structure learning and hypergraph message passing enables the model to aggregate information in a global region. With flexible global-level messages, a self-augmented contrastive learning component is designed to constrain hypergraph structure learning and enhance the discrimination of drug/gene representations. Experiments conducted on three datasets show that DGCL is superior to eight state-of-the-art methods and notably gains a 7.6% performance improvement on the DGIdb dataset. Further analyses verify the robustness of DGCL for alleviating data sparsity and over-smoothing issues. Wen Tao, Yuansheng Liu, Xuan Lin, Bosheng Song, Xiangxiang Zeng |
Briefings Bioinform. | 1 |
| 2018 | Deep Neural Network Based Sparse Measurement Matrix for Image Compressed SensingabstractGaussian random matrix (GRM) has been widely used to generate linear measurements in compressed sensing (CS) of natural images. However, there actually exist two disadvantages with GRM in practice. One is that GRM has large memory requirement and high computational complexity, which restrict the applications of CS. Another is that the CS measurements randomly obtained by GRM cannot provide sufficient reconstruction performances. In this paper, a Deep neural network based Sparse Measurement Matrix (DSMM) is learned by the proposed convolutional network to reduce the sampling computational complexity and improve the CS reconstruction performance. Two sub-networks are included in the proposed network, which are the sampling sub-network and the reconstruction sub-network. In the sampling sub-network, the sparsity and the normalization are both considered by the limitation of the storage and the computational complexity. In order to improve the CS reconstruction performance, a reconstruction sub-network are introduced to help enhance the sampling sub-network. So by the offline iterative training of the proposed end-to-end network, the DSMM is generated for accurate measurement and excellent reconstruction. Experimental results demonstrate that the proposed DSMM outperforms GRM greatly on representative CS reconstruction methods. Wenxue Cui, Feng Jiang 0001, Xinwei Gao, Wen Tao, Debin Zhao |
ICIP | 4 |
| 2018 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractDeep learning, e.g., convolutional neural networks (CNNs), has achieved great success in image processing and computer vision especially in high-level vision applications, such as recognition and understanding. However, it is rarely used to solve low-level vision problems such as image compression studied in this paper. Here, we move forward a step and propose a novel compression framework based on CNNs. To achieve high-quality image compression at low bit rates, two CNNs are seamlessly integrated into an end-to-end compression framework. The first CNN, named compact convolutional neural network (ComCNN), learns an optimal compact representation from an input image, which preserves the structural information and is then encoded using an image codec (e.g., JPEG, JPEG2000, or BPG). The second CNN, named reconstruction convolutional neural network (RecCNN), is used to reconstruct the decoded image with high quality in the decoding end. To make two CNNs effectively collaborate, we develop a unified end-to-end learning algorithm to simultaneously learn ComCNN and RecCNN, which facilitates the accurate reconstruction of the decoded image using RecCNN. Such a design also makes the proposed compression framework compatible with existing image coding standards. Experimental results validate that the proposed compression framework greatly outperforms several compression frameworks that use existing image coding standards with the state-of-the-art deblocking or denoising post-processing methods. Feng Jiang 0001, Wen Tao, Shaohui Liu, Jie Ren 0016, Xun Guo 0002, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractSummary form only given. Traditional image coding standards (such as JPEG and JPEG2000) make the decoded image suffer from many blocking artifacts or noises since the use of big quantization steps. To overcome this problem, we proposed an end-to-end compression framework based on two CNNs, as shown in Figure 1, which produce a compact representation for encoding using a third party coding standard and reconstruct the decoded image, respectively. To make two CNNs effectively collaborate, we develop a unified end-to-end learning framework to simultaneously learn CrCNN and ReCNN such that the compact representation obtained by CrCNN preserves the structural information of the image, which facilitates to accurately reconstruct the decoded image using ReCNN and also makes the proposed compression framework compatible with existing image coding standards. Wen Tao, Feng Jiang 0001, Shengping Zhang, Jie Ren 0016, Wuzhen Shi, Wangmeng Zuo, Xun Guo 0002, Debin Zhao |
DCC | 1 |