EDBT 2026 Demo / reviewers in the wild / expert
Yongjun Bao
dblp:52/5055
· DBLP profile ↗
35ranked-venue papers
0as first author
22since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 11 since 2021Databases, data management, data science and information retrieval · 12 · 4 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TempMe: Video Temporal Token Merging for Efficient Text-Video RetrievalabstractMost text-video retrieval methods utilize the text-image pre-trained models like CLIP as a backbone. These methods process each sampled frame independently by the image encoder, resulting in high computational overhead and limiting practical deployment. Addressing this, we focus on efficient text-video retrieval by tackling two key challenges: 1. From the perspective of trainable parameters, current parameter-efficient fine-tuning methods incur high inference costs; 2. From the perspective of model complexity, current token compression methods are mainly designed for images to reduce spatial redundancy but overlook temporal redundancy in consecutive frames of a video. To tackle these challenges, we propose Temporal Token Merging (TempMe), a parameter-efficient and training-inference efficient text-video retrieval architecture that minimizes trainable parameters and model complexity. Specifically, we introduce a progressive multi-granularity framework. By gradually combining neighboring clips, we reduce spatio-temporal redundancy and enhance temporal modeling across different frames, leading to improved efficiency and performance. Extensive experiments validate the superiority of our TempMe. Compared to previous parameter-efficient text-video retrieval methods, TempMe achieves superior performance with just 0.50M trainable parameters. It significantly reduces output tokens by 95% and GFLOPs by 51%, while achieving a 1.8X speedup and a 4.4% R-Sum improvement. With full fine-tuning, TempMe achieves a significant 7.9% R-Sum improvement, trains 1.57X faster, and utilizes 75.2% GPU memory usage. The code is available at https://github.com/LunarShen/TempMe. Leqi Shen, Tianxiang Hao 0001, Sicheng Zhao, Pengzhang Liu, Yongjun Bao, Guiguang Ding |
ICLR | 7 |
| 2025 | Spatio-Temporal Attention for Text-Video RetrievalabstractText-video retrieval, a fundamental task for associating textual descriptions with video content, has become increasingly important in the video domain. Most existing methods focus on the single-modality features only considering the knowledge within individual video or text modalities, often neglecting cross-modal interactions. However, a text description corresponds to a specific spatio-temporal content within a video, involving a certain segment of a frame sequence and distinct sub-regions within these frames. Therefore, we focus on the text-conditioned video features to bridge the modality gap. In this article, we propose Spatio-Temporal Attention for video-text retrieval, termed STAttn, which utilizes textual information to focus on the spatio-temporal video content. Our final text-conditioned video features are generated from the text-related video frames and the text-related regions within these frames. First, we propose the Spatial Text-Attention Module (STAM) to learn the spatial information within video frames. STAM introduces the text-related salient patches to capture more fine-grained details. Second, we propose the Temporal Text-Attention Module (TTAM) to learn the temporal relationships between video frames. Temporal Triplet loss is proposed in TTAM to enhance the attention toward the text-related frames. Thus, the two modules learn the text-related spatio-temporal content from both intra-frame and inter-frame aspects. Extensive experiments on three benchmark datasets, MSRVTT, ActivityNet, and DiDeMo, demonstrate that our STAttn outperforms the state-of-the-art methods. Leqi Shen, Sicheng Zhao, Pengzhang Liu, Yongjun Bao, Guiguang Ding |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Patch-Aware Sample Selection for Efficient Masked Image ModelingabstractNowadays sample selection is drawing increasing attention. By extracting and training only on the most informative subset, sample selection can effectively reduce the training cost. Although sample selection is effective in conventional supervised learning, applying it to Masked Image Modeling (MIM) still poses challenges due to the gap between sample-level selection and patch-level pre-training. In this paper, we inspect the sample selection in MIM pre-training and find the basic selection suffers from performance degradation. We attribute this degradation primarily to 2 factors: the random mask strategy and the simple averaging function. We then propose Patch-Aware Sample Selection (PASS), including a low-cost Dynamic Trained Mask Predictor (DTMP) and Weighted Selection Score (WSS). DTMP consistently masks the informative patches in samples, ensuring a relatively accurate representation of selection score. WSS enhances the selection score using patch-level disparity. Extensive experiments show the effectiveness of PASS in selecting the most informative subset and accelerating pretraining. PASS exhibits superior performance across various datasets, MIM methods, and downstream tasks. Particularly, PASS improves MAE by 0.7% on ImageNet-1K while utilizing only 37% data budget and achieves ~1.7x speedup. Zhengyang Zhuge, Yongjun Bao, Peisong Wang 0001, Jian Cheng 0001 |
AAAI | 4 |
| 2024 | PYRA: Parallel Yielding Re-activation for Training-Inference Efficient Task Adaptation
Yizhe Xiong, Hui Chen 0013, Tianxiang Hao 0001, Zijia Lin, Jungong Han, Yuesong Zhang, Yongjun Bao, Guiguang Ding |
ECCV (9) | 8 |
| 2024 | GraphHI: Boosting Graph Neural Networks for Large-Scale GraphsabstractTo analyze and process graph data, researchers have proposed Graph Neural Network (GNN) models. In this paper, we focus on methods for boosting the performance of existing GNN models and propose GraphHI, a GNN framework that integrates Hidden Insights to enhance the performance of a given GNN model. We propose to utilize both inter-model and intra-model hidden insights. The inter-model hidden insights encompass the embedding vectors and logit vectors derived from other pretrained models using the same graph data. The intra-model hidden insights incorporate the embedding vectors of other nodes from the same GNN model. To optimize the suitability of hidden insights for GNN model training, we conduct a theoretical analysis of the influence of various forms of the transformed logits and the parameter$T$in the data transformation function. Based on this analysis, a method for setting dynamic personalized parameters in the data transformation is proposed, which is tailored to the current state of each node in the GNN model. To integrate multiple sources of hidden insights, we propose ALC, an algorithm that dynamically sets appropriate combination coefficients for various loss terms. The experimental results show that GraphHI can boost the performance of GNN models using different pretrained models in four different tasks. Hao Feng 0007, Chaokun Wang, Ziyang Liu 0004, Yunkai Lou, Xiaokun Zhu, Yongjun Bao, Weipeng Yan |
ICDE | 7 |
| 2024 | TaD: A Plug-and-Play Task-Aware Decoding Method to Better Adapt LLMs on Downstream Tasks
Hui Chen 0013, Zijia Lin, Jungong Han, Lixing Gong, Yongjun Bao, Guiguang Ding |
IJCAI | 7 |
| 2024 | Multi-Label Learning with Block Diagonal LabelsabstractCollecting large-scale multi-label data with full labels is difficult for real-world scenarios. Many existing studies have tried to address the issue of missing labels caused by annotation but ignored the difficulties encountered during the annotation process. We find that the high annotation workload can be attributed to two reasons: (1) Annotators are required to identify labels on widely varying visual concepts. (2) Exhaustively annotating the entire dataset with all the labels becomes notably difficult and time-consuming. In this paper, we propose a new setting, i.e. block diagonal labels, to reduce the workload on both sides. The numerous categories can be divided into different subsets based on semantics and relevance. Each annotator can only focus on its own subset of labels so that only a small set of highly relevant labels are required to be annotated per image. To deal with the issue of such missing labels, we introduce a simple yet effective method that does not require any prior knowledge of the dataset. In practice, we propose an Adaptive Pseudo-Labeling method to predict the unknown labels with less noise. Formal analysis is conducted to evaluate the superiority of our setting. Extensive experiments are conducted to verify the effectiveness of our method on multiple widely used benchmarks. Leqi Shen, Sicheng Zhao, Hui Chen 0013, Jundong Zhou, Pengzhang Liu, Yongjun Bao, Guiguang Ding |
ACM Multimedia | 7 |
| 2024 | Parquet-Based CTR Model Training in Production EnvironmentabstractCTR(click through rate) model has played an important role in modern recommendation systems. Most of the recommendation models in industrial scenario are trained by TensorFlow. However, we observed that, TFRecord, the native k-v data format in TensorFlow, is not the best choice for CTR training. Those keys take up to 54% of the storage space in TFRecord formatted training data. To overcome this, we introduce Apache Parquet, a column-oriented data format, into CTR tasks to improve spatial efficiency. Besides, to use in production environment, we further give some high performance implementations of Parquet training scheme. Firstly, GPU data preprocessing method is adopted in replace of original Spark based solution to generate Parquet training data and accelerate data preprocessing. Secondly, we modify data loader in TensorFlow to consume the Parquet training data with high efficiency. Experimental results show that, the size of preprocessed Criteo dataset is 95.29% smaller in comparison to TFRecord and the data preprocessing time also reduces 99.6%. Without any model performance damage, we speed up the training process by 1.45x. Our scheme has applied to our internal business and has obtained similar performance benefits. Jinrong Guo, Biyu Zhou, Xiaokun Zhu, Yongjun Bao, Jizhong Han, Songlin Hu 0001 |
SMC | 5 |
| 2023 | Blending Advertising with Organic Content in E-commerce via Virtual BidsabstractIt has become increasingly common that sponsored content (i.e., paid ads) and non-sponsored content are jointly displayed to users, especially on e-commerce platforms. Thus, both of these contents may interact together to influence their engagement behaviors. In general, sponsored content helps brands achieve their marketing goals and provides ad revenue to the platforms. In contrast, non-sponsored content contributes to the long-term health of the platform through increasing users' engagement. A key conundrum to platforms is learning how to blend both of these contents allowing their interactions to be considered and balancing these business objectives. This paper proposes a system built for this purpose and applied to product detail pages of JD.COM, an e-commerce company. This system achieves three objectives: (a) Optimization of competing business objectives via Virtual Bids allowing the expressiveness of the valuation of the platform for these objectives. (b) Modeling the users' click behaviors considering explicitly the influence exerted by the sponsored and non-sponsored content displayed alongside through a deep learning approach. (c) Consideration of a Vickrey-Clarke-Groves (VCG) Auction design compatible with the allocation of ads and its induced externalities. Experiments are presented demonstrating the performance of the proposed system. Moreover, our approach is fully deployed and serves all traffic through JD.COM's mobile application. Carlos Carrion, Harikesh S. Nair, Xianghong Luo, Yulin Lei, Peiqin Gu, Xiliang Lin, Junsheng Jin, Fanan Zhu, Changping Peng, Yongjun Bao, Zhangang Lin, Weipeng Yan, Jingping Shao |
AAAI | 12 |
| 2023 | Exploring Structured Semantic Prior for Multi Label Recognition with Incomplete LabelsabstractMulti-label recognition (MLR) with incomplete labels is very challenging. Recent works strive to explore the image-to-label correspondence in the vision-language model, i.e., CLIP [22], to compensate for insufficient annotations. In spite of promising performance, they generally overlook the valuable prior about the label-to-label correspondence. In this paper, we advocate remedying the deficiency of label supervision for the MLR with incomplete labels by deriving a structured semantic prior about the label-to-label corre-spondence via a semantic prior prompter. We then present a novel Semantic Correspondence Prompt Network (SCP-Net), which can thoroughly explore the structured semantic prior. A Prior-Enhanced Self-Supervised Learning method is further introduced to enhance the use of the prior. Comprehensive experiments and analyses on several widely used benchmark datasets show that our method significantly out-performs existing methods on all datasets, well demonstrating the effectiveness and the superiority of our method. Our code will be available at https://github.com/jameslahm/SCPNet. Zixuan Ding, Hui Chen 0013, Qiang Zhang 0020, Pengzhang Liu, Yongjun Bao, Weipeng Yan, Jungong Han |
CVPR | 6 |
| 2023 | DynaMS: Dyanmic Margin Selection for Efficient Deep Learning
Jingwei Zhuo, Xupeng Shi, Lixing Gong, Tong Tao, Pengzhang Liu, Yongjun Bao, Weipeng Yan |
ICLR | 9 |
| 2023 | Hierarchical Prompt Learning Using CLIP for Multi-label Classification with Single Positive LabelsabstractCollecting full annotations to construct multi-label datasets is difficult and labor-consuming. As an effective solution to relieve the annotation burden, single positive multi-label learning (SPML) draws increasing attention from both academia and industry. It only annotates each image with one positive label, leaving other labels unobserved. Therefore, existing methods strive to explore the cue of unobserved labels to compensate for the insufficiency of label supervision. Though achieving promising performance, they generally consider labels independently, leaving out the inherent hierarchical semantic relationship among labels which reveals that labels can be clustered into groups. In this paper, we propose a hierarchical prompt learning method with a novel Hierarchical Semantic Prompt Network (HSPNet) to harness such hierarchical semantic relationships using a large-scale pretrained vision and language model, i.e., CLIP, for SPML. We first introduce a Hierarchical Conditional Prompt (HCP) strategy to grasp the hierarchical label-group dependency. Then we equip a Hierarchical Graph Convolutional Network (HGCN) to capture the high-order inter-label and inter-group dependencies. Comprehensive experiments and analyses on several benchmark datasets show that our method significantly outperforms the state-of-the-art methods, well demonstrating its superiority and effectiveness. Our code will be available at https://github.com/jameslahm/HSPNet. Hui Chen 0013, Zijia Lin, Zixuan Ding, Pengzhang Liu, Yongjun Bao, Weipeng Yan, Guiguang Ding |
ACM Multimedia | 6 |
| 2022 | LEGO-ABSA: A Prompt-based Task Assemblable Unified Generative Framework for Multi-task Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) has received increasing attention recently. ABSA can be divided into multiple tasks according to the different extracted elements. Existing generative methods usually treat the output as a whole string rather than the combination of different elements and only focus on a single task at once. This paper proposes a unified generative multi-task framework that can solve multiple ABSA tasks by controlling the type of task prompts consisting of multiple element prompts. Further, the proposed approach can train on simple tasks and transfer to difficult tasks by assembling task prompts, like assembling Lego bricks. We conduct experiments on six ABSA tasks across multiple benchmarks. Our proposed multi-task approach achieves new state-of-the-art results in almost all tasks and competitive results in task transfer scenarios. Tianhao Gao, Pengzhang Liu, Yongjun Bao, Weipeng Yan |
COLING | 7 |
| 2022 | Dynamic Graph Modeling for Weakly-Supervised Temporal Action LocalizationabstractWeakly supervised action localization is a challenging task that aims to localize action instances in untrimmed videos given only video-level supervision. Existing methods mostly distinguish action from background via attentive feature fusion with RGB and optical flow modalities. Unfortunately, this strategy fails to retain the distinct characteristics of each modality, leading to inaccurate localization under hard-to-discriminate cases such as action-context interference and in-action stationary period. As an action is typically comprised of multiple stages, an intuitive solution is to model the relation between the finer-grained action segments to obtain a more detailed analysis. In this paper, we propose a dynamic graph-based method, namely DGCNN, to explore the two-stream relation between action segments. To be specific, segments within a video which are likely to be actions are dynamically selected to construct an action graph. For each graph, a triplet adjacency matrix is devised to explore the temporal and contextual correlations between the pseudo action segments, which consists of three components, i.e., mutual importance, feature similarity, and high-level contextual similarity. The two-stream dynamic pseudo graphs, along with the pseudo background segments, are used to derive more detailed video representation. For action localization, a non-local based temporal refinement module is proposed to fully leverage the temporal consistency between consecutive segments. Experimental results on three datasets, i.e., THUMOS14, ActivityNet v1.2 and v1.3, demonstrate that our method is superior to the state-of-the-arts. Haichao Shi, Xiaoyu Zhang 0002, Lixing Gong, Yong Li 0034, Yongjun Bao |
ACM Multimedia | 6 |
| 2022 | Adaptive Experimentation with Delayed Binary FeedbackabstractConducting experiments with objectives that take significant delays to materialize (e.g. conversions, add-to-cart events, etc.) is challenging. Although the classical “split sample testing” is still valid for the delayed feedback, the experiment will take longer to complete, which also means spending more resources on worse-performing strategies due to their fixed allocation schedules. Alternatively, adaptive approaches such as “multi-armed bandits” are able to effectively reduce the cost of experimentation. But these methods generally cannot handle delayed objectives directly out of the box. This paper presents an adaptive experimentation solution tailored for delayed binary feedback objectives by estimating the real underlying objectives before they materialize and dynamically allocating variants based on the estimates. Experiments show that the proposed method is more efficient for delayed feedback compared to various other approaches and is robust in different settings. In addition, we describe an experimentation product powered by this algorithm. This product is currently deployed in the online experimentation platform of JD.com, a large e-commerce company and a publisher of digital ads. Carlos Carrion, Xiliang Lin, Fuhua Ji, Yongjun Bao, Weipeng Yan |
WWW | 5 |
| 2022 | Bidirectional difference locating and semantic consistency reasoning for change captioningabstractChange captioning is an emerging task to describe the changes between a pair of images. The difficulty in this task is to discover the differences between the two images. Recently, some methods have been proposed to address this problem. However, they all employ unidirectional difference localization to identify the changes. This can lead to ambiguity about the nature of the changes. Instead, we propose a framework with bidirectional difference localization and semantic consistency reasoning to describe the image changes. First, we locate the changes in the two images by capturing bidirectional differences. Then we design a decoder with spatial-channel attention to generate the change caption. Finally, we introduce semantic consistency reasoning to constrain our bidirectional difference localization module and spatial-channel attention module. Extensive experiments on three public data sets show that the performance of our proposed model outperforms the state-of-the-art change captioning models by a large margin. Yaoqi Sun, Liang Li 0003, Tongyv Lu, Bolun Zheng, Chenggang Yan 0001, Yongjun Bao, Guiguang Ding, Gregory Slabaugh |
Int. J. Intell. Syst. | 8 |
| 2021 | Probing Product Description Generation via Posterior DistillationabstractIn product description generation (PDG), the user-cared aspect is critical for the recommendation system, which can not only improve user's experiences but also obtain more clicks. High-quality customer reviews can be considered as an ideal source to mine user-cared aspects. However, in reality, a large number of new products (known as long-tailed commodities) cannot gather sufficient amount of customer reviews, which brings a big challenge in the product description generation task. Existing works tend to generate the product description solely based on item information, i.e., product attributes or title words, which leads to tedious contents and cannot attract customers effectively. To tackle this problem, we propose an adaptive posterior network based on Transformer architecture that can utilize user-cared information from customer reviews. Specifically, we first extend the self-attentive Transformer encoder to encode product titles and attributes. Then, we apply an adaptive posterior distillation module to utilize useful review information, which integrates user-cared aspects to the generation process. Finally, we apply a Transformer-based decoding phase with copy mechanism to automatically generate the product description. Besides, we also collect a large-scare Chinese product description dataset to support our work and further research in this field. Experimental results show that our model is superior to traditional generative models in both automatic indicators and human evaluation. Haolan Zhan, Hainan Zhang 0001, Hongshen Chen, Lei Shen 0001, Zhuoye Ding, Yongjun Bao, Weipeng Yan, Yanyan Lan |
AAAI | 6 |
| 2021 | Augmenting Knowledge-grounded Conversations with Sequential Knowledge TransitionabstractHaolan Zhan, Hainan Zhang, Hongshen Chen, Zhuoye Ding, Yongjun Bao, Yanyan Lan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Haolan Zhan, Hainan Zhang 0001, Hongshen Chen, Zhuoye Ding, Yongjun Bao, Yanyan Lan |
NAACL-HLT | 5 |
| 2021 | Underestimation Refinement: A General Enhancement Strategy for Exploration in Recommendation SystemsabstractClick-through rate (CTR) prediction based on deep neural networks has made significant progress in recommendation systems. However, these methods often suffer from CTR underestimation due to insufficient impressions for long-tail items. When formalizing CTR prediction as a contextual bandit problem, exploration methods provide a natural solution addressing this issue. In this paper, we first benchmark state-of-the-art exploration methods in the recommendation system setting. We find that the combination of gradient-based uncertainty modeling and Thompson Sampling achieves a significant advantage. On the basis of the benchmark, we further propose a general enhancement strategy, Underestimation Refinement (UR), which explicitly incorporates the prior knowledge that insufficient impressions likely leads to CTR underestimation. This strategy is applicable to almost all the existing exploration methods. Experimental results validate UR's effectiveness, achieving consistent improvement across all baseline exploration methods. Yuhai Song, Lu Wang 0031, Haoming Dang, Jing Guan, Xiwei Zhao, Changping Peng, Yongjun Bao, Jingping Shao |
SIGIR | 8 |
| 2021 | Trustworthy authorization method for security in Industrial Internet of Things
Yang Zhao 0027, Yongjun Bao, Houbing Song |
Ad Hoc Networks | 3 |
| 2021 | Dynamic Selective Network for RGB-D Salient Object DetectionabstractRGB-D saliency detection is receiving more and more attention in recent years. There are many efforts have been devoted to this area, where most of them try to integrate the multi-modal information, i.e. RGB images and depth maps, via various fusion strategies. However, some of them ignore the inherent difference between the two modalities, which leads to the performance degradation when handling some challenging scenes. Therefore, in this paper, we propose a novel RGB-D saliency model, namely Dynamic Selective Network (DSNet), to perform salient object detection (SOD) in RGB-D images by taking full advantage of the complementarity between the two modalities. Specifically, we first deploy a cross-modal global context module (CGCM) to acquire the high-level semantic information, which can be used to roughly locate salient objects. Then, we design a dynamic selective module (DSM) to dynamically mine the cross-modal complementary information between RGB images and depth maps, and to further optimize the multi-level and multi-scale information by executing the gated and pooling based selection, respectively. Moreover, we conduct the boundary refinement to obtain high-quality saliency maps with clear boundary details. Extensive experiments on eight public RGB-D datasets show that the proposed DSNet achieves a competitive and excellent performance against the current 17 state-of-the-art RGB-D SOD models. Hongfa Wen, Chenggang Yan 0001, Xiaofei Zhou 0003, Runmin Cong, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Yongjun Bao, Guiguang Ding |
IEEE Trans. Image Process. | 8 |
| 2021 | Scene Segmentation With Dual Relation-Aware Attention NetworkabstractIn this article, we propose a Dual Relation-aware Attention Network (DRANet) to handle the task of scene segmentation. How to efficiently exploit context is essential for pixel-level recognition. To address the issue, we adaptively capture contextual information based on the relation-aware attention mechanism. Especially, we append two types of attention modules on the top of the dilated fully convolutional network (FCN), which model the contextual dependencies in spatial and channel dimensions, respectively. In the attention modules, we adopt a self-attention mechanism to model semantic associations between any two pixels or channels. Each pixel or channel can adaptively aggregate context from all pixels or channels according to their correlations. To reduce the high cost of computation and memory caused by the abovementioned pairwise association computation, we further design two types of compact attention modules. In the compact attention modules, each pixel or channel is built into association only with a few numbers of gathering centers and obtains corresponding context aggregation over these gathering centers. Meanwhile, we add a cross-level gating decoder to selectively enhance spatial details that boost the performance of the network. We conduct extensive experiments to validate the effectiveness of our network and achieve new state-of-the-art segmentation performance on four challenging scene segmentation data sets, i.e., Cityscapes, ADE20K, PASCAL Context, and COCO Stuff data sets. In particular, a Mean IoU score of 82.9% on the Cityscapes test set is achieved without using extra coarse annotated data. Jun Fu 0005, Jing Liu 0001, Jie Jiang 0016, Yong Li 0034, Yongjun Bao, Hanqing Lu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Decoupled Graph Convolution Network for Inferring Substitutable and Complementary ItemsabstractInferring substitutable and complementary items is an important and fundamental concern for recommendation in e-commerce websites. However, the item relationships in real-world are usually heterogeneous, posing great challenges to conventional methods that can only deal with homogeneous relationships. More specifically, for this problem, there is a lack of in-depth investigation on 1) decoupling item semantics for modeling heterogeneous item relationships, and at the same time, 2) incorporating mutual influence between different relationships. To fill this gap, we propose a novel solution, namely Decoupled Graph Convolutional Network (DecGCN), to solve the problem of inferring substitutable and complementary items. DecGCN is designed to model item substitutability and complementarity in separated embedding spaces, and is equipped with a two-step integration scheme,where inherent influences between 1) different graph structures and 2) different item semantics are captured. Our experiments on three real-world datasets demonstrate that DecGCN is more effective than the state-of-the-art baselines for the problem at hand. We also conduct offline and online A/B tests on large-scale industrial data, where the results show that DecGCN is effective to be deployed in real-world applications. We release the codes at https://github.com/liuyiding1993/CIKM2020_DecGCN. Yulong Gu, Zhuoye Ding, Junchao Gao, Yongjun Bao, Weipeng Yan |
CIKM | 6 |
| 2020 | Dimension Relation Modeling for Click-Through Rate PredictionabstractEmbedding mechanism plays an important role in Click-Through-Rate (CTR) prediction. Essentially, it tries to learn a new feature space with some learned latent properties as the basis, and maps the high dimensional and categorical raw data to dense, rich and expressive representations, i.e., the embedding features. Current researches usually focus on learning the interactions through operations on the whole embedding features without considering the relations among the learned latent properties. In this paper, we find it has clear positive effects on CTR prediction to model such relations and propose a novel Dimension Relation Module (DRM) to capture them through dimension recalibration. We show that DRM can improve the performance of existing models consistently and the improvements are more obvious when the embedding dimension is higher. We further boost Field-wise and Element-wise embedding methods with our DRM and name this new model FED network. Extensive experiments demonstrate that FED is very powerful in CTR prediction task and achieves new state-of-the-art results on Criteo, Avazu and JD.com datasets. Zihao Zhao 0008, Zhiwei Fang, Yong Li 0034, Changping Peng, Yongjun Bao, Weipeng Yan |
CIKM | 5 |
| 2020 | An Attention-based Model for Conversion Rate Prediction with Delayed Feedback via Post-click CalibrationabstractConversion rate (CVR) prediction is becoming increasingly important in the multi-billion dollar online display advertising industry. It has two major challenges: firstly, the scarce user history data is very complicated and non-linear; secondly, the time delay between the clicks and the corresponding conversions can be very large, e.g., ranging from seconds to weeks. Existing models usually suffer from such scarce and delayed conversion behaviors. In this paper, we propose a novel deep learning framework to tackle the two challenges. Specifically, we extract the pre-trained embedding from impressions/clicks to assist in conversion models and propose an inner/self-attention mechanism to capture the fine-grained personalized product purchase interests from the sequential click data. Besides, to overcome the time-delay issue, we calibrate the delay model by learning dynamic hazard function with the abundant post-click data more in line with the real distribution. Empirical experiments with real-world user behavior data prove the effectiveness of the proposed method. Yumin Su, Liang Zhang 0042, Quanyu Dai, Bo Zhang 0086, Jinyao Yan, Dan Wang 0002, Yongjun Bao, Sulong Xu, Weipeng Yan |
IJCAI | 7 |
| 2020 | Category-Specific CNN for Visual-aware CTR Prediction at JD.comabstractAs one of the largest B2C e-commerce platforms in China, JD.com also powers a leading advertising system, serving millions of advertisers with fingertip connection to hundreds of millions of customers. In our system, as well as most e-commerce scenarios, ads are displayed with images. This makes visual-aware Click Through Rate (CTR) prediction of crucial importance to both business effectiveness and user experience. Existing algorithms usually extract visual features using off-the-shelf Convolutional Neural Networks (CNNs) and late fuse the visual and non-visual features for the finally predicted CTR. Despite being extensively studied, this field still face two key challenges. First, although encouraging progress has been made in offline studies, applying CNNs in real systems remains non-trivial, due to the strict requirements for efficient end-to-end training and low-latency online serving. Second, the off-the-shelf CNNs and late fusion architectures are suboptimal. Specifically, off-the-shelf CNNs were designed for classification thus never take categories as input features. While in e-commerce, categories are precisely labeled and contain abundant visual priors that will help the visual modeling. Unaware of the ad category, these CNNs may extract some unnecessary category-unrelated features, wasting CNN's limited expression ability. To overcome the two challenges, we propose Category-specific CNN (CSCNN) specially for CTR prediction. CSCNN early incorporates the category knowledge with a light-weighted attention-module on each convolutional layer. This enables CSCNN to extract expressive category-specific visual patterns that benefit the CTR prediction. Offline experiments on benchmark and a 10 billion scale real production dataset from JD, together with an Online A/B test show that CSCNN outperforms all compared state-of-the-art algorithms. We also build a highly efficient infrastructure to accomplish end-to-end training with CNN on the 10 billion scale real production dataset within 24 hours, and meet the low latency requirements of online system (20ms on CPU). CSCNN is now deployed in the search advertising system of JD, serving the main traffic of hundreds of millions of active users. Hao Yang 0030, Xiwei Zhao, Sulong Xu, Wenjie Niu, Xiaokun Zhu, Yongjun Bao, Weipeng Yan |
KDD | 10 |
| 2020 | Kalman Filtering Attention for User Behavior Modeling in CTR PredictionabstractClick-through rate (CTR) prediction is one of the fundamental tasks for e-commerce search engines. As search becomes more personalized, it is necessary to capture the user interest from rich behavior data. Existing user behavior modeling algorithms develop different attention mechanisms to emphasize query-relevant behaviors and suppress irrelevant ones. Despite being extensively studied, these attentions still suffer from two limitations. First, conventional attentions mostly limit the attention field only to a single user's behaviors, which is not suitable in e-commerce where users often hunt for new demands that are irrelevant to any historical behaviors. Second, these attentions are usually biased towards frequent behaviors, which is unreasonable since high frequency does not necessarily indicate great importance. To tackle the two limitations, we propose a novel attention mechanism, termed Kalman Filtering Attention (KFAtt), that considers the weighted pooling in attention as a maximum a posteriori (MAP) estimation. By incorporating a priori, KFAtt resorts to global statistics when few user behaviors are relevant. Moreover, a frequency capping mechanism is incorporated to correct the bias towards frequent behaviors. Offline experiments on both benchmark and a 10 billion scale real production dataset, together with an Online A/B test, show that KFAtt outperforms all compared state-of-the-arts. KFAtt has been deployed in the ranking system of JD.com, one of the largest B2C e-commerce websites in China, serving the main traffic of hundreds of millions of active users. Xiwei Zhao, Sulong Xu, Junsheng Jin, Yongjun Bao, Weipeng Yan |
NeurIPS | 10 |
| 2020 | Smart Targeting: A Relevance-driven and Configurable Targeting Framework for Advertising SystemabstractTargeting system is an essential part of computational advertising. It allows advertisers to select and reach their targeted users. Due to various advertising goals and the demand for making budget plans, advertisers have a strong will to configure the final targeting results, or they can become very cautious in spending money on advertising campaigns. Meanwhile, to guarantee the advertising performance, the targeted users should also be relevant to the ads of the advertisers. Recent targeting methods are mainly based on tags produced by the Data Management Platform (DMP) which is easy for the advertisers to configure the targeting results. However, in such methods, the relevance between the targeted users and ads is not technically evaluated and cannot be guaranteed. The biggest challenge is that it is hard for a machine learning model to both model the relevance and take account of the advertiser’s configuration demands. In this paper, we propose a novel relevance-driven and configurable targeting framework called Smart Targeting to solve the problem. Specifically, different from Tag-wise Targeting, we first use a relevance model to retrieve the most relevant users for the ads. To further enable the advertisers to configure the final results, we develop a Delay Intervention Mechanism to leverage the power of DMP. As far as we know, this is the first attempt of combining relevance modeling and advertiser intervention into a unified targeting system. We implement and evaluate our framework on JD.com platform with over 300 million users and the results show that it can bring significant improvements to the core indicators such as CTR and eCPM. The long term monitoring also demonstrates that Smart Targeting gradually becomes the most popular targeting tool after its release. Yong Li 0034, Zihao Zhao 0008, Zhiwei Fang, Yafei Yao, Changping Peng, Yongjun Bao, Weipeng Yan |
RecSys | 7 |
| 2020 | A Joint Dynamic Ranking System with DNN and Vector-based Clustering BanditabstractThe ad-ranking module is the core of the advertising recommender system. Existing ad-ranking modules are mainly based on the deep neural network click-through rate prediction model. Recently an innovative ad-ranking paradigm called DNN-MAB has been introduced to address DNN-only paradigms’ weakness in perceiving highly dynamic user intent over time. We introduce the DNN-MAB paradigm into our ad-ranking system to alleviate the Matthew effect that harms the user experience. Due to data sparsity, however, the actual performance of DNN-MAB is lower than expected. In this paper, we propose an innovative ad-ranking paradigm called DNN-VMAB to solve these problems. Based on vectorization and clustering, it utilizes latent collaborative information in user behavior data to find a set of ads with higher relativity and diversity. As an integration of the essences of classical collaborative filtering, deep click-through rate prediction model, and contextual multi-armed bandit, it can improve platform revenue and user experience. Both offline and online experiments show the advantage of our new algorithm over DNN-MAB and some other existing algorithms. Yong Li 0034, Changping Peng, Yongjun Bao, Weipeng P. Yan |
RecSys | 6 |
| 2020 | MaHRL: Multi-goals Abstraction Based Deep Hierarchical Reinforcement Learning for RecommendationsabstractAs huge commercial value of the recommender system, there has been growing interest to improve its performance in recent years. The majority of existing methods have achieved great improvement on the metric of click, but perform poorly on the metric of conversion possibly due to its extremely sparse feedback signal. To track this challenge, we design a novel deep hierarchical reinforcement learning based recommendation framework to model consumers' hierarchical purchase interest. Specifically, the high-level agent catches long-term sparse conversion interest, and automatically sets abstract goals for low-level agent, while the low-level agent follows the abstract goals and catches short-term click interest via interacting with real-time environment. To solve the inherent problem in hierarchical reinforcement learning, we propose a novel multi-goals abstraction based deep hierarchical reinforcement learning algorithm (MaHRL). Our proposed algorithm contains three contributions: 1) the high-level agent generates multiple goals to guide the low-level agent in different sub-periods, which reduces the difficulty of approaching high-level goals; 2) different goals share the same state encoder structure and its parameters, which increases the update frequency of the high-level agent and thus accelerates the convergence of our proposed algorithm; 3) an appreciated reward assignment mechanism is designed to allocate rewards in each goal so as to coordinate different goals in a consistent direction. We evaluate our proposed algorithm based on a real-world e-commerce dataset and validate its effectiveness. Dongyang Zhao, Liang Zhang 0042, Bo Zhang 0010, Lizhou Zheng, Yongjun Bao, Weipeng Yan |
SIGIR | 5 |
| 2020 | Contextual deconvolution network for semantic segmentation
Jun Fu 0005, Jing Liu 0001, Yong Li 0034, Yongjun Bao, Weipeng Yan, Zhiwei Fang, Hanqing Lu |
Pattern Recognit. | 4 |
| 2019 | Regularized Adversarial Sampling and Deep Time-aware Attention for Click-Through Rate PredictionabstractImproving the performance of click-through rate (CTR) prediction remains one of the core tasks in online advertising systems. With the rise of deep learning, CTR prediction models with deep networks remarkably enhance model capacities. In deep CTR models, exploiting users' historical data is essential for learning users' behaviors and interests. As existing CTR prediction works neglect the importance of the temporal signals when embed users' historical clicking records, we propose a time-aware attention model which explicitly uses absolute temporal signals for expressing the users' periodic behaviors and relative temporal signals for expressing the temporal relation between items. Besides, we propose a regularized adversarial sampling strategy for negative sampling which eases the classification imbalance of CTR data and can make use of the strong guidance provided by the observed negative CTR samples. The adversarial sampling strategy significantly improves the training efficiency, and can be co-trained with the time-aware attention model seamlessly. Experiments are conducted on real-world CTR datasets from both in-station and out-station advertising places. Yikai Wang 0001, Liang Zhang 0042, Quanyu Dai, Fuchun Sun 0001, Bo Zhang 0010, Weipeng Yan, Yongjun Bao |
CIKM | 8 |
| 2019 | Dual Attention Network for Scene SegmentationabstractIn this paper, we address the scene segmentation task by capturing rich contextual dependencies based on the self-attention mechanism. Unlike previous works that capture contexts by multi-scale features fusion, we propose a Dual Attention Networks (DANet) to adaptively integrate local features with their global dependencies. Specifically, we append two types of attention modules on top of traditional dilated FCN, which model the semantic interdependencies in spatial and channel dimensions respectively. The position attention module selectively aggregates the features at each position by a weighted sum of the features at all positions. Similar features would be related to each other regardless of their distances. Meanwhile, the channel attention module selectively emphasizes interdependent channel maps by integrating associated features among all channel maps. We sum the outputs of the two attention modules to further improve feature representation which contributes to more precise segmentation results. We achieve new state-of-the-art segmentation performance on three challenging scene segmentation datasets, i.e., Cityscapes, PASCAL Context and COCO Stuff dataset. In particular, a Mean IoU score of 81.5% on Cityscapes test set is achieved without using coarse data. Jun Fu 0005, Jing Liu 0001, Haijie Tian, Yong Li 0034, Yongjun Bao, Zhiwei Fang, Hanqing Lu |
CVPR | 5 |
| 2019 | Adaptive Context Network for Scene ParsingabstractRecent works attempt to improve scene parsing performance by exploring different levels of contexts, and typically train a well-designed convolutional network to exploit useful contexts across all pixels equally. However, in this paper, we find that the context demands are varying from different pixels or regions in each image. Based on this observation, we propose an Adaptive Context Network (ACNet) to capture the pixel-aware contexts by a competitive fusion of global context and local context according to different per-pixel demands. Specifically, when given a pixel, the global context demand is measured by the similarity between the global feature and its local feature, whose reverse value can also be used to measure the local context demand. We model the two demanding measurements by the proposed global context module and local context module, respectively, to generate their adaptive contextual features. Furthermore, we import multiple such modules to build several adaptive context blocks in different levels of network to obtain a coarse-to-fine result. Finally, comprehensive experimental evaluations demonstrate the effectiveness of the proposed ACNet, and new state-of-the-arts performances are achieved on all four public datasets, i.e. Cityscapes, ADE20K, PASCAL Context, and COCO Stuff. Jun Fu 0005, Jing Liu 0001, Yong Li 0034, Yongjun Bao, Jinhui Tang 0001, Hanqing Lu |
ICCV | 5 |
| 2018 | A Practical Deep Online Ranking System in E-commerce Recommendation
Yan Yan 0024, Zitao Liu 0001, Wentao Guo 0008, Weipeng P. Yan, Yongjun Bao |
ECML/PKDD (3) | 6 |