EDBT 2026 Demo / reviewers in the wild / expert
Yipeng Yu
dblp:129/5062
· DBLP profile ↗
16ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RecGPT-Mobile: On-Device Large Language Models for User Intent Understanding in Taobao Feed RecommendationabstractPredicting a user's next search query from recent interaction behaviors is a critical problem in modern e-commerce systems, particularly in scenarios where user intent evolves rapidly. Large Language Models (LLMs) offer strong semantic reasoning capabilities and have recently been adopted to enhance training data construction for next-query prediction. However, due to resource constraints on mobile devices, existing applications are deployed on cloud servers, resulting in high inference costs. In this paper, we propose RecGPT-Mobile, a framework that designs a lightweight LLM-based intent understanding agent to improve recommendation quality in mobile e-commerce scenarios. By deploying LLM directly on mobile devices, our approach can capture the evolving interests of users more quickly and adjust the recommendation results in real time. Extensive offline analyzes and online experiments demonstrate that our method significantly improves the accuracy of recommendation results, laying a practical path for LLM deployment in production-scale recommendation systems on mobile devices, as well as a scalable solution for integrating LLMs into real-world next-query prediction systems. Weipeng Huang, Dimin Wang, Yuning Jiang 0001, Zhaode Wang, Chengfei Lv, Junqing Wu, Yipeng Yu |
SIGIR | 12 |
| 2026 | SIT-KGED: Simply Inject Topology into LLM for Knowledge Graph Error Detection
Xingyi Mao, Yipeng Yu |
WWW | 3 |
| 2026 | Beyond content: Multimodal emotional responses predict online moral contagion across laboratory and real-world contexts
Xu Liu 0039, Yipeng Yu, Waxun Su, Yueqin Hu |
Int. J. Hum. Comput. Stud. | 3 |
| 2025 | Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image GenerationabstractPersonalized text-to-image generation methods can generate customized images based on the reference images, which have garnered wide research interest. Recent methods propose a finetuning-free approach with a decoupled cross-attention mechanism to generate personalized images requiring no test-time finetuning. However, when multiple reference images are provided, the current decoupled cross-attention mechanism encounters the object confusion problem and fails to map each reference image to its corresponding object, thereby seriously limiting its scope of application. To address the object confusion problem, in this work we investigate the relevance of different positions of the latent image features to the target object in diffusion model, and accordingly propose a weighted-merge method to merge multiple reference image features into the corresponding objects. Next, we integrate this weighted-merge method into existing pre-trained models and continue to train the model on a multi-object dataset constructed from the open-sourced SA-1B dataset. To mitigate object confusion and reduce training costs, we propose an object quality score to estimate the image quality for the selection of high-quality training samples. Furthermore, our weighted-merge training framework can be employed on single-object generation when a single object has multiple reference images. The experiments verify that our method achieves superior performance to the state-of-the-arts on the Concept101 dataset and DreamBooth dataset of multi-object personalized image generation, and remarkably improves the performance on single-object personalized image generation. Qihan Huang, Siming Fu, Hao Jiang 0014, Yipeng Yu, Jie Song 0011 |
AAAI | 5 |
| 2025 | MHSNet: An MoE-based Hierarchical Semantic Representation Network for Accurate Duplicate Resume Detection with Large Language ModelabstractTo maintain the company's talent pool, recruiters need to continuously search for resumes from third-party websites (e.g., LinkedIn, Indeed). However, fetched resumes are often incomplete and inaccurate. To improve the quality of third-party resumes and enrich the company's talent pool, it is essential to conduct duplication detection between the fetched resumes and those already in the company's talent pool. Such duplication detection is challenging due to the semantic complexity, structural heterogeneity, and information incompleteness of resume texts. To this end, we propose MHSNet, an multi-level identity verification framework that fine-tunes BGE-M3 using contrastive learning. With the fine-tuned BGE-M3, MHSNet generates multi-level sparse and dense representations for resumes, enabling the computation of corresponding multi-level semantic similarities. Moreover, the state-aware Mixture-of-Experts (MoE) is employed in MHSNet to handle diverse incomplete resumes. Experimental results verify the effectiveness of MHSNet. Yu Li 0015, Zulong Chen, Wenjian Xu, Hong Wen 0002, Yipeng Yu, Man Lung Yiu, Yuyu Yin |
CIKM | 5 |
| 2025 | SEAL: Structure and Element Aware Learning Improves Long Structured Document RetrievalabstractIn long structured document retrieval, existing methods typically fine-tune pre-trained language models (PLMs) using contrastive learning on datasets lacking explicit structural information.This practice suffers from two critical issues: 1) current methods fail to leverage structural features and element-level semantics effectively, and 2) the lack of datasets containing structural metadata.To bridge these gaps, we propose SEAL, a novel contrastive learning framework.It leverages structure-aware learning to preserve semantic hierarchies and masked element alignment for fine-grained semantic discrimination.Furthermore, we release StructDocRetrieval, a long structured document retrieval dataset with rich structural annotations.Extensive experiments on both released and industrial datasets across various modern PLMs, along with online A/B testing, demonstrate consistent performance improvements, boosting NDCG@10 from 79.41% to 82.59% on BGE-M3.The resources are available at this URL. Xinhao Huang, Zhibo Ren, Yipeng Yu, Zulong Chen, Zeyi Wen |
EMNLP | 3 |
| 2023 | VideoMaster: A Multimodal Micro Game Video RecreatorabstractTo free human from laborious video production, this paper proposes the building of VideoMaster, a multimodal system equipped with four capabilities: highlight extraction, video describing, video dubbing and video editing. It extracts interesting episodes from long game videos, generates subtitles for each episode, reads the subtitles through synthesized speech, and finally re-creates a better short video through video editing. Notably, VideoMaster takes a combination of deep learning and traditional computer vision techniques to extract highlights with fine-to-coarse labels, utilizes a novel framework named PCSG-v (probabilistic context sensitive grammar for video) for video description generation, and imitates a target speaker's voice to read the description. To the best of our knowledge, VideoMaster is the first multimedia system that can automatically produce product-level micro-videos without heavy human annotation. Yipeng Yu, Hui Zhan |
IJCAI | 1 |
| 2022 | Relational Graph Reasoning Transformer for Image CaptioningabstractThe current published methods of image captioning are directly inputting the features of objects in image into model, and introduced a variety of attention mechanisms to capture the associations between the objects and specific words. But the relationships of vision and semantic between objects are not sufficiently concerned. In this paper, we propose a relational graph reasoning Transformer which explicitly incorporates the relationships of vision and semantic between objects to construct an object relational graph in Transformer. Specifically, besides the detected object features, the global spatial relationships and the semantic context between different objects is attended. Meanwhile, a graph structures feature which correlates object features, their spatial and semantic information is reasoned by a learned grafting mechanism. Finally, the contextual graph feature is integrated into the proposed Transformer decoder. Experimental results demonstrate the significance of our relationship reasoning Transformer model. Xinyu Xiao, Zixun Sun, Tingtian Li, Yipeng Yu |
ICME | 4 |
| 2022 | Learning high-order structural and attribute information by knowledge graph attention networks for enhancing knowledge graph embedding
Hongyun Cai 0001, Sifa Xie, Yipeng Yu, dukehyzhang |
Knowl. Based Syst. | 5 |
| 2021 | Hierarchical Multilabel Text Classification via Multitask LearningabstractHierarchical multilabel classification is a variant of classification where instances might belong to multiple labels and these labels come from a hierarchy. In this paper, we solve the hierarchical multilabel text classification problem of professionally-generated content via multitask learning. More specifically, we focus on (1) how to build models that can share features well in multitask learning, (2) how to incorporate the label dependence into the training procedure of the models, and (3) how to combine the predicted labels of different levels in the hierarchy. To make the experiments simple and comparable, we bring in the state-of-art BERT model as the base model in our work. Experiment results show that the multitask models we build are competitive, the penalty loss we propose is able to improve the performance, and the union operation is the best choice to handle prediction contradiction. In other words, the time cost is reduced but performance is improved via our multitask learning approach. Yipeng Yu, Zixun Sun, Chi Sun |
ICTAI | 1 |
| 2021 | Text2Video: Automatic Video Generation Based on Text ScriptsabstractTo make video creation simpler, in this paper we present Text2Video, a novel system to automatically produce videos using only text-editing for novice users. Given an input text script, the director-like system can generate game-related engaging videos which illustrate the given narrative, provide diverse multi-modal content, and follow video editing guidelines. The system involves five modules: (1) A material manager extracts highlights from raw live game videos, and tags each video highlight, image and audio with labels. (2) A natural language processor extracts entities and semantics from the input text scripts. (3) A refined cross-modal retrieval searches for matching candidate shots from the material manager. (4) A text to speech speaker reads the processed text scripts with synthesized human voice. (5) The selected material shots and synthesized speech are assembled artistically through appropriate video editing techniques. Yipeng Yu, Zirui Tu, Longyu Lu, Hui Zhan, Zixun Sun |
ACM Multimedia | 1 |
| 2021 | Deep Music Retrieval for Fine-Grained Videos by Exploiting Cross-Modal-Encoded Voice-OversabstractRecently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. However, existing video-music retrieval methods only based on the visual modality cannot show promising performance regarding videos with fine-grained virtual contents. In this paper, we also investigate the widely added voice-overs in short videos and propose a novel framework to retrieve BGM for fine-grained short videos. In our framework, we use the self-attention (SA) and the cross-modal attention (CMA) modules to explore the intra- and the inter-relationships of different modalities respectively. For balancing the modalities, we dynamically assign different weights to the modal features via a fusion gate. For paring the query and the BGM embeddings, we introduce a triplet pseudo-label loss to constrain the semantics of the modal embeddings. As there are no existing virtual-content video-BGM retrieval datasets, we build and release two virtual-content video datasets HoK400 and CFM400. Experimental results show that our method achieves superior performance and outperforms other state-of-the-art methods with large margins. Tingtian Li, Zixun Sun, Haoruo Zhang, Ziming Wu, Hui Zhan, Yipeng Yu, Hengcan Shi |
SIGIR | 7 |
| 2020 | When and Who? Conversation Transition Based on Bot-Agent Symbiosis Learning NetworkabstractIn online customer service applications, multiple chatbots that are specialized in various topics are typically developed separately and are then merged with other human agents to a single platform, presenting to the users with a unified interface.Ideally the conversation can be transparently transferred between different sources of customer support so that domain-specific questions can be answered timely and this is what we coined as a Bot-Agent symbiosis.Conversation transition is a major challenge in such online customer service and our work formalises the challenge as two core problems, namely, when to transfer and which bot or agent to transfer to and introduces a deep neural networks based approach that addresses these problems.Inspired by the net promoter score (NPS), our research reveals how the problems can be effectively solved by providing user feedback and developing deep neural networks that predict the conversation category distribution and the NPS of the dialogues.Experiments on realistic data generated from an online service support platform demonstrate that the proposed approach outperforms state-of-the-art methods and shows promising perspective for transparent conversation transition. Yipeng Yu, Ran Guan, Zhuoxuan Jiang, Jingchang Huang |
COLING | 1 |
| 2019 | A General Planning-Based Framework for Goal-Driven Conversation AssistantabstractWe propose a general framework for goal-driven conversation assistant based on Planning methods. It aims to rapidly build a dialogue agent with less handcrafting and make the more interpretable and efficient dialogue management in various scenarios. By employing the Planning method, dialogue actions can be efficiently defined and reusable, and the transition of the dialogue are managed by a Planner. The proposed framework consists of a pipeline of Natural Language Understanding (intent labeler), Planning of Actions (with a World Model), and Natural Language Generation (learned by an attention-based neural network). We demonstrate our approach by creating conversational agents for several independent domains. Zhuoxuan Jiang, GuangYuan Yu, Yipeng Yu, Shaochun Li |
AAAI | 5 |
| 2019 | A Crowdsource-Based Sensing System for Monitoring Fine-Grained Air Quality in Urban EnvironmentsabstractNowadays more and more urban residents are aware of the importance of the air quality to their health, especially who are living in the large cities that are seriously threatened by air pollution. Meanwhile, being limited by the spare sense nodes, the air quality information is very coarse in resolution, which brings urgent demands for high-resolution air quality data acquisition. In this paper, we refer the real-time and fine-gained air quality data in city-scale by employing the crowdsource automobiles as well as their built-in sensors, which significantly improves the sensing system's feasibility and practicability. The main idea of this paper is motivated by that the air component concentration within a vehicle is very similar to that of its nearby environment when the vehicle's windows are open, given the fact that the air will exchange between the inside and outside of the vehicle though the opening window. Therefore, this paper first develops an intelligent algorithm to detect vehicular air exchange state, then extracts the concentration of pollutant in the condition that the concentration trend is convergent after opening the windows, finally, the sensed convergent value is denoted as the equivalent air quality level of the surrounding environment. Based on our Internet of Things cloud platform, real-time air quality data streams from all over the city are collected and analyzed in our data center, and then a fine-gained city level air quality map can be exhibited elaborately. In order to demonstrate the effective- ness of the proposed method, experiments crowdsourcing 500 floating vehicles are conducted in Beijing city for three months to ubiquitously sample the air quality data. Evaluations of the algorithm's performance in comparison with the ground truth indicate the proposed system is practical for collecting air quality data in urban environments. Jingchang Huang, Ning Duan, Chunyang Ma, Yuanyuan Ding, Yipeng Yu, Qianwei Zhou |
IEEE Internet Things J. | 7 |
| 2012 | FlyingBuddy2: a brain-controlled assistant for the handicappedabstractThe motor impaired people have much limit in moving. The devices augmenting their mobility will be much helpful for improving their living experiences. This poster develops a brain-controlled assistive system, called FlyingBuddy2, to aid the handicapped in mobility. It uses the brain EEG signals to directly control a quadrotor. Signals from an EEG headset are transmitted wirelessly to a computer, then the decoded brain signals are converted to trigger the quadrotor to move in 3D space. Three applications are developed: thinking to play games, thinking to see, and thinking to take pictures. Yipeng Yu, Weidong Hua, Shijian Li, Yueming Wang 0001, Gang Pan 0001 |
UbiComp | 1 |