EDBT 2026 Demo / reviewers in the wild / expert
Yexin Wang
dblp:51/2047
· DBLP profile ↗
18ranked-venue papers
2as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RMP-adapter: A region-based Multiple Prompt Adapter for multi-concept customization in text-to-image diffusion modelabstractThis paper introduces a novel framework for multi-concept customization in text-to-image diffusion models . At its core is a Multiple Prompt Adapter (MP-Adapter) capable of processing multiple image prompts in parallel, extracting features from target concepts and projecting them into the same latent space as the text prompt. This enables simultaneous handling of multiple concepts using just one reference image per concept. To address challenges in fusing multiple concepts with complex interactions, we propose a Region-based Denoising Framework (RDF) that dynamically generates concept-specific regions of interest during inference, allowing spatially decoupled injection of concept features. By integrating the MP-Adapter and RDF, our end-to-end pipeline enables multi-concept customization with intricate occlusions and interactions while preserving concept identities. This approach surpasses current methods by resolving concept conflicts, identity degradation, and occlusion issues, allowing flexible customization without concept-specific retraining. Both qualitative and quantitative evaluations demonstrate that our framework outperforms state-of-the-art approaches in multi-concept customization tasks, while ablation studies validate the effectiveness of each proposed component. This work significantly advances text-to-image generation capabilities for complex, user-defined concept combinations. Code and models will be released at https://github.com/baojudezeze/RMP-Adapter . Lai-Man Po, Xuyuan Xu, Yexin Wang, Haoxuan Wu, Kun Li 0015 |
Expert Syst. Appl. | 4 |
| 2025 | Cross-Site Visual Localization of Zhurong Mars Rover Based on Self-Supervised Keypoint Extraction and Robust MatchingabstractHigh-precision localization of the Mars rovers is fundamental for path planning and safe navigation toward exploration targets during Mars missions. In cross-site visual localization, image matching is the key step to obtain corresponding points connecting images from different sites. The cross-site visual localization method based on Affine SIFT (ASIFT) is used in Tianwen-1 mission but is constrained in regions of Mars with poor texture and large viewpoint invariance. In this article, we propose a cross-site visual localization methodology of Mars rover based on self-supervised keypoint extraction and robust matching. The self-supervised keypoint extraction network, which is called MRSS-Net, uses multiscale deformable structures (MSDSs) during the feature encoding stage to enhance the network’s ability of extracting invariant features in regions with large viewpoint variations and improve the rate of identical points for cross-site images with poor texture. In addition, we develop self-attention descriptor enhancement mechanism (SADEM) to distinguish local features in repetitive patterns. The robust matching, which is called adaptive 2-D–3-D matching, uses GNC dead-reckoning (3-D priori information) to construct the initial coarse matching domain and homography matrix (2-D information) to construct a progressively shrinking refined matching domain. We compared our method against ASIFT based cross-site visual localization model and advanced deep learning algorithms and evaluate the performance using NaTeCam images collected during the traversal of four long-distance traversals (a total of 44 Martian sol sites) by Zhurong rover. The experimental results show that our framework reduces the localization error by 12.5% and improves localization robustness by 50.8%, compared with ASIFT-based cross-site visual localization method used in Zhurong rover. In addition, our method outperforms state-of-the-art deep learning techniques and ensures the current accuracy of cross-site visual localization for Mars rover, while significantly increasing the level of automation. Yuke Kou, Wenhui Wan, Kaichang Di, Zhaoqin Liu, Man Peng, Yexin Wang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | NC-ALG: Graph-Based Active Learning Under Noisy CrowdabstractGraph Neural Networks (GNNs) have achieved great success in various data mining tasks but they heavily rely on a large number of annotated nodes, requiring considerable human efforts. Despite the effectiveness of existing GNN-based Active Learning (AL) methods, they assume that the annotated labels are always correct, which is contradictory to the error-prone labeling process in a practical crowdsourcing environment. Besides, due to this impractical assumption, existing works only focus on optimizing the node selection in AL but neglect optimizing the labeling process. Therefore, we present NC-ALG, the first GNN-based AL framework that optimizes both the node selection and node labeling process under a noisy crowd. For node selection, NC-ALG introduces a new measurement to model influence reliability and an effective influence maximization objective to select nodes. For node labeling, NC-ALG significantly reduces the labeling cost by considering the model-predicted labels and the labels of mirror nodes. To the best of our knowledge, this is the first attempt to consider GNN-based AL under the practical noisy crowd. Empirical studies on public datasets demonstrate that NC-ALG significantly outperforms existing methods in terms labeling efficiency. Notably, it only takes NC-ALG one-third of the labeling budget that the competitive baseline GRAIN needs to achieve an accuracy of 70.7 % on PubMed. Wentao Zhang 0001, Yexin Wang, Zhenbang You, Yang Li 0106, Gang Cao 0003, Zhi Yang 0001, Bin Cui 0001 |
ICDE | 2 |
| 2024 | SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
Chaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi, Wei-Yi Pei, Yexin Wang, Ying Shan, Wei-Shi Zheng 0001, Jianfang Hu |
ACM Multimedia | 7 |
| 2023 | Darwinian Model Upgrades: Model Evolving with Selective CompatibilityabstractThe traditional model upgrading paradigm for retrieval requires recomputing all gallery embeddings before deploying the new model (dubbed as "backfilling"), which is quite expensive and time-consuming considering billions of instances in industrial applications. BCT presents the first step towards backward-compatible model upgrades to get rid of backfilling. It is workable but leaves the new model in a dilemma between new feature discriminativeness and new-to-old compatibility due to the undifferentiated compatibility constraints. In this work, we propose Darwinian Model Upgrades (DMU), which disentangle the inheritance and variation in the model evolving with selective backward compatibility and forward adaptation, respectively. The old-to-new heritable knowledge is measured by old feature discriminativeness, and the gallery features, especially those of poor quality, are evolved in a lightweight manner to become more adaptive in the new latent space. We demonstrate the superiority of DMU through comprehensive experiments on large-scale landmark retrieval and face recognition benchmarks. DMU effectively alleviates the new-to-new degradation at the same time improving new-to-old compatibility, rendering a more proper model upgrading paradigm in large-scale retrieval systems.Code: https://github.com/TencentARC/OpenCompatible. Binjie Zhang, Shupeng Su, Yixiao Ge, Xuyuan Xu, Yexin Wang, Chun Yuan 0003, Zheng Shou 0001, Ying Shan |
AAAI | 5 |
| 2023 | Binary Embedding-based Retrieval at TencentabstractLarge-scale embedding-based retrieval (EBR) is the cornerstone of search-related industrial applications. Given a user query, the system of EBR aims to identify relevant information from a large corpus of documents that may be tens or hundreds of billions in size. The storage and computation turn out to be expensive and inefficient with massive documents and high concurrent queries, making it difficult to further scale up. Yukang Gan, Yixiao Ge, Chang Zhou 0008, Shupeng Su, Zhouchuan Xu, Xuyuan Xu, Quanchao Hui, Yexin Wang, Ying Shan |
KDD | 9 |
| 2023 | Scapin: Scalable Graph Structure Perturbation by Augmented Influence MaximizationabstractGenerating data perturbations to graphs has become a useful tool for analyzing the robustness of Graph Neural Networks (GNNs). However, existing model-driven methodologies can be prohibitively expensive to apply in large graphs, which hinders the understanding of GNN robustness at scale. In this paper, we present Scapin, a data-driven methodology that opens up a new perspective by connecting graph structure perturbation for GNNs with augmented influence maximization-to either facilitate desirable spreads or curtail undesirable ones by adding or deleting a small set of edges. This connection not only allows us to perform data perturbation on GNNs with computation scalability but also provides nice interpretations. To transform such connections into efficient perturbation approaches for the new GNN setting, Scapin introduces a novel edge influence model, decomposed influence maximization objectives, and a principled algorithm for edge addition by exploiting submodularity of the objectives. Empirical studies demonstrate that Scapin can give orders of magnitude improvement over state-of-art methods in terms of runtime and memory efficiency, with comparable or even better performance. Yexin Wang, Zhi Yang 0001, Wentao Zhang 0001, Bin Cui 0001 |
Proc. ACM Manag. Data | 1 |
| 2022 | Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video RepresentationabstractSpatio-temporal representation learning is critical for video self-supervised representation. Recent approaches mainly use contrastive learning and pretext tasks. However, these approaches learn representation by discriminating sampled instances via feature similarity in the latent space while ignoring the intermediate state of the learned representations, which limits the overall performance. In this work, taking into account the degree of similarity of sampled instances as the intermediate state, we propose a novel pretext task - spatio-temporal overlap rate (STOR) prediction. It stems from the observation that humans are capable of discriminating the overlap rates of videos in space and time. This task encourages the model to discriminate the STOR of two generated samples to learn the representations. Moreover, we employ a joint optimization combining pretext tasks with contrastive learning to further enhance the spatio-temporal representation learning. We also study the mutual influence of each component in the proposed scheme. Extensive experiments demonstrate that our proposed STOR task can favor both contrastive learning and pretext tasks and the joint optimization scheme can significantly improve the spatio-temporal representation in video understanding. The code is available at https://github.com/Katou2/CSTP. Yujia Zhang 0002, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing Yin Yu |
AAAI | 5 |
| 2022 | Aesthetic Text Logo Synthesis via Content-aware Layout InferringabstractText logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, few attention has been paid to this task which needs to take many factors (e.g., fonts, linguistics, topics, etc.) into consideration. In this paper, we propose a content-aware layout generation network which takes glyph images and their corresponding text as input and synthesizes aesthetic layouts for them automatically. Specifically, we develop a dual-discriminator module, including a sequence discriminator and an image discriminator, to evaluate both the character placing trajectories and rendered shapes of synthesized text logos, respectively. Furthermore, we fuse the information of linguistics from texts and visual semantics from glyphs to guide layout prediction, which both play important roles in professional layout design. To train and evaluate our approach, we construct a dataset named as TextLogo3K, consisting of about 3,500 text logo images and their pixel-level annotations. Experimental studies on this dataset demonstrate the effectiveness of our approach for synthesizing visually-pleasing text logos and verify its superiority against the state of the art. Guo Pu, Wenhan Luo, Yexin Wang, Pengfei Xiong, Hongwen Kang, Zhouhui Lian |
CVPR | 4 |
| 2022 | Hot-Refresh Model Upgrades with Regression-Free Compatible Training in Image Retrieval
Binjie Zhang, Yixiao Ge, Yantao Shen 0003, Yu Li 0003, Chun Yuan 0003, Xuyuan Xu, Yexin Wang, Ying Shan |
ICLR | 7 |
| 2022 | Information Gain Propagation: a New Way to Graph Active Learning with Soft Labels
Wentao Zhang 0001, Yexin Wang, Zhenbang You, Jiulong Shan, Zhi Yang 0001, Bin Cui 0001 |
ICLR | 2 |
| 2022 | Towards Universal Backward-Compatible Representation LearningabstractConventional model upgrades for visual search systems require offline refresh of gallery features by feeding gallery images into new models (dubbed as “backfill”), which is time-consuming and expensive, especially in large-scale applications. The task of backward-compatible representation learning is therefore introduced to support backfill-free model upgrades, where the new query features are interoperable with the old gallery features. Despite the success, previous works only investigated a close-set training scenario (i.e., the new training set shares the same classes as the old one), and are limited by more realistic and challenging open-set scenarios. To this end, we first introduce a new problem of universal backward-compatible representation learning, covering all possible data split in model upgrades. We further propose a simple yet effective method, dubbed as Universal Backward-Compatible Training (UniBCT) with a novel structural prototype refinement algorithm, to learn compatible representations in all kinds of model upgrading benchmarks in a unified manner. Comprehensive experiments on the large-scale face recognition datasets MS1Mv3 and IJB-C fully demonstrate the effectiveness of our method. Source code is available at https://github.com/TencentARC/OpenCompatible. Binjie Zhang, Yixiao Ge, Yantao Shen 0003, Shupeng Su, Fanzi Wu, Chun Yuan 0003, Xuyuan Xu, Yexin Wang, Ying Shan |
IJCAI | 8 |
| 2022 | Blind Robust Video Watermarking Based on Adaptive Region Selection and Channel ReferenceabstractDigital watermarking technology has a wide range of applications in video distribution and copyright protection due to its excellent invisibility and convenient traceability. This paper proposes a robust blind watermarking algorithm using adaptive region selection and channel reference. By designing a combinatorial selection algorithm using texture information and feature points, the method realizes automatically selecting stable blocks which can avoid being destroyed during video encoding and complex attacks. In addition, considering human's insensitivity to some specific color components, a channel-referenced watermark embedding method is designed for less impact on video quality. Moreover, compared with other methods' embedding watermark only at low frequencies, our method tends to modify low-frequency coefficients close to mid frequencies, further ensuring stable retention of the watermark information in the video encoding process. Experimental results show that the proposed method achieves excellent video quality and high robustness against geometric attacks, compression, transcoding and camcorder recordings attacks. Qinwei Chang, Leichao Huang, Shaoteng Liu, Hualuo Liu, Yexin Wang |
ACM Multimedia | 6 |
| 2021 | RIM: Reliable Influence-based Active Learning on GraphsabstractMessage passing is the core of most graph models such as Graph Convolutional Network (GCN) and Label Propagation (LP), which usually require a large number of clean labeled data to smooth out the neighborhood over the graph. However, the labeling process can be tedious, costly, and error-prone in practice. In this paper, we propose to unify active learning (AL) and message passing towards minimizing labeling costs, e.g., making use of few and unreliable labels that can be obtained cheaply. We make two contributions towards that end. First, we open up a perspective by drawing a connection between AL enforcing message passing and social influence maximization, ensuring that the selected samples effectively improve the model performance. Second, we propose an extension to the influence model that incorporates an explicit quality factor to model label noise. In this way, we derive a fundamentally new AL selection criterion for GCN and LP--reliable influence maximization (RIM)--by considering quantity and quality of influence simultaneously. Empirical studies on public datasets show that RIM significantly outperforms current AL methods in terms of accuracy and efficiency. Wentao Zhang 0001, Yexin Wang, Zhenbang You, Jiulong Shan, Zhi Yang 0001, Bin Cui 0001 |
NeurIPS | 2 |
| 2021 | Grain: Improving Data Efficiency of Graph Neural Networks via Diversified Influence MaximizationabstractData selection methods, such as active learning and core-set selection, are useful tools for improving the data efficiency of deep learning models on large-scale datasets. However, recent deep learning models have moved forward from independent and identically distributed data to graph-structured data, such as social networks, e-commerce user-item graphs, and knowledge graphs. This evolution has led to the emergence of Graph Neural Networks (GNNs) that go beyond the models existing data selection methods are designed for. Therefore, we present GRAIN, an efficient framework that opens up a new perspective through connecting data selection in GNNs with social influence maximization. By exploiting the common patterns of GNNs, GRAIN introduces a novel feature propagation concept, a diversified influence maximization objective with novel influence and diversity functions, and a greedy algorithm with an approximation guarantee into a unified framework. Empirical studies on public datasets demonstrate that GRAIN significantly improves both the performance and efficiency of data selection (including active learning and core-set selection) for GNNs. To the best of our knowledge, this is the first attempt to bridge two largely parallel threads of research, data selection, and social influence maximization, in the setting of GNNs, paving new ways for improving data efficiency. Wentao Zhang 0001, Zhi Yang 0001, Yexin Wang, Yu Shen 0003, Yang Li 0106, Liang Wang 0001, Bin Cui 0001 |
Proc. VLDB Endow. | 3 |
| 2020 | Landing site topographic mapping and rover localization for Chang'e-4 mission
Zhaoqin Liu, Kaichang Di, Jianfeng Xie, Xiaofeng Cui, Luhua Xi, Wenhui Wan, Man Peng, Bin Liu 0049, Yexin Wang, Sheng Gou, Zongyu Yue, Lichun Li, Jia Wang 0044, Chuankai Liu, Mengna Jia, Zheng Bo, Jia Liu 0047, Runzhi Wang 0002, Shengli Niu, Kuan Zhang 0005, Yi You |
Sci. China Inf. Sci. | 10 |
| 2009 | MagicCube: choosing the best snippet for each aspect of an entityabstractWikis are currently used in business to provide knowledge management systems, especially for individual organizations. However, building wikis manually is a laborious and time-consuming work. To assist founding wikis, we propose a methodology in this paper to automatically select the best snippets for entities as their initial explanations. Our method consists of two steps. First, we focus on extracting snippets from a given set of web pages for each entity. Starting from a seed sentence, a snippet grows up by adding the most relevant neighboring sentences into itself. The sentences are chosen by the Snippet Growth Model, which employs a distance function and an influence function to make decisions. Secondly, we pick out the best snippet for each aspect of an entity. The combination of all the selected snippets serves as the primary description of the entity. We present three ever-increasing methods to handle selection process. Experimental results based on a real data set show that our proposed method works effectively in producing primary descriptions for entities such as employee names. Yexin Wang, Yan Zhang 0004 |
CIKM | 1 |
| 2008 | Weighting Links Using Lexical and Positional Analysis in Web RankingabstractLink analysis has been widely used to evaluate the importance of Web pages. Popular link analysis algorithms are mainly based on the link structure between pages. However, a Web page usually contains various links such as for navigation, decoration or nepotism, which are irrelevant to the topic of the Web page and can not reflect the actual voting relations between pages. In order to improve the performance of Web ranking, we bring out one filtering algorithm to recognize and eliminate these unrelated links using Content Lexical and Positional analysis. Experimental results on different Web domains show that our filtering model can efficiently detect the irrelevant links and effectively help to build a good link graph for the ranking calculation. Yi Zhang 0012, Yexin Wang, Lidong Bing, Yan Zhang 0004 |
WAIM | 2 |