VLDB 2026 Research / reviewers in the wild / expert
Junfeng Kang
dblp:226/9833
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TMT: A Tri-Modal Transformer for Non-histone Lysine Acetylation Site Prediction
Shuang Cheng, Junfeng Kang, Yuehui Chen |
ICIC (16) | 4 |
| 2025 | Distribution-Driven Dense Retrieval: Modeling Many-to-One Query-Document RelationshipabstractDense retrieval has emerged as the leading approach in information retrieval, aiming to find semantically relevant documents based on natural language queries. Given that a single document can be retrieved by multiple distinct queries, existing methods aim to represent a document with multiple vectors. Each vector is aligned with a different query to model the many-to-one relationship between queries and documents. However, these multiple vector-based approaches encounter challenges such as Increased Storage, Vector Collapse, and Search Efficiency. To address these issues, we introduce the Distribution-Driven Dense Retrieval framework (DDR). Specifically, we use vectors to represent queries and distributions to represent documents. This approach not only captures the relationships between multiple queries corresponding to the same document but also avoids the need to use multiple vectors to represent the document. Furthermore, to ensure search efficiency for DDR, we propose a dot product-based computation method to calculate the similarity between documents represented by distributions and queries represented by vectors. This allows for seamless integration with existing approximate nearest neighbor (ANN) search algorithms for efficient search. Finally, we conduct extensive experiments on real-world datasets, which demonstrate that our method significantly outperforms traditional dense retrieval methods. Junfeng Kang, Rui Li 0093, Qi Liu 0003, Zhenya Huang, Zheng Zhang 0048, Yanjiang Chen, Linbo Zhu, Yu Su 0002 |
AAAI | 1 |
| 2025 | PQR: Improving Dense Retrieval via Potential Query ModelingabstractDense retrieval has now become the mainstream paradigm in information retrieval. The core idea of dense retrieval is to align document embeddings with their corresponding query embeddings by maximizing their dot product. The current training data is quite sparse, with each document typically associated with only one or a few labeled queries. However, a single document can be retrieved by multiple different queries. Aligning a document with just one or a limited number of labeled queries results in a loss of its semantic information. In this paper, we propose a training-free Potential Query Retrieval (PQR) framework to address this issue. Specifically, we use a Gaussian mixture distribution to model all potential queries for a document, aiming to capture its comprehensive semantic information. To obtain this distribution, we introduce three sampling strategies to sample a large number of potential queries for each document and encode them into a semantic space. Using these sampled queries, we employ the Expectation-Maximization algorithm to estimate parameters of the distribution. Finally, we also propose a method to calculate similarity scores between user queries and documents under the PQR framework. Extensive experiments demonstrate the effectiveness of the proposed method. Junfeng Kang, Rui Li 0093, Qi Liu 0003, Yanjiang Chen, Zheng Zhang 0048, Junzhe Jiang 0001, Yu Su 0002 |
ACL (1) | 1 |
| 2025 | MGS3: A Multi-Granularity Self-Supervised Code Search FrameworkabstractIn the pursuit of enhancing software reusability and developer productivity, code search has emerged as a key area, aimed at retrieving code snippets relevant to functionalities based on natural language queries. Despite significant progress in self-supervised code pre-training utilizing the vast amount of code data in repositories, existing methods have primarily focused on leveraging contrastive learning to align natural language with function-level code snippets. These studies have overlooked the abundance of fine-grained (such as block-level and statement-level) code snippets prevalent within the function-level code snippets, which results in suboptimal performance across all levels of granularity. To address this problem, we first construct a multi-granularity code search dataset called MGCodeSearchNet, which contains 536K+ pairs of natural language and code snippets. Subsequently, we introduce a novel Multi-Granularity Self-Supervised contrastive learning code Search framework (MGS3). First, MGS3 features a Hierarchical Multi-Granularity Representation module (HMGR), which leverages syntactic structural relationships for hierarchical representation and aggregates fine-grained information into coarser-grained representations. Then, during the contrastive learning phase, we endeavor to construct positive samples of the same granularity for fine-grained code, and introduce in-function negative samples for fine-grained code. Finally, we conduct extensive experiments on code search benchmarks across various granularities, demonstrating that the framework exhibits outstanding performance in code search tasks of multiple granularities. These experiments also showcase its model-agnostic nature and compatibility with existing pre-trained code representation models. Rui Li 0093, Junfeng Kang, Qi Liu 0003, Liyang He, Zheng Zhang 0048, Yunhao Sha, Linbo Zhu, Zhenya Huang |
KDD (1) | 2 |
| 2025 | GEAR: Generalized Alternating Regressor for Multi-Behavior Sequential RecommendationabstractModern recommender systems face a critical challenge in modeling the intricate interplay between multi-behavior interactions of users (e.g., clicks, adds-to-cart and purchases) and temporal dynamics that drive evolving preferences. While existing multi-behavior sequential recommendation methods attempt to capture these signals, they often suffer from fragmented modeling, such as decoupling behaviors and items into separate sequences, neglecting time-aware transitions, or relying on computationally intensive architectures that hinder real-world scalability. To address these limitations, we propose GEneralized Alternating Regressor (GEAR), a novel framework that unifies behaviors, items, and temporal contexts into a single autoregressive sequence through an alternating architecture. At its core, GEAR represents user interactions as triplets and processes them through a modular transformer architecture. In this architecture, each triplet is alternately modeled at lower layers to disentangle fine-grained patterns, while upper layers jointly learn cross-signal dependencies. This design mimics the interlocking mechanism of gears, enabling the seamless transitions between multi-behavior dynamics and item transitions. Additionally, we incorporate a time-bias term to quantify the decay of behavioral influence across both short- and long-term horizons. Extensive experiments on real-world datasets validate the effectiveness, generalizability, and computational efficiency of the proposed framework. Junzhe Jiang 0001, Kai Zhang 0038, Junfeng Kang, Yucong Luo, Min Gao 0017 |
SIGIR | 3 |
| 2022 | Integration of Internet search data to predict tourism trends using spatial-temporal XGBoost composite modelabstractTourism trend prediction facilitates estimation of tourism investment and revenue. Studies on tourism prediction have primarily relied on linear models and historical visitors; however, relationships between tourism trends and their factors may be nonlinear. This study constructed factors from internet search data and predicted tourism trends using a spatiotemporal framework based on the extreme gradient boosting (XGBoost) method. The study first sorted Baidu index data that is computed by weighting the search frequency. The spatial cluster analysis was conducted to incorporate spatial characteristics, and principal component analysis was further performed to identify factors. The next step derived variables using the weighted moving average method to reduce the lag effect between tourism internet search and actual behavior. We applied the proposed spatiotemporal XGBoost composite model to predict Beijing’s tourism trends. The R2 scores of the simple XGBoost model, the autoregressive integrated moving average model, the spatial XGBoost model, and the spatiotemporal XGBoost composite model were 0.517, 0.625, 0.791, and 0.940, respectively. Compared to predictions from different models, the spatiotemporal XGBoost composite model has the best prediction ability. The findings also suggest that machine learning methods may not perform well without considering spatial properties, such as spatial autocorrelation and spatial heterogeneity. Junfeng Kang, Xingyu Guo, Zhengqiu Fan |
Int. J. Geogr. Inf. Sci. | 1 |