Jingwei Wu

dblp:62/3245 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author
YearPublicationVenuePosition
2026 Policy-Grounded Dynamic Facet Suggestions for Job Search
abstract
Job seekers often initiate search with short, underspecified queries. At LinkedIn, over 80% of job-related queries contain three or fewer keywords, making accurate user intent inference and relevant job retrieval particularly challenging. We present dynamic facet suggestion (DFS), an interactive query-refinement mechanism that facilitates intent disambiguation by surfacing personalized semantic attributes conditioned on the joint user-query context in real time. We propose a policy-grounded, retrieval-augmented ranking framework for facet suggestion, comprising offline taxonomy curation, embedding-based retrieval of top-K candidates, and a distilled small language model (SLM) based candidate scoring. The system is optimized for real-time serving via point-wise single-token scoring and batching/prefix caching. Offline evaluation demonstrates high precision for generated suggestions, and online A/B tests show significant lifts in suggestion engagement and job search outcomes.
Baofen Zheng, Qianqi Shen, Jianqiang Shen, Wenqiong Liu, Chunnan Yao, Ping Liu 0002, Rajat Arora 0002, Kevin Kao, Hsiang Lin, Wanjun Jiang, Yusuke Takebuchi, Jingwei Wu
SIGIR13
2026 Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
Baofen Zheng, Jianqiang Shen, Benjamin Le, Wen Pu, Neha Saraf, Alice Leung, Qianqi Shen, Liangjie Hong, Jingwei Wu
SIGIR13
2025 Powering Job Search at Scale: LLM-Enhanced Query Understanding in Job Matching Systems
abstract
Query understanding is essential in modern relevance systems, where user queries are often short, ambiguous, and highly context-dependent. Traditional approaches often rely on multiple task-specific Named Entity Recognition models to extract structured facets as seen in job search applications. However, this fragmented architecture is brittle, expensive to maintain, and slow to adapt to evolving taxonomies and language patterns. In this paper, we introduce a unified query understanding framework powered by a Large Language Model (LLM), designed to address these limitations. Our approach jointly models the user query and contextual signals such as profile attributes to generate structured interpretations that drive more accurate and personalized recommendations. The framework improves relevance quality in online A/B testing while significantly reducing system complexity and operational overhead. The results demonstrate that our solution provides a scalable and adaptable foundation for query understanding in dynamic web applications.
Ping Liu 0002, Jianqiang Shen, Qianqi Shen, Chunnan Yao, Kevin Kao, Rajat Arora 0002, Baofen Zheng, Caleb Johnson, Liangjie Hong, Jingwei Wu
CIKM11
2025 Advancing video self-supervised learning via image foundation models
Jingwei Wu, Zhewei Huang, Chang Liu 0047
Pattern Recognit. Lett.1
2009 Oracle Streams: A High Performance Implementation for Near Real Time Asynchronous Replication
abstract
We present the architectural design and recent performance optimizations of a state of the art commercial database replication technology provided in Oracle Streams. The underlying design of streams replication is a pipeline of components that are responsible for capturing, propagating, and applying logical change records (LCRs) from a source database to a destination database. Each LCR encapsulates a database change. The communication in this pipeline is now latch-free to increase the throughput of LCRs. In addition, the apply component now bypasses SQL whenever possible and uses a new latch-free metadata cache. We outline the algorithms behind these optimizations and quantify the replication performance improvement from each optimization. Finally, we demonstrate that these optimizations improve the replication performance by more than a factor of four and achieve replication throughput of over 20,000 LCRs per second with sub-second latency on commodity hardware.
Lik Wong, Nimar S. Arora, Thuvan Hoang, Jingwei Wu
ICDE5
2007 Empirical Evidence for SOC Dynamics in Software Evolution
abstract
We examine eleven large open source software systems and present empirical evidence for the existence of fractal structures in software evolution. In our study, fractal structures are measured as power laws throughout the lifetime of each software system. We describe two specific power law related phenomena: the probability distribution of software changes decreases as a power function of change sizes; and the time series of software change exhibits long range correlations with power law behavior. The existence of such spatial (across the system) and temporal (over the system lifetime) power laws suggests that self-organized criticality (SOC) occurs in the evolution of open source systems. As a result, SOC may be useful as a conceptual framework for understanding software evolution dynamics (the cause and mechanism of change or growth). We also discuss the implications of SOC to software practices.
Jingwei Wu, Richard C. Holt, Ahmed E. Hassan
ICSM1
2005 Comparison of Clustering Algorithms in the Context of Software Evolution
abstract
To aid software analysis and maintenance tasks, a number of software clustering algorithms have been proposed to automatically partition a software system into meaningful subsystems or clusters. However, it is unknown whether these algorithms produce similar meaningful clusterings for similar versions of a real-life software system under continual change and growth. This paper describes a comparative study of six software clustering algorithms. We applied each of the algorithms to subsequent versions from five large open source systems. We conducted comparisons based on three criteria respectively: stability (Does the clustering change only modestly as the system undergoes modest updating?), authoritative-ness (Does the clustering reasonably approximate the structure an authority provides?) and extremity of cluster distribution (Does the clustering avoid huge clusters and many very small clusters?). Experimental results indicate that the studied algorithms exhibit distinct characteristics. For example, the clusterings from the most stable algorithm bear little similarity to the implemented system structure, while the clusterings from the least stable algorithm has the best cluster distribution. Based on obtained results, we claim that current automatic clustering algorithms need significant improvement to provide continual support for large software projects.
Jingwei Wu, Ahmed E. Hassan, Richard C. Holt
ICSM1