Jason Zhu

dblp:37/6928 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
2since 2021 · last 2025
0009-0004-3416-0798ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Multi-agent systems · 46% Question answering and dialogue systems · 46% Graph learning · 8%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
conversational search
0.912025
Insight Agents: An LLM-Based Multi-Agent System for Data Insights · SIGIR 2025
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems
0.912025
Insight Agents: An LLM-Based Multi-Agent System for Data Insights · SIGIR 2025
Information retrieval › online advertising
sponsored search
0.512021
TextGNN: Improving Text Encoder via Graph Neural Network in Sponsored Search · WWW 2021
Machine learning › Graph learning
graph neural network
0.112021
TextGNN: Improving Text Encoder via Graph Neural Network in Sponsored Search · WWW 2021

Methods — techniques the papers use, named apart from their topics

out-of-domain detection · 1.7large language model · 1.7agent routing · 1.7BERT-based classifier · 1.7twin tower encoder · 1.0graph neural network · 1.0
YearPublicationVenuePosition
2025 Insight Agents: An LLM-Based Multi-Agent System for Data Insights
abstract
Today, E-commerce sellers face several key challenges, including difficulties in discovering and effectively utilizing available programs and tools, and struggling to understand and utilize rich data from various tools. We therefore aim to develop Insight Agents (IA), a conversational multi-agent Data Insight system, to provide E-commerce sellers with personalized data and business insights through automated information retrieval. Our hypothesis is that IA will serve as a force multiplier for sellers, thereby driving incremental seller adoption by reducing the effort required and increase speed at which sellers make good business decisions. In this paper, we introduce this new LLM-backed end-to-end agentic workflow designed for comprehensive coverage, high accuracy, and low latency. It features a hierarchical multi-agent structure, consisting of manager agent and two worker agents: data presentation and insight generation, for efficient information retrieval and problem-solving. We design a simple yet effective ML solution for manager agent that combines Out-of-Domain (OOD) detection using a lightweight encoder-decoder model and agent routing through a BERT-based classifier, optimizing both accuracy and latency. Within the two worker agents, a strategic planning is designed for API-based data model that breaks down queries into granular components to generate more accurate responses, and domain knowledge is dynamically injected to to enhance the insight generator. IA has been launched for Amazon sellers in US, which has achieved high accuracy of 89.5% based on human evaluation, with latency of P90 below 15s.
Jincheng Bai, Zhenyu Zhang 0021, Jennifer Zhang, Jason Zhu
SIGIR4
2021 TextGNN: Improving Text Encoder via Graph Neural Network in Sponsored Search
abstract
Text encoders based on C-DSSM or transformers have demonstrated strong performance in many Natural Language Processing (NLP) tasks. Low latency variants of these models have also been developed in recent years in order to apply them in the field of sponsored search which has strict computational constraints. However these models are not the panacea to solve all the Natural Language Understanding (NLU) challenges as the pure semantic information in the data is not sufficient to fully identify the user intents. We propose the TextGNN model that naturally extends the strong twin tower structured encoders with the complementary graph information from user historical behaviors, which serves as a natural guide to help us better understand the intents and hence generate better language representations. The model inherits all the benefits of twin tower models such as C-DSSM and TwinBERT so that it can still be used in the low latency environment while achieving a significant performance gain than the strong encoder-only counterpart baseline models in both offline evaluations and online production system. In offline experiments, the model achieves a 0.14% overall increase in ROC-AUC with a 1% increased accuracy for long-tail low-frequency Ads, and in the online A/B testing, the model shows a 2.03% increase in Revenue Per Mille with a 2.32% decrease in Ad defect rate.
Jason Zhu, Yanling Cui, Hao Sun 0015, Markus Pelger, Liangjie Zhang, Ruofei Zhang, Huasha Zhao
WWW1
1997 Image-based keyword recognition in oriental language document images
Jason Zhu, Tao Hong 0001, Jonathan J. Hull
Pattern Recognit.1
1994 Image-based word recognition in oriental language document images
abstract
An algorithm for word recognition in oriental languages such as Chinese, Japanese, and Korean is presented. The objective is to recognize words, that are composed of a number of consecutive characters, in document images where there are no explicit visually defined word boundaries. The technique exploits the redundancy in these languages that is expressed by the difference between the number of possible character strings of a fixed length and the number of legal words of that length. Sequences of character images are matched simultaneously to lists of legal words and illegal strings that are likely to occur. A word is located if its image is more likely to occur in the current context than any of the illegal strings that are visually similar to it. No intermediate character recognition step is used. The application of contextual information directly to the interpretation of features extracted from the image overcomes noise that could have made isolated character recognition impossible and the location of words with conventional postprocessing algorithms difficult. Experimental results are presented that show the ability of this algorithm to correctly recognize text in the presence of noise.
Jason Zhu, Jonathan J. Hull
ICPR (2)1