Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shenghui Xu

dblp:160/8892 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0004-6495-2825ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 87% Recommender systems · 13%
Computer networks
1 paper
Network measurement and analytics · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Interconnection networks and networks-on-chip · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
e-commerce search
1.012026
Demand-Calibrated Facet Diversification with Deployable Soft Quotas for E-Commerce Search · SIGIR 2026
Information retrieval
search result diversification
1.012026
Demand-Calibrated Facet Diversification with Deployable Soft Quotas for E-Commerce Search · SIGIR 2026
Network measurement and analytics
per-flow measurement
0.612022
Short-Term Memory Sampling for Spread Measurement in High-Speed Networks · INFOCOM 2022
Network measurement and analytics › traffic measurement
spread estimation
0.612022
Short-Term Memory Sampling for Spread Measurement in High-Speed Networks · INFOCOM 2022
Network measurement and analytics
traffic measurement
0.612022
Short-Term Memory Sampling for Spread Measurement in High-Speed Networks · INFOCOM 2022
Interconnection networks and networks-on-chip
high-speed networks
0.212022
Short-Term Memory Sampling for Spread Measurement in High-Speed Networks · INFOCOM 2022

Methods — techniques the papers use, named apart from their topics

short-term memory sampling · 1.1duplicate filtering · 1.1soft-quota score adjustment · 1.0cohort-level distribution estimation · 1.0
YearPublicationVenuePosition
2026 Demand-Calibrated Facet Diversification with Deployable Soft Quotas for E-Commerce Search
abstract
E-commerce search rankers optimized for engagement or conversion often homogenize first-page results, under-serving ambiguous or exploratory queries. Prior diversification methods typically assume per-query intent/aspect supervision or heavy re-ranking, which is costly and brittle under production constraints. We proposed a supervision-free, deployable framework for demand-calibrated facet diversification. From independent holdout traffic, we estimate a cohort-level target facet distribution using session-attributed downstream engagement, capturing how much diversity users actually demand. At serving time, we steer the first-page slate toward this target with a lightweight soft-quota score adjustment over a short candidate prefix, preserving the base ranker's within-facet ordering and adding negligible latency. Large-scale online experiments on high-volume traffic show consistent improvements in revenue and engagement while reducing abandonment. We further find that delivered first-page facet mixes move closer to the engagement-derived target, indicating effectiveness of the proposed mechanism.
Shenghui Xu
SIGIR1
2022 Cross-Scale Context Extracted Hashing for Fine-Grained Image Binary Encoding
Xuetong Xue, Jiaying Shi, Xinxue He, Shenghui Xu, Zhaoming Pan
ACML4
2022 Short-Term Memory Sampling for Spread Measurement in High-Speed Networks
abstract
Per-flow spread measurement in high-speed networks can provide indispensable information to many practical applications. However, it is challenging to measure millions of flows at line speed because on-chip memory modules cannot simultaneously provide large capacity and large bandwidth. The prior studies address this mismatch by entirely using on-chip compact data structures or utilizing off-chip space to assist limited on-chip memory. Nevertheless, their on-chip data structures record massive transient elements, each of which only appears in a short time interval in a long-period measurement task, and thus waste significant on-chip space. This paper presents short-term memory sampling, a novel spread estimator that samples new elements while only holding elements for short periods. Our estimator can work with tiny on-chip space and provide accurate estimations for online queries. The key of our design is a short-term memory duplicate filter that reports new elements and filters duplicates effectively while allowing incoming elements to override the stale elements to reduce on-chip memory usage. We implement our approach on a NetFPGA-equipped prototype. Experimental results based on real Internet traces show that, compared to the state-of-the-art, short-term memory sampling reduces up to 99% of on-chip memory usage when providing the same probabilistic assurance on spread-estimation error.
Yang Du 0006, He Huang 0001, Yu-e Sun, Shigang Chen, Guoju Gao, Xiaoyu Wang 0004, Shenghui Xu
INFOCOM7
2021 An Efficient Adaptive Noise Correction Framework for Size Measurement over Data Streams
abstract
With the rapid development of the Internet of Things (IoT), massive high-speed data streams are produced every moment, making accurate size estimation a challenging task. Many sketches have been proposed to summarize real-time high-speed data streams and provide per-flow size estimations. However, sketches have to share the memory units to fit in limited on-chip space, inevitably introducing noises to all flows and resulting in over-estimation problems. Prior work adopts an average denoising strategy to remove the same noise from raw sketch estimations. However, they overlook that the noise distribution is highly skewed, leading to inaccurate results for most flows. This paper proposes an efficient Adaptive Noise Correction (ANC) framework, which analyzes the noise of each flow on a case-by-case basis and provides accurate size estimations. The key of our design is to build an ML model to predict a weight coefficient that indicates the noises in raw estimations, which is conducted for each flow by analyzing the neighbor flows whose memory units overlap with the given flow. Then we introduce a novel Probabilistic Cold Filter to block the tiny flows and assist in noise correction. Experimental results based on real Internet traces show that our framework can effectively remove the noises for different sketches, showing better estimation accuracy than the state-of-the-art.
Shenghui Xu, He Huang 0001, Yu-e Sun, Yang Du 0006, Guoju Gao, Xiaoyu Wang 0004, Shiping Chen 0002
ICPADS1
2021 Online Anomalous Taxi Trajectory Detection Based on Multidimensional Criteria
abstract
Online anomalous taxi trajectory detection, which identifies anomalies from ongoing taxi trajectories, has become an important and fundamental concern in many real-world applications. Most of the existing studies define the anomalous trajectories as the ones deviating from the majority of routes or showing abnormal driving time and distance at the same time. However, due to the complexity of road conditions and the variety of passenger preferences, those methods have large false-positive rates, i.e., reporting many normal routes as anomalies. A high false-positive rate is harmful since false alarms will 1) bring unnecessary panic to passengers, 2) cause fiscally punishments to normal drivers, and 3) wastes human resource to deal with drivers' complaints. To this end, this paper proposes an online anomalous trajectory detection method, namely multidimensional criteria based anomalous trajectory (MCAT), to identify anomalous trajectories online. It judges anomalies by considering multidimensional criteria (similarity, time, distance) at the same time, reducing the false positives without sacrificing false negative rates. We evaluate the proposed method based on the real-world taxi data collected from Shanghai, China. The experimental results demonstrate that our method can outperform state-of-the-art methods in terms of accuracy, false-negative rate, and false-positive rate.
Yang Du 0006, Shenghui Xu, Yu-e Sun, He Huang 0001, Guoju Gao
IJCNN3