EDBT 2026 Demo / reviewers in the wild / expert
Gensheng Zhang
dblp:137/7229
· DBLP profile ↗
8ranked-venue papers
4as first author
1since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 3 first-authorArtificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 33% Knowledge graphs · 33% Data mining · 31% |
Topics — the 5 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
pattern mining |
0.4 | 2 | 2018 | Maverick: Discovering Exceptional Facts from Knowledge Graphs · SIGMOD Conference 2018 Maverick: A System for Discovering Exceptional Facts from Knowledge Graphs · Proc. VLDB Endow. 2018 |
Knowledge graphs
knowledge graph mining |
0.3 | 1 | 2018 | Maverick: Discovering Exceptional Facts from Knowledge Graphs · SIGMOD Conference 2018 |
Information retrieval › fact-checking
claim detection |
0.3 | 1 | 2017 | ClaimBuster: The First-ever End-to-end Fact-checking System · Proc. VLDB Endow. 2017 |
Information retrieval
fact-checking |
0.3 | 1 | 2017 | ClaimBuster: The First-ever End-to-end Fact-checking System · Proc. VLDB Endow. 2017 |
Information retrieval
text analysis |
0.1 | 1 | 2017 | ClaimBuster: The First-ever End-to-end Fact-checking System · Proc. VLDB Endow. 2017 |
Methods — techniques the papers use, named apart from their topics
beam search · 0.7set enumeration tree · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Uncertainty Calibrations of Deep-Learning Schemes for Full-Wave Inverse Scattering ProblemsabstractRecently, deep learning methods have attracted intensive attentions on solving inverse scattering problems (ISPs). However, different with traditional physical-model based methods, the reliability of the learning methods is low due to the low explainability of deep neural networks (DNNs). Thus, the uncertainty quantification is extremely important, which provides the “confidence level" of the reconstructed results in ISPs. In this paper, four uncertainty calibration methods are introduced to better quantify the uncertainties and improve the reconstructions in deep-learning based ISP solvers, which includes Deep-Ensemble, Drop-Block, Drop-Channel, and Loss-Balance. Different calibration methods are evaluated in terms of their impacts on the accuracies of reconstructed results in solving ISPs and how well they quantify the uncertainties of the full-wave reconstructed results, which are further compared with previously proposed uncertainty quantification methods including MC-Dropout and generative models in ISPs. Our synthetic and experimental results show that the methods of Deep-Ensemble and Drop-block can achieve better performance than other methods quantitatively in terms of uncertainty quantifications. Besides, Loss-Balance has a significant effect on the preservation of reconstruction accuracy. It is expected that the results and conclusions are helpful to better quantify uncertainties and push learning methods to a more reliable way in full-wave nonlinear reconstructions of ISPs. Gensheng Zhang, Zhun Wei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Maverick: Discovering Exceptional Facts from Knowledge GraphsabstractWe present Maverick, a general, extensible framework that discovers exceptional facts about entities in knowledge graphs. To the best of our knowledge, there was no previous study of the problem. We model an exceptional fact about an entity of interest as a context-subspace pair, in which a subspace is a set of attributes and a context is defined by a graph query pattern of which the entity is a match. The entity is exceptional among the entities in the context, with regard to the subspace. The search spaces of both patterns and subspaces are exponentially large. Maverick conducts beam search on the patterns which uses a match-based pattern construction method to evade the evaluation of invalid patterns. It applies two heuristics to select promising patterns to form the beam in each iteration. Maverick traverses and prunes the subspaces organized as a set enumeration tree by exploiting the upper bound properties of exceptionality scoring functions. Results of experiments and user studies using real-world datasets demonstrated substantial performance improvement of the proposed framework over the baselines as well as its effectiveness in discovering exceptional facts. Gensheng Zhang, Damian Jimenez, Chengkai Li 0001 |
SIGMOD Conference | 1 |
| 2018 | Maverick: A System for Discovering Exceptional Facts from Knowledge GraphsabstractThis paper presents Maverick, a system for discovering exceptional facts about entities in knowledge graphs. Maverick is built upon a beam-search based algorithmic framework which we proposed in a research paper that is published in SIGMOD 2018. In this demonstration proposal, we showcase an end-to-end system that includes a user-facing portal and a cache server. In Maverick, an exceptional fact about an entity of interest is modeled as a context-subspace pair, in which the subspace is a set of attributes and the context is defined by a graph query pattern of which the entity is a match, together with other matching entities. The entity is exceptional among the entities in the context, with regard to the sub-space. The portal allows users to search entities in a knowledge graph and explores exceptional facts about the entities of interest. It presents exceptional facts to users in forms of natural language sentences and illustration charts, for better interpretability of the discovered exceptional facts. The cache server stores intermediate computation results, such as pattern evaluations, exceptionality calculations, and candidate patterns. It is built for sharing computation across entities, such that repetitive computation across entities can be avoided. Gensheng Zhang, Chengkai Li 0001 |
Proc. VLDB Endow. | 1 |
| 2017 | ClaimBuster: The First-ever End-to-end Fact-checking SystemabstractOur society is struggling with an unprecedented amount of falsehoods, hyperboles, and half-truths. Politicians and organizations repeatedly make the same false claims. Fake news floods the cyberspace and even allegedly influenced the 2016 election. In fighting false information, the number of active fact-checking organizations has grown from 44 in 2014 to 114 in early 2017. 1 Fact-checkers vet claims by investigating relevant data and documents and publish their verdicts. For instance, PolitiFact.com, one of the earliest and most popular fact-checking projects, gives factual claims truthfulness ratings such as True, Mostly True, Half true, Mostly False, False, and even "Pants on Fire". In the U.S., the election year made fact-checking a part of household terminology. For example, during the first presidential debate on September 26, 2016, NPR.org's live fact-checking website drew 7.4 million page views and delivered its biggest traffic day ever. Naeemul Hassan, Gensheng Zhang, Fatma Arslan, Josue Caraballo, Damian Jimenez, Siddhant Gawsane, Shohedul Hasan, Minumol Joseph, Aaditya Kulkarni, Anil Kumar Nayak, Vikas Sable, Chengkai Li 0001, Mark Tremayne |
Proc. VLDB Endow. | 2 |
| 2015 | Crowdsourcing Pareto-Optimal Object Finding By Pairwise ComparisonsabstractThis is the first study of crowdsourcing Pareto-optimal object finding over partial orders and by pairwise comparisons, which has applications in public opinion collection, group decision making, and information exploration. Departing from prior studies on crowdsourcing skyline and ranking queries, it considers the case where objects do not have explicit attributes and preference relations on objects are strict partial orders. The partial orders are derived by aggregating crowdsourcers' responses to pairwise comparison questions. The goal is to find all Pareto-optimal objects by the fewest possible questions. It employs an iterative question-selection framework. Guided by the principle of eagerly identifying non-Pareto optimal objects, the framework only chooses candidate questions which must satisfy three conditions. This design is both sufficient and efficient, as it is proven to find a short terminal question sequence. The framework is further steered by two ideas---macro-ordering and micro-ordering. By different micro-ordering heuristics, the framework is instantiated into several algorithms with varying power in pruning questions. Experiment results using both real crowdsourcing marketplace and simulations exhibited not only orders of magnitude reductions in questions when compared with a brute-force approach, but also close-to-optimal performance from the most efficient instantiation. Abolfazl Asudeh, Gensheng Zhang, Naeemul Hassan, Chengkai Li 0001, Gergely V. Záruba |
CIKM | 2 |
| 2015 | Fourier irregularity index: A new approach to measure tumor mass irregularity in breast mammogram images
Gensheng Zhang, Wei Wang 0015, Sung Y. Shin, Carrie B. Hruska, Seong-Ho Son |
Multim. Tools Appl. | 1 |
| 2014 | Data In, Fact Out: Automated Monitoring of Facts by FactWatcherabstractTowards computational journalism, we present FactWatcher, a system that helps journalists identify data-backed, attention-seizing facts which serve as leads to news stories. FactWatcher discovers three types of facts, including situational facts, one-of-the-few facts, and prominent streaks, through a unified suite of data model, algorithm framework, and fact ranking measure. Given an append-only database, upon the arrival of a new tuple, FactWatcher monitors if the tuple triggers any new facts. Its algorithms efficiently search for facts without exhaustively testing all possible ones. Furthermore, FactWatcher provides multiple features in striving for an end-to-end system, including fact ranking, fact-to-statement translation and keyword-based fact search. Naeemul Hassan, Afroza Sultana, You Wu 0001, Gensheng Zhang, Chengkai Li 0001, Jun Yang 0001, Cong Yu 0001 |
Proc. VLDB Endow. | 4 |
| 2013 | Discovering General Prominent Streaks in Sequence DataabstractThis article studies the problem of prominent streak discovery in sequence data. Given a sequence of values, a prominent streak is a long consecutive subsequence consisting of only large (small) values, such as consecutive games of outstanding performance in sports, consecutive hours of heavy network traffic, and consecutive days of frequent mentioning of a person in social media. Prominent streak discovery provides insightful data patterns for data analysis in many real-world applications and is an enabling technique for computational journalism. Given its real-world usefulness and complexity, the research on prominent streaks in sequence data opens a spectrum of challenging problems. A baseline approach to finding prominent streaks is a quadratic algorithm that exhaustively enumerates all possible streaks and performs pairwise streak dominance comparison. For more efficient methods, we make the observation that prominent streaks are in fact skyline points in two dimensions—streak interval length and minimum value in the interval. Our solution thus hinges on the idea to separate the two steps in prominent streak discovery: candidate streak generation and skyline operation over candidate streaks. For candidate generation, we propose the concept of local prominent streak (LPS). We prove that prominent streaks are a subset of LPSs and the number of LPSs is less than the length of a data sequence, in comparison with the quadratic number of candidates produced by the brute-force baseline method. We develop efficient algorithms based on the concept of LPS. The nonlinear local prominent streak (NLPS)-based method considers a superset of LPSs as candidates, and the linear local prominent streak (LLPS)-based method further guarantees to consider only LPSs. The proposed properties and algorithms are also extended for discovering general top- k , multisequence, and multidimensional prominent streaks. The results of experiments using multiple real datasets verified the effectiveness of the proposed methods and showed orders of magnitude performance improvement against the baseline method. Gensheng Zhang, Ping Luo 0001, Min Wang 0001, Chengkai Li 0001 |
ACM Trans. Knowl. Discov. Data | 1 |