EDBT 2026 Demo / reviewers in the wild / expert
Jinsong Guo
dblp:91/10056
· DBLP profile ↗
16ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-1142-3610ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 72% Language models and text generation · 11% Information extraction and text analysis · 10% | |
| Databases, data mining, and information retrieval
3 papers |
Data integration and cleaning · 94% Information retrieval · 6% | |
| Network and information security
1 paper |
Systems and software security · 100% | |
| Theoretical computer science
1 paper |
Computational complexity · 50% Automated reasoning and model checking · 50% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
machine unlearning |
0.9 | 1 | 2025 | Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models · AAAI 2025 |
Machine learning › Trustworthy machine learning
privacy |
0.9 | 1 | 2025 | Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models · AAAI 2025 |
Systems and software security › vulnerability discovery
vulnerability classification |
0.9 | 1 | 2025 | Enabling Generalized Zero-Shot Vulnerability Classification · IEEE Trans. Dependable Secur. Comput. 2025 |
Systems and software security
vulnerability discovery |
0.9 | 1 | 2025 | Enabling Generalized Zero-Shot Vulnerability Classification · IEEE Trans. Dependable Secur. Comput. 2025 |
Data integration and cleaning
entity relationship discovery |
0.7 | 1 | 2023 | When Automatic Filtering Comes to the Rescue: Pre-Computing Company Competitor Pairs in Owler · Proc. ACM Manag. Data 2023 |
Data integration and cleaning › data extraction
web data extraction |
0.5 | 2 | 2019 | RED: Redundancy-Driven Data Extraction from Result Pages? · WWW 2019 Robust and Noise Resistant Wrapper Induction · SIGMOD Conference 2016 |
Data integration and cleaning
data extraction |
0.4 | 1 | 2019 | RED: Redundancy-Driven Data Extraction from Result Pages? · WWW 2019 |
Natural language and speech › Information extraction and text analysis › web information extraction
wrapper induction |
0.2 | 1 | 2016 | Robust and Noise Resistant Wrapper Induction · SIGMOD Conference 2016 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › global constraints
table constraints |
0.2 | 1 | 2013 | Making Simple Tabular ReductionWorks on Negative Table Constraints · AAAI 2013 |
Computational complexity › constraint satisfaction
constraint propagation |
0.2 | 1 | 2013 | Making Simple Tabular ReductionWorks on Negative Table Constraints · AAAI 2013 |
Automated reasoning and model checking › constraint solving
generalized arc consistency |
0.2 | 1 | 2013 | Making Simple Tabular ReductionWorks on Negative Table Constraints · AAAI 2013 |
Information retrieval
search interfaces |
0.1 | 1 | 2019 | RED: Redundancy-Driven Data Extraction from Result Pages? · WWW 2019 |
Methods — techniques the papers use, named apart from their topics
sensitive span annotation · 0.9selective unlearning · 0.9generalized zero-shot learning · 0.9description-based class embedding · 0.9LLM-based verification · 0.9inference from existing competitors · 0.7empirical evidence validation · 0.7query language induction · 0.5XPath · 0.5unsupervised learning · 0.4simple tabular reduction · 0.3negative table constraints · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GaV: Guess and Verification of Column Semantics
Davide Di Stefano, Jinsong Guo, Matteo Capalbo, Davide M. Longo, Georg Gottlob |
ICDE | 2 |
| 2026 | A multi-feature alignment fusion neural network model for red blood cell aggregation classification using ultrasonic radiofrequency data of blood
Jinsong Guo, Bingbing He, Xun Lang |
Artif. Intell. Medicine | 1 |
| 2025 | Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language ModelsabstractThis paper explores Machine Unlearning (MU), an emerging field that is gaining increased attention due to concerns about neural models unintentionally remembering personal or sensitive information. We present SeUL, a novel method that enables selective and fine-grained unlearning for language models. Unlike previous work that employs a fully reversed training objective in unlearning, SeUL minimizes the negative impact on the capability of language models, particularly in terms of generation. Furthermore, we introduce two innovative evaluation metrics, sensitive extraction likelihood (S-EL) and sensitive memorization accuracy (S-MA), specifically designed to assess the effectiveness of forgetting sensitive information. In support of the unlearning framework, we propose efficient automatic online and offline sensitive span annotation methods. The online selection method, based on language probability scores, ensures computational efficiency, while the offline annotation involves a two-stage LLM-based process for robust verification. In summary, this paper contributes a novel selective unlearning method (SeUL), introduces specialized evaluation metrics (S-EL and S-MA) for assessing sensitive information forgetting, and proposes automatic online and offline sensitive span annotation methods to support the overall unlearning framework and evaluation. Lingzhi Wang 0001, Xingshan Zeng, Jinsong Guo, Kam-Fai Wong, Georg Gottlob |
AAAI | 3 |
| 2025 | Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human PerceptionabstractThe pervasive spread of misinformation and disinformation in social media underscores the critical importance of detecting media bias. While robust Large Language Models (LLMs) have emerged as foundational tools for bias prediction, concerns about inherent biases within these models persist. In this work, we investigate the presence and nature of bias within LLMs and its consequential impact on media bias detection. Departing from conventional approaches that focus solely on bias detection in media content, we delve into biases within the LLM systems themselves. Through meticulous examination, we probe whether LLMs exhibit biases, particularly in political bias prediction and text continuation tasks. Additionally, we explore bias across diverse topics, aiming to uncover nuanced variations in bias expression within the LLM framework. Importantly, we propose debiasing strategies, including prompt engineering and model fine-tuning. Extensive analysis of bias tendencies across different LLMs sheds light on the broader landscape of bias propagation in language models. This study advances our understanding of LLM bias, offering critical insights into its implications for bias detection tasks and paving the way for more robust and equitable AI systems Luyang Lin, Lingzhi Wang 0001, Jinsong Guo, Kam-Fai Wong |
COLING | 3 |
| 2025 | IndiTag: An Online Media Bias Analysis System Using Fine-Grained Bias IndicatorsabstractIn the age of information overload and polarized discourse, understanding media bias has become imperative for informed decision-making and fostering a balanced public discourse. However, without the experts' analysis, it is hard for the readers to distinguish bias from the news articles. This paper presents IndiTag, an innovative online media bias analysis system that leverages fine-grained bias indicators to dissect and distinguish bias in digital content. IndiTag offers a novel approach by incorporating large language models, bias indicators, and vector database to detect and interpret bias automatically. Complemented by a user-friendly interface facilitating automated bias analysis for readers, IndiTag offers a comprehensive platform for in-depth bias examination. We demonstrate the efficacy and versatility of IndiTag through experiments on four datasets encompassing news articles from diverse platforms. Furthermore, we discuss potential applications of IndiTag in fostering media literacy, facilitating fact-checking initiatives, and enhancing the transparency and accountability of digital media platforms. IndiTag stands as a valuable tool in the pursuit of fostering a more informed, discerning, and inclusive public discourse in the digital age. We release an online system for end users and the source code is available at https://github.com/lylin0/IndiTag. Luyang Lin, Lingzhi Wang 0001, Jinsong Guo, Jing Li 0049, Kam-Fai Wong |
ICPADS | 3 |
| 2025 | Enabling Generalized Zero-Shot Vulnerability ClassificationabstractRegarding computer security, the growth of code vulnerability types presents a persistent challenge. These vulnerabilities, which may cause severe consequences, necessitate precise classification for effective mitigation. However, the rapid emergence of new vulnerability types complicates the classification process. Traditional methodologies, which often involve human expertise and the manual labeling or generation of example instances, are not only resource-intensive but also struggle to adapt to the dynamic nature of these vulnerabilities. This article introducesVulnSense, an innovative method that harnesses the capabilities of Generalized Zero-Shot Learning (GZSL) to address the vulnerability classification problem.VulnSenselearns about unseen vulnerability classes from the descriptions of these unseen classes, while not requiring to see any instances of these unseen classes. Specifically,VulnSenselearns from three main resources: 1) seen classes with labeled code instances; 2) descriptions of these seen classes; and 3) descriptions of “unseen” classes which have no labeled instances. Our experiments underscoreVulnSense's superiority over existing GZSL methods in classifying instances of unseen classes. Concurrently, it maintains a performance parity with traditional labeled-example based learning methods in classifying instances of seen vulnerabilities.VulnSensedemonstrates the potential of using GZSL for vulnerability classification, while also highlighting challenges that inspire future work. Jinghao Hu 0001, Jinsong Guo, Chen Luo 0003, Matthias Lanzinger, Zhanshan Li |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | When Automatic Filtering Comes to the Rescue: Pre-Computing Company Competitor Pairs in OwlerabstractCompetitor data constitutes information significantly valuable for many business applications. Meltwater provides users with access to a large Company Information System (CIS), Owler, which contains competitor pairs and other useful information about companies. Meltwater has been seeking a practical solution to discover more competitor pairs in Owler. The first attempt, a fully-manual workflow (called MW_Manual) for finding more competitor pairs in Owler consisted of two manual steps: a filtering step that excludes obvious non-competitor company pairs, and a further inspection process that inspects each left company pair after the filtering step. MW_Manual was cost prohibitive because the results of the filtering step contained too many non-competitor pairs. Inspecting such non-competitor pairs caused an overhead to the overall workload. To reduce the manual workload, especially the required human effort in the manual inspection process, Meltwater has transformed MW_Manual into a semi-automatic workflow (called MW_CPFilter) by replacing the manual filtering with an automatic yet more precise process that adopts a system called CPFilter. This paper presents CPFilter, a system used in the filtering process of MW_CPFilter. CPFilter automatically pre-computes likely competitor pairs from existing competitor pairs in Owler. CPFilter combines (i) the generation of new competitor candidate pairs by inference from existing competitors and other company-specific knowledge, with (ii) the validation of each candidate competitor pair of two companies by checking whether or not empirical evidence that indicates the competitor relationships of these two companies can be found. CPFilter has three key advantages compared with the manual filtering process and previous works: (i) it resulted in a high workload reduction rate of 0.81, (ii) it is domain-independent so that it can be applied to different sectors in Owler, and (iii) its results are explainable so that humans can easily understand its results. Jinsong Guo, Aditya Jami, Markus Kröll, Lukas Schweizer, Sergey Paramonov 0001, Eric Aichinger, Stefano Sferrazza, Mattia Scaccia, Stéphane Reissfelder, Eda Cicek, Giovanni Grasso 0001, Georg Gottlob |
Proc. ACM Manag. Data | 1 |
| 2021 | Revisiting the efficacy of weak consistencies: a study of forward checking
Zhe Li 0017, Zhezhou Yu, Hongbo Li 0005, Jinsong Guo, Zhanshan Li |
Sci. China Inf. Sci. | 4 |
| 2019 | RED: Redundancy-Driven Data Extraction from Result Pages?abstractData-driven websites are mostly accessed through search interfaces. Such sites follow a common publishing pattern that, surprisingly, has not been fully exploited for unsupervised data extraction yet: the result of a search is presented as a paginated list of result records. Each result record contains the main attributes about one single object, and links to a page dedicated to the details of that object. Jinsong Guo, Valter Crescenzi, Tim Furche, Giovanni Grasso 0001, Georg Gottlob |
WWW | 1 |
| 2016 | Robust and Noise Resistant Wrapper InductionabstractWrapper induction is the problem of automatically inferring a query from annotated web pages of the same template. This query should not only select the annotated content accurately but also other content following the same template. Beyond accurately matching the template, we consider two additional requirements: (1) wrappers should be robust against a large class of changes to the web pages, and (2) the induction process should be noise resistant, i.e., tolerate slightly erroneous (e.g., machine generated) samples. Key to our approach is a query language that is powerful enough to permit accurate selection, but limited enough to force noisy samples to be generalized into wrappers that select the likely intended items. We introduce such a language as subset of XPATH and show that even for such a restricted language, inducing optimal queries according to a suitable scoring is infeasible. Nevertheless, our wrapper induction framework infers highly robust and noise resistant queries. We evaluate the queries on snapshots from web pages that change over time as provided by the Internet Archive, and show that the induced queries are as robust as the human-made queries. The queries often survive hundreds sometimes thousands of days, with many changes to the relative position of the selected nodes (including changes on template level). This is due to the few and discriminative anchor (intermediately selected) nodes of the generated queries. The queries are highly resistant against positive noise (up to 50%) and negative noise (up to 20%). Tim Furche, Jinsong Guo, Sebastian Maneth, Christian Schallhart |
SIGMOD Conference | 2 |
| 2013 | Making Simple Tabular ReductionWorks on Negative Table ConstraintsabstractSimple Tabular Reduction algorithms (STR) work well to establish Generalized Arc Consistency (GAC) on positive table constraints. However, the existing STR algorithms are useless for negative table constraints. In this work, we propose a novel STR algorithm and its improvement, which work on negative table constraints. Our preliminary experiments are performed on some random instances and a certain benchmark instances. The results show that the new algorithms outperform GAC-valid and the MDD-based GAC algorithm. Hongbo Li 0005, Yanchun Liang 0001, Jinsong Guo, Zhanshan Li |
AAAI | 3 |
| 2013 | Reducing consistency checks in generating corrective explanations for interactive constraint satisfaction
Hongbo Li 0005, Haijiao Shen, Zhanshan Li, Jinsong Guo |
Knowl. Based Syst. | 4 |
| 2012 | Partial Max-restricted Path ConsistencyabstractFiltering techniques are essential in the search algorithms solving constraint satisfaction problems (CSPs). Arc consistency (AC) is the most often used filtering technique because it cheaply removes some values that cannot belong to any solutions. Comparing with AC, max-Restricted Path Consistency (maxRPC) has a stronger pruning power while it is not suited for use during search because of the prohibitive time cost. Thus, light maxRPC which is the light version of maxRPC was proposed. Comparing with maxRPC, it has a lower time cost and the search algorithm maintaining light maxRPC (MlmaxRPC) can outperform the search algorithm maintaining AC (MAC) on some problems. However, MlmaxRPC suffers from the time waste in the cases that applying a stricter checking standard does not intrigue any value deletion. In order to avoid the time waste in MlmaxRPC, in this paper, partial maxRPC which is a new approximation of maxRPC is proposed. It only applies the stricter checking standard when the value deletion is of high possibility. MpmaxRPC which is the search algorithm maintaining partial maxRPC has a better average performance than MAC and MlmaxRPC. Jinsong Guo, Zhanshan Li, Hongbo Li 0005 |
ICTAI | 1 |
| 2012 | Efficient Singleton Consistency by Combining Forward Checking and Bound ConsistencyabstractMaintaining local consistencies can improve the efficiencies of the search algorithms solving constraint satisfaction problems (CSPs). Comparing with arc consistency which is the most widely used local consistency, stronger local consistencies can make the search space smaller while they require higher computational cost. In this paper, we make an attempt on the compromise between the pruning ability and the computational cost. A new local consistency called singleton strong bound consistency (SSBC) and its light version, light SSBC, are proposed. The search algorithm maintaining light SSBC can outperform MAC on a considerable number of problems. Jinsong Guo, Zhanshan Li, Yonggang Zhang 0002 |
ICTAI | 1 |
| 2011 | MaxRPC Algorithms Based on Bitwise Operations
Jinsong Guo, Zhanshan Li, Xuena Geng |
CP | 1 |
| 2011 | Disassembling and Reconstructing Algorithms for Discrete Event SystemsabstractThis paper addresses the problem of failure diagnosis in component-based discrete event systems. In this paper we propose a method to obtain the set of components when dealing with diagnosis in large complex discrete event systems. In the new method, before disassembling the system into components, we need to identify whether insert communication events into the system or not. When analyzing the diagnosability, we treat the system containing communication events as a distributed discrete event system. Otherwise we treat the system as a decentralized discrete event system. For the components which are not diagnosable, we propose a method to reconstruct them by utilizing some other components sharing the same communication events with them. This algorithm provides more accurate information of the diagnosability of the system. Xuena Geng, Dantong Ouyang, Jinsong Guo |
ICTAI | 3 |