Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jinsong Guo

dblp:91/10056 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-1142-3610ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 72% Language models and text generation · 11% Information extraction and text analysis · 10%
Databases, data mining, and information retrieval
3 papers
Data integration and cleaning · 94% Information retrieval · 6%
Network and information security
1 paper
Systems and software security · 100%
Theoretical computer science
1 paper
Computational complexity · 50% Automated reasoning and model checking · 50%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
machine unlearning
0.912025
Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models · AAAI 2025
Machine learning › Trustworthy machine learning
privacy
0.912025
Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models · AAAI 2025
Systems and software security › vulnerability discovery
vulnerability classification
0.912025
Enabling Generalized Zero-Shot Vulnerability Classification · IEEE Trans. Dependable Secur. Comput. 2025
Systems and software security
vulnerability discovery
0.912025
Enabling Generalized Zero-Shot Vulnerability Classification · IEEE Trans. Dependable Secur. Comput. 2025
Data integration and cleaning
entity relationship discovery
0.712023
When Automatic Filtering Comes to the Rescue: Pre-Computing Company Competitor Pairs in Owler · Proc. ACM Manag. Data 2023
Data integration and cleaning › data extraction
web data extraction
0.522019
RED: Redundancy-Driven Data Extraction from Result Pages? · WWW 2019
Robust and Noise Resistant Wrapper Induction · SIGMOD Conference 2016
Data integration and cleaning
data extraction
0.412019
RED: Redundancy-Driven Data Extraction from Result Pages? · WWW 2019
Natural language and speech › Information extraction and text analysis › web information extraction
wrapper induction
0.212016
Robust and Noise Resistant Wrapper Induction · SIGMOD Conference 2016
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › global constraints
table constraints
0.212013
Making Simple Tabular ReductionWorks on Negative Table Constraints · AAAI 2013
Computational complexity › constraint satisfaction
constraint propagation
0.212013
Making Simple Tabular ReductionWorks on Negative Table Constraints · AAAI 2013
Automated reasoning and model checking › constraint solving
generalized arc consistency
0.212013
Making Simple Tabular ReductionWorks on Negative Table Constraints · AAAI 2013
Information retrieval
search interfaces
0.112019
RED: Redundancy-Driven Data Extraction from Result Pages? · WWW 2019

Methods — techniques the papers use, named apart from their topics

sensitive span annotation · 0.9selective unlearning · 0.9generalized zero-shot learning · 0.9description-based class embedding · 0.9LLM-based verification · 0.9inference from existing competitors · 0.7empirical evidence validation · 0.7query language induction · 0.5XPath · 0.5unsupervised learning · 0.4simple tabular reduction · 0.3negative table constraints · 0.3
YearPublicationVenuePosition
2026 GaV: Guess and Verification of Column Semantics
Davide Di Stefano, Jinsong Guo, Matteo Capalbo, Davide M. Longo, Georg Gottlob
ICDE2
2026 A multi-feature alignment fusion neural network model for red blood cell aggregation classification using ultrasonic radiofrequency data of blood
Jinsong Guo, Bingbing He, Xun Lang
Artif. Intell. Medicine1
2025 Selective Forgetting: Advancing Machine Unlearning Techniques and Evaluation in Language Models
abstract
This paper explores Machine Unlearning (MU), an emerging field that is gaining increased attention due to concerns about neural models unintentionally remembering personal or sensitive information. We present SeUL, a novel method that enables selective and fine-grained unlearning for language models. Unlike previous work that employs a fully reversed training objective in unlearning, SeUL minimizes the negative impact on the capability of language models, particularly in terms of generation. Furthermore, we introduce two innovative evaluation metrics, sensitive extraction likelihood (S-EL) and sensitive memorization accuracy (S-MA), specifically designed to assess the effectiveness of forgetting sensitive information. In support of the unlearning framework, we propose efficient automatic online and offline sensitive span annotation methods. The online selection method, based on language probability scores, ensures computational efficiency, while the offline annotation involves a two-stage LLM-based process for robust verification. In summary, this paper contributes a novel selective unlearning method (SeUL), introduces specialized evaluation metrics (S-EL and S-MA) for assessing sensitive information forgetting, and proposes automatic online and offline sensitive span annotation methods to support the overall unlearning framework and evaluation.
Lingzhi Wang 0001, Xingshan Zeng, Jinsong Guo, Kam-Fai Wong, Georg Gottlob
AAAI3
2025 Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception
abstract
The pervasive spread of misinformation and disinformation in social media underscores the critical importance of detecting media bias. While robust Large Language Models (LLMs) have emerged as foundational tools for bias prediction, concerns about inherent biases within these models persist. In this work, we investigate the presence and nature of bias within LLMs and its consequential impact on media bias detection. Departing from conventional approaches that focus solely on bias detection in media content, we delve into biases within the LLM systems themselves. Through meticulous examination, we probe whether LLMs exhibit biases, particularly in political bias prediction and text continuation tasks. Additionally, we explore bias across diverse topics, aiming to uncover nuanced variations in bias expression within the LLM framework. Importantly, we propose debiasing strategies, including prompt engineering and model fine-tuning. Extensive analysis of bias tendencies across different LLMs sheds light on the broader landscape of bias propagation in language models. This study advances our understanding of LLM bias, offering critical insights into its implications for bias detection tasks and paving the way for more robust and equitable AI systems
Luyang Lin, Lingzhi Wang 0001, Jinsong Guo, Kam-Fai Wong
COLING3
2025 IndiTag: An Online Media Bias Analysis System Using Fine-Grained Bias Indicators
abstract
In the age of information overload and polarized discourse, understanding media bias has become imperative for informed decision-making and fostering a balanced public discourse. However, without the experts' analysis, it is hard for the readers to distinguish bias from the news articles. This paper presents IndiTag, an innovative online media bias analysis system that leverages fine-grained bias indicators to dissect and distinguish bias in digital content. IndiTag offers a novel approach by incorporating large language models, bias indicators, and vector database to detect and interpret bias automatically. Complemented by a user-friendly interface facilitating automated bias analysis for readers, IndiTag offers a comprehensive platform for in-depth bias examination. We demonstrate the efficacy and versatility of IndiTag through experiments on four datasets encompassing news articles from diverse platforms. Furthermore, we discuss potential applications of IndiTag in fostering media literacy, facilitating fact-checking initiatives, and enhancing the transparency and accountability of digital media platforms. IndiTag stands as a valuable tool in the pursuit of fostering a more informed, discerning, and inclusive public discourse in the digital age. We release an online system for end users and the source code is available at https://github.com/lylin0/IndiTag.
Luyang Lin, Lingzhi Wang 0001, Jinsong Guo, Jing Li 0049, Kam-Fai Wong
ICPADS3
2025 Enabling Generalized Zero-Shot Vulnerability Classification
abstract
Regarding computer security, the growth of code vulnerability types presents a persistent challenge. These vulnerabilities, which may cause severe consequences, necessitate precise classification for effective mitigation. However, the rapid emergence of new vulnerability types complicates the classification process. Traditional methodologies, which often involve human expertise and the manual labeling or generation of example instances, are not only resource-intensive but also struggle to adapt to the dynamic nature of these vulnerabilities. This article introducesVulnSense, an innovative method that harnesses the capabilities of Generalized Zero-Shot Learning (GZSL) to address the vulnerability classification problem.VulnSenselearns about unseen vulnerability classes from the descriptions of these unseen classes, while not requiring to see any instances of these unseen classes. Specifically,VulnSenselearns from three main resources: 1) seen classes with labeled code instances; 2) descriptions of these seen classes; and 3) descriptions of “unseen” classes which have no labeled instances. Our experiments underscoreVulnSense's superiority over existing GZSL methods in classifying instances of unseen classes. Concurrently, it maintains a performance parity with traditional labeled-example based learning methods in classifying instances of seen vulnerabilities.VulnSensedemonstrates the potential of using GZSL for vulnerability classification, while also highlighting challenges that inspire future work.
Jinghao Hu 0001, Jinsong Guo, Chen Luo 0003, Matthias Lanzinger, Zhanshan Li
IEEE Trans. Dependable Secur. Comput.2
2023 When Automatic Filtering Comes to the Rescue: Pre-Computing Company Competitor Pairs in Owler
abstract
Competitor data constitutes information significantly valuable for many business applications. Meltwater provides users with access to a large Company Information System (CIS), Owler, which contains competitor pairs and other useful information about companies. Meltwater has been seeking a practical solution to discover more competitor pairs in Owler. The first attempt, a fully-manual workflow (called MW_Manual) for finding more competitor pairs in Owler consisted of two manual steps: a filtering step that excludes obvious non-competitor company pairs, and a further inspection process that inspects each left company pair after the filtering step. MW_Manual was cost prohibitive because the results of the filtering step contained too many non-competitor pairs. Inspecting such non-competitor pairs caused an overhead to the overall workload. To reduce the manual workload, especially the required human effort in the manual inspection process, Meltwater has transformed MW_Manual into a semi-automatic workflow (called MW_CPFilter) by replacing the manual filtering with an automatic yet more precise process that adopts a system called CPFilter. This paper presents CPFilter, a system used in the filtering process of MW_CPFilter. CPFilter automatically pre-computes likely competitor pairs from existing competitor pairs in Owler. CPFilter combines (i) the generation of new competitor candidate pairs by inference from existing competitors and other company-specific knowledge, with (ii) the validation of each candidate competitor pair of two companies by checking whether or not empirical evidence that indicates the competitor relationships of these two companies can be found. CPFilter has three key advantages compared with the manual filtering process and previous works: (i) it resulted in a high workload reduction rate of 0.81, (ii) it is domain-independent so that it can be applied to different sectors in Owler, and (iii) its results are explainable so that humans can easily understand its results.
Jinsong Guo, Aditya Jami, Markus Kröll, Lukas Schweizer, Sergey Paramonov 0001, Eric Aichinger, Stefano Sferrazza, Mattia Scaccia, Stéphane Reissfelder, Eda Cicek, Giovanni Grasso 0001, Georg Gottlob
Proc. ACM Manag. Data1
2021 Revisiting the efficacy of weak consistencies: a study of forward checking
Zhe Li 0017, Zhezhou Yu, Hongbo Li 0005, Jinsong Guo, Zhanshan Li
Sci. China Inf. Sci.4
2019 RED: Redundancy-Driven Data Extraction from Result Pages?
abstract
Data-driven websites are mostly accessed through search interfaces. Such sites follow a common publishing pattern that, surprisingly, has not been fully exploited for unsupervised data extraction yet: the result of a search is presented as a paginated list of result records. Each result record contains the main attributes about one single object, and links to a page dedicated to the details of that object.
Jinsong Guo, Valter Crescenzi, Tim Furche, Giovanni Grasso 0001, Georg Gottlob
WWW1
2016 Robust and Noise Resistant Wrapper Induction
abstract
Wrapper induction is the problem of automatically inferring a query from annotated web pages of the same template. This query should not only select the annotated content accurately but also other content following the same template. Beyond accurately matching the template, we consider two additional requirements: (1) wrappers should be robust against a large class of changes to the web pages, and (2) the induction process should be noise resistant, i.e., tolerate slightly erroneous (e.g., machine generated) samples. Key to our approach is a query language that is powerful enough to permit accurate selection, but limited enough to force noisy samples to be generalized into wrappers that select the likely intended items. We introduce such a language as subset of XPATH and show that even for such a restricted language, inducing optimal queries according to a suitable scoring is infeasible. Nevertheless, our wrapper induction framework infers highly robust and noise resistant queries. We evaluate the queries on snapshots from web pages that change over time as provided by the Internet Archive, and show that the induced queries are as robust as the human-made queries. The queries often survive hundreds sometimes thousands of days, with many changes to the relative position of the selected nodes (including changes on template level). This is due to the few and discriminative anchor (intermediately selected) nodes of the generated queries. The queries are highly resistant against positive noise (up to 50%) and negative noise (up to 20%).
Tim Furche, Jinsong Guo, Sebastian Maneth, Christian Schallhart
SIGMOD Conference2
2013 Making Simple Tabular ReductionWorks on Negative Table Constraints
abstract
Simple Tabular Reduction algorithms (STR) work well to establish Generalized Arc Consistency (GAC) on positive table constraints. However, the existing STR algorithms are useless for negative table constraints. In this work, we propose a novel STR algorithm and its improvement, which work on negative table constraints. Our preliminary experiments are performed on some random instances and a certain benchmark instances. The results show that the new algorithms outperform GAC-valid and the MDD-based GAC algorithm.
Hongbo Li 0005, Yanchun Liang 0001, Jinsong Guo, Zhanshan Li
AAAI3
2013 Reducing consistency checks in generating corrective explanations for interactive constraint satisfaction
Hongbo Li 0005, Haijiao Shen, Zhanshan Li, Jinsong Guo
Knowl. Based Syst.4
2012 Partial Max-restricted Path Consistency
abstract
Filtering techniques are essential in the search algorithms solving constraint satisfaction problems (CSPs). Arc consistency (AC) is the most often used filtering technique because it cheaply removes some values that cannot belong to any solutions. Comparing with AC, max-Restricted Path Consistency (maxRPC) has a stronger pruning power while it is not suited for use during search because of the prohibitive time cost. Thus, light maxRPC which is the light version of maxRPC was proposed. Comparing with maxRPC, it has a lower time cost and the search algorithm maintaining light maxRPC (MlmaxRPC) can outperform the search algorithm maintaining AC (MAC) on some problems. However, MlmaxRPC suffers from the time waste in the cases that applying a stricter checking standard does not intrigue any value deletion. In order to avoid the time waste in MlmaxRPC, in this paper, partial maxRPC which is a new approximation of maxRPC is proposed. It only applies the stricter checking standard when the value deletion is of high possibility. MpmaxRPC which is the search algorithm maintaining partial maxRPC has a better average performance than MAC and MlmaxRPC.
Jinsong Guo, Zhanshan Li, Hongbo Li 0005
ICTAI1
2012 Efficient Singleton Consistency by Combining Forward Checking and Bound Consistency
abstract
Maintaining local consistencies can improve the efficiencies of the search algorithms solving constraint satisfaction problems (CSPs). Comparing with arc consistency which is the most widely used local consistency, stronger local consistencies can make the search space smaller while they require higher computational cost. In this paper, we make an attempt on the compromise between the pruning ability and the computational cost. A new local consistency called singleton strong bound consistency (SSBC) and its light version, light SSBC, are proposed. The search algorithm maintaining light SSBC can outperform MAC on a considerable number of problems.
Jinsong Guo, Zhanshan Li, Yonggang Zhang 0002
ICTAI1
2011 MaxRPC Algorithms Based on Bitwise Operations
Jinsong Guo, Zhanshan Li, Xuena Geng
CP1
2011 Disassembling and Reconstructing Algorithms for Discrete Event Systems
abstract
This paper addresses the problem of failure diagnosis in component-based discrete event systems. In this paper we propose a method to obtain the set of components when dealing with diagnosis in large complex discrete event systems. In the new method, before disassembling the system into components, we need to identify whether insert communication events into the system or not. When analyzing the diagnosability, we treat the system containing communication events as a distributed discrete event system. Otherwise we treat the system as a decentralized discrete event system. For the components which are not diagnosable, we propose a method to reconstruct them by utilizing some other components sharing the same communication events with them. This algorithm provides more accurate information of the diagnosability of the system.
Xuena Geng, Dantong Ouyang, Jinsong Guo
ICTAI3