Zongxiong Chen

dblp:256/6334 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0003-2452-0572ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 46% Trustworthy machine learning · 16% Vision and language · 14%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 67% Data stream processing · 33%
Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
entity linking
1.012026
NERdME: A Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories · WWW 2026
Natural language and speech › Information extraction and text analysis
named entity recognition
1.012026
NERdME: A Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories · WWW 2026
Natural language and speech › Information extraction and text analysis › document analysis › scholarly text analysis
scientific information extraction
1.012026
NERdME: A Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories · WWW 2026
Natural language and speech › Language models and text generation
hallucination detection
0.912025
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs · ACL (1) 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update · AAAI 2025
Computer vision › Vision and language
vision-language model
0.912025
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update · AAAI 2025
Machine learning › Efficient and distributed learning
dataset distillation
0.712023
A Survey on Dataset Distillation: Approaches, Applications and Future Directions · IJCAI 2023
Query processing and optimization › query compilation
just-in-time compilation
0.412020
Grizzly: Efficient Stream Processing Through Adaptive Query Compilation · SIGMOD Conference 2020
Query processing and optimization
query compilation
0.412020
Grizzly: Efficient Stream Processing Through Adaptive Query Compilation · SIGMOD Conference 2020
Data stream processing
stream processing systems
0.412020
Grizzly: Efficient Stream Processing Through Adaptive Query Compilation · SIGMOD Conference 2020
Software maintenance and evolution
software ecosystems
0.312026
NERdME: A Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories · WWW 2026
Security and privacy of machine learning › adversarial attack
jailbreak attack
0.312025
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update · AAAI 2025
Machine learning › Trustworthy machine learning
privacy and data protection
0.212023
A Survey on Dataset Distillation: Approaches, Applications and Future Directions · IJCAI 2023

Methods — techniques the papers use, named apart from their topics

large language model · 2.0fine-tuned transformer · 2.0contrastive sample construction · 1.7activation revision · 1.7neural differential equation · 0.9just-in-time compilation · 0.9adaptive compilation · 0.9
YearPublicationVenuePosition
2026 NERdME: A Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories
abstract
Existing scholarly information extraction (SIE) datasets focus on scientific papers and overlook implementation-level details in code repositories. README files describe datasets, source code, and other implementation-level artifacts, however, their free-form Markdown offers little semantic structure, making automatic information extraction difficult. To address this gap, NERdME is introduced: 200 manually annotated README files with over νm10000 labeled spans and 10 entity types. Baseline results using large language models and fine-tuned transformers show clear differences between paper-level and implementation-level entities, indicating the value of extending SIE benchmarks with entity types available in README files. A downstream entity-linking experiment was conducted to demonstrate that entities derived from READMEs can support artifact discovery and metadata integration.
Genet Asefa Gesese, Zongxiong Chen, Shufan Jiang 0001, Mary Ann Tan, Zhaotai Liu, Sonja Schimmler, Harald Sack
WWW2
2025 Internal Activation Revision: Safeguarding Vision Language Models Without Parameter Update
abstract
Warning: This paper contains offensive content that may disturb some readers. Vision-language models (VLMs) demonstrate strong multimodal capabilities but have been found to be more susceptible to generating harmful content compared to their backbone large language models (LLMs). Our investigation reveals that the integration of images significantly shifts the model's internal activations during the forward pass, diverging from those triggered by textual input. Moreover, the safety alignments of LLMs embedded within VLMs are not sufficiently robust to handle the activations discrepancies, making the models vulnerable to even the simplest jailbreaking attacks. To address this issue, we propose an internal activation revision approach that efficiently revises activations during generation, steering the model toward safer outputs. Our framework incorporates revisions at both the layer and head levels, offering control over the model's generation at varying levels of granularity. In addition, we explore three strategies for constructing positive and negative samples and two approaches for extracting revision vectors, resulting in different variants of our method. Comprehensive experiments demonstrate that the internal activation revision method significantly improves the safety of widely used VLMs, reducing attack success rates by an average of 48.94%, 34.34%, 43.92%, and 52.98% on SafeBench, Safe-Unsafe, Unsafe, and MM-SafetyBench, respectively, while minimally impacting model helpfulness.
Qing Li 0038, Jiahui Geng, Derui Zhu, Zongxiong Chen, Kun Song 0001, Lei Ma 0003, Fakhri Karray
AAAI4
2025 HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
abstract
6173
Qing Li 0038, Jiahui Geng, Zongxiong Chen, Derui Zhu, Yuxia Wang 0003, Congbo Ma, Chenyang Lyu, Fakhri Karray
ACL (1)3
2025 GDCK: Efficient Large-Scale Graph Distillation Utilizing a Model-Free Kernelized Approach
Yue Zhang 0069, Zongxiong Chen, Sonja Schimmler, Manfred Hauswirth
PAKDD (7)2
2024 Towards Trustworthy Dataset Distillation: A Benchmark of Privacy, Fairness and Robustness
abstract
Dataset distillation is an increasingly prevalent technique for condensing a large-scale dataset into more compact versions while preserving their intrinsic utility. However, very few studies have investigated the trustworthiness of data distillation, i.e., privacy, robustness, and fairness. The deficiency is particularly striking given the existing research that underscores the vulnerabilities in current AI models, including privacy breaches, biased predictions against underrepresented subgroups, and susceptibility to imperceptible attacks. To bridge the gap, we propose a trustworthy benchmark for assessing representative dataset distillation solutions across the benchmark CIFAR10 with comprehensive evaluation metrics. Through extensive experiments, we uncover vulnerabilities inherent in the application of dataset distillation, offering valuable insights for practitioners. Our work aims to drive the development of more transparent, reliable, and responsible machine learning models, fostering AI systems that align with trustworthy principles.
Zongxiong Chen, Jiahui Geng, Derui Zhu, Qing Li 0038, Sonja Schimmler, Manfred Hauswirth
IJCNN1
2023 A Survey on Dataset Distillation: Approaches, Applications and Future Directions
abstract
Dataset distillation is attracting more attention in machine learning as training sets continue to grow and the cost of training state-of-the-art models becomes increasingly high. By synthesizing datasets with high information density, dataset distillation offers a range of potential applications, including support for continual learning, neural architecture search, and privacy protection. Despite recent advances, we lack a holistic understanding of the approaches and applications. Our survey aims to bridge this gap by first proposing a taxonomy of dataset distillation, characterizing existing approaches, and then systematically reviewing the data modalities, and related applications. In addition, we summarize the challenges and discuss future directions for this field of research.
Jiahui Geng, Zongxiong Chen, Yuandou Wang, Herbert Woisetschlaeger, Sonja Schimmler, Ruben Mayer, Zhiming Zhao, Chunming Rong
IJCAI2
2020 Grizzly: Efficient Stream Processing Through Adaptive Query Compilation
abstract
Stream Processing Engines (SPEs) execute long-running queries on unbounded data streams. They follow an interpretation-based processing model and do not perform runtime optimizations. This limits the utilization of modern hardware and neglects changing data characteristics at runtime. In this paper, we present Grizzly, a novel adaptive query compilation-based SPE, to enable highly efficient query execution. We extend query compilation and task-based parallelization for the unique requirements of stream processing and apply adaptive compilation to enable runtime re-optimizations. The combination of light-weight statistic gathering with just-in-time compilation enables Grizzly to adjust to changing data-characteristics dynamically at runtime. Our experiments show that Grizzly outperforms state-of-the-art SPEs by up to an order of magnitude in throughput.
Philipp M. Grulich, Sebastian Breß, Steffen Zeuch, Jonas Traub, Janis von Bleichert, Zongxiong Chen, Tilmann Rabl, Volker Markl
SIGMOD Conference6