Zian Jia

dblp:330/8167 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Graph learning · 48% Language models and text generation · 28% Knowledge representation and reasoning · 24%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 90% Data mining · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Network and information security
1 paper
Network security · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector Perspective · WWW 2026
Information retrieval › retrieval-augmented generation
context compression
1.012026
Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector Perspective · WWW 2026
Information retrieval
retrieval-augmented generation
1.012026
Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector Perspective · WWW 2026
Machine learning › Graph learning › graph neural network
dynamic graph neural network
0.912025
Unifying Text Semantics and Graph Structures for Temporal Text-attributed Graphs with Large Language Models · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
temporal knowledge graph
0.912025
Unifying Text Semantics and Graph Structures for Temporal Text-attributed Graphs with Large Language Models · NeurIPS 2025
Machine learning › Graph learning
text-attributed graph
0.912025
Unifying Text Semantics and Graph Structures for Temporal Text-attributed Graphs with Large Language Models · NeurIPS 2025
Computational science and engineering › inverse problem
inverse design
0.912025
MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial Discovery · KDD (2) 2025
Computational science and engineering
materials informatics
0.912025
MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial Discovery · KDD (2) 2025
Visualization and visual analytics › interactive visualization
interactive visual interfaces
0.912025
MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial Discovery · KDD (2) 2025
Network security › intrusion detection and prevention › intrusion detection › attack detection
advanced persistent threat detection
0.812024
MAGIC: Detecting Advanced Persistent Threats via Masked Graph Representation Learning · USENIX Security Symposium 2024
Network security › intrusion detection and prevention
intrusion detection
0.812024
MAGIC: Detecting Advanced Persistent Threats via Masked Graph Representation Learning · USENIX Security Symposium 2024
Data mining › representation learning
graph representation learning
0.212024
MAGIC: Detecting Advanced Persistent Threats via Masked Graph Representation Learning · USENIX Security Symposium 2024

Methods — techniques the papers use, named apart from their topics

soft compression · 2.0query-conditioned selection · 2.0machine learning benchmarking · 1.7finite element analysis · 1.7masked graph representation learning · 1.5semantic-structural co-encoding · 0.9large language model · 0.9
YearPublicationVenuePosition
2026 Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector Perspective
abstract
Retrieval-Augmented Generation (RAG) effectively grounds Large Language Models (LLMs) with external knowledge and is widely applied to Web-related tasks. However, its scalability is hindered by excessive context length and redundant retrievals. Recent research on soft context compression aims to address this by encoding long documents into compact embeddings, yet they often underperform non-compressed RAG due to their reliance on auto-encoder-like full-compression that forces the encoder to compress all document information regardless of relevance to the input query.
Zian Jia, Kanjun Xu, Yun Xiong
WWW2
2025 MetamatBench: Integrating Heterogeneous Data, Computational Tools, and Visual Interface for Metamaterial Discovery
abstract
Metamaterials, engineered materials with architected structures across multiple length scales, offer unprecedented and tunable mechanical properties that surpass those of conventional materials. However, leveraging advanced machine learning (ML) for metamaterial discovery is hindered by three fundamental challenges: (C1) Data Heterogeneity Challenge arises from heterogeneous data sources, heterogeneous composition scales, and heterogeneous structure categories; (C2) Model Complexity Challenge stems from the intricate geometric constraints of ML models, which complicate their adaptation to metamaterial structures; and (C3) Human-AI Collaboration Challenge comes from the ''dual black-box'' nature of sophisticated ML models and the need for intuitive user interfaces. To tackle these challenges, we introduce a unified framework, named MetamatBench, that operates on three levels. (1) At the data level, we integrate and standardize 5 heterogeneous, multi-modal metamaterial datasets. (2) The ML level provides a comprehensive toolkit that adapts 17 state-of-the-art ML methods for metamaterial discovery. It also includes a comprehensive evaluation suite with 12 novel performance metrics plus a finite element-based assessment to ensure accurate and reliable model validation. (3) The user level features a visual-interactive interface that bridges the gap between complex ML techniques and non-ML researchers, advancing property prediction and inverse design of metamaterials for research and applications. MetamatBench offers a unified platform that enables machine learning researchers and practitioners to develop and evaluate new methodologies in metamaterial discovery. For accessibility and reproducibility, we open-source our benchmark and the codebase at https://github.com/cjpcool/Metamaterial-Benchmark.
Jianpeng Chen, Wangzhi Zhan, Haohui Wang, Zian Jia, Jingru Gan, Jingyuan Qi, Lifu Huang, Muhao Chen 0001, Wei Wang 0010, Dawei Zhou 0003
KDD (2)4
2025 Unifying Text Semantics and Graph Structures for Temporal Text-attributed Graphs with Large Language Models
abstract
Temporal graph neural networks (TGNNs) have shown remarkable performance in temporal graph modeling. However, real-world temporal graphs often possess rich textual information, giving rise to temporal text-attributed graphs (TTAGs). Such combination of dynamic text semantics and evolving graph structures introduces heightened complexity. Existing TGNNs embed texts statically and rely heavily on encoding mechanisms that biasedly prioritize structural information, overlooking the temporal evolution of text semantics and the essential interplay between semantics and structures for synergistic reinforcement. To tackle these issues, we present $\textbf{CROSS}$, a flexible framework that seamlessly extends existing TGNNs for TTAG modeling. CROSS is designed by decomposing the TTAG modeling process into two phases: (i) temporal semantics extraction; and (ii) semantic-structural information unification. The key idea is to advance the large language models (LLMs) to $\textit{dynamically}$ extract the temporal semantics in text space and then generate $\textit{cohesive}$ representations unifying both semantics and structures. Specifically, we propose a Temporal Semantics Extractor in the CROSS framework, which empowers LLMs to offer the temporal semantic understanding of node's evolving contexts of textual neighborhoods, facilitating semantic dynamics. Subsequently, we introduce the Semantic-structural Co-encoder, which collaborates with the above Extractor for synthesizing illuminating representations by jointly considering both semantic and structural information while encouraging their mutual reinforcement. Extensive experiments show that CROSS achieves state-of-the-art results on four public datasets and one industrial dataset, with 24.7\% absolute MRR gain on average in temporal link prediction and 3.7\% AUC gain in node classification of industrial application.
Siwei Zhang 0001, Yun Xiong, Yateng Tang, Jiarong Xu, Xi Chen 0072, Zehao Gu, Xuehao Zheng, Zian Jia, Jiawei Zhang 0001
NeurIPS8
2024 MAGIC: Detecting Advanced Persistent Threats via Masked Graph Representation Learning
Zian Jia, Yun Xiong, Yuhong Nan, Yao Zhang 0009, Jinjing Zhao, Mi Wen
USENIX Security Symposium1