Jiaxin Ding 0001

dblp:165/9711-1 · DBLP profile ↗
← Back
14ranked-venue papers in the field
4as first author
11since 2021 · last 2026
0000-0002-0009-9237ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 9 (4 first)Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2026 RARD: Rationale-First Blockwise Autoregressive Diffusion in Rationale?Dominated Graph Generation
abstract
Graph generation underlies many critical applications, from social network modeling to knowledge graph reasoning. Across these diverse domains, many graphs are rationale–dominated : a small, semantically meaningful subgraph determines the property of interest, while the remaining edges contribute largely noisy variation. Despite the significance of this inherent structure, existing generative methods often fail to preserve these task–critical substructures. We introduce RARD (Rationale-first blockwise AutoRegressive Diffusion), a topology-guided framework that learns to separate and prioritize the rationale. RARD employs a persistent-homology-based learning objective to discover an optimal graph filtration, an edge ordering that explicitly separates rationale from noise. Building upon this learned filtration, RARD generates graphs blockwise: it adds filtration-aligned blocks autoregressively and refines each new block with a shared local discrete diffusion module, ensuring the rationale appears early while peripheral structure is added later. We provide theoretical analysis showing that maximizing the topological gap yields rationale-first ordering and collapses to a two-level filtration. Comprehensive experiments across seven benchmarks demonstrate that RARD achieves state-of-the-art performance on widely used metrics.
Jiaxin Ding 0001, Luoyi Fu, Xinbing Wang
WWW4
2025 CELLM: Curvature Enhanced Large Language Models for Graph Structure Learning
Jiaxin Ding 0001, Haochen Han, Yule Xie, Luoyi Fu, Xinbing Wang
DASFAA (2)2
2025 Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge Graph
abstract
Large language models (LLMs) have demonstrated exceptional performance across a wide variety of domains. Nonetheless, generalist LLMs continue to fall short in reasoning tasks necessitating specialized knowledge, e.g., emotional sociology and medicine. Prior investigations into specialized LLMs focused on domain-specific training, which entails substantial efforts in domain data acquisition and model parameter fine-tuning. To address these challenges, this paper proposes the Way-to-Specialist (WTS) framework, which synergizes retrieval-augmented generation with knowledge graphs (KGs) to enhance the specialized capability of LLMs in the absence of specialized training. In distinction to existing paradigms that merely utilize external knowledge from general KGs or static domain KGs to prompt LLM for enhanced domain-specific reasoning, WTS proposes an innovative ''LLM↻KG'' paradigm, which achieves bidirectional enhancement between specialized LLM and domain knowledge graph (DKG). The proposed paradigm encompasses two closely coupled components: the DKG-Augmented LLM and the LLM-Assisted DKG Evolution. The former retrieves question-relevant domain knowledge from DKG and uses it to prompt LLM to enhance the reasoning capability for domain-specific tasks; the latter leverages LLM to generate new domain knowledge from processed tasks and use it to evolve DKG. WTS closes the loop between DKG-Augmented LLM and LLM-Assisted DKG Evolution, enabling continuous improvement in the domain specialization as it progressively answers and learns from domain-specific questions. We validate the performance of WTS on 7 datasets (e.g., TweetQA, ChatDoctor5k) spanning 6 domains, e.g., emotional sociology, medical, ect. The experimental results show that WTS surpasses the previous SOTA in 5 specialized domains, and achieves a maximum performance improvement of 11.3%.
Yutong Zhang 0003, Lixing Chen, Shenghong Li 0001, Nan Cao 0001, Yang Shi 0007, Jiaxin Ding 0001, Pan Zhou 0001, Yang Bai 0010
KDD (1)6
2025 Leveraging Homophily Under Local Differential Privacy for Effective Graph Neural Networks
Yule Xie, Jiaxin Ding 0001, Pengyu Xue, Haochen Han, Luoyi Fu, Xinbing Wang
ECML/PKDD (5)2
2025 DeepReport: An AI-assisted Idea Generation System for Scientific Research
abstract
Nowadays, the explosive growth of academic literature has been going far beyond scientists' limited capability to read through, making it increasingly difficult for them to absorb disciplinary insights and extract intellectual essences critical for generating novel research ideas in interdisciplinary studies. To address this, we develop DeepReport, an AI-assisted scientific idea generation system to alleviate the research burden. Technically, DeepReport maintains evolving concept co-occurrence graphs to extract core insights from over 260 million publications across all disciplines. These concepts are periodically collected and updated, enabling the automatic extraction of hidden cross-domain connections. Combining temporal link prediction and analysis techniques with large language models, DeepReport is able to further transform these patterns of insights into actionable ideas. With the function of integrating up-to-date academic databases, visualizing dynamic relationships of concepts, and automatically generating new ideas, DeepReport empowers researchers to navigate complex knowledge landscapes, reduce cognitive burdens, and accelerate the generation of groundbreaking concepts. This work provides an in-depth exploration of DeepReport's architecture, functionalities, and applications, highlighting its transformative potential for advancing interdisciplinary research and fostering innovation. DeepReport is available at https://idea.acemap.cn/.
Yi Xu 0004, Luoyi Fu, Shuqian Sheng, Jiaxin Ding 0001, Lei Zhou 0016, Xinbing Wang, Chenghu Zhou
SIGIR5
2024 Characterizing the Influence of Topology on Graph Learning Tasks
Kailong Wu, Yule Xie, Jiaxin Ding 0001, Yuxiang Ren, Luoyi Fu, Xinbing Wang, Chenghu Zhou
DASFAA (2)3
2024 Adversarial Reconstruction of Trajectories: Privacy Risks and Attack Models in Trajectory Embedding
abstract
Human trajectories, representing sequences of location points over time, are extensively collected and analyzed for various real-world applications such as urban planning, transportation management, and personalized location-based services. Trajectory embedding transforms raw trajectories into vector representations, capturing the underlying patterns and structures in the data. However, the abstraction provided by vector representations introduces significant security and privacy risks. These embeddings, often shared between entities or organizations, can be exploited by adversaries to reconstruct original trajectories, thereby compromising individual privacy. In this paper, we investigate the privacy issues of trajectory embeddings from an adversary's perspective. We propose two types of attacks to reconstruct original trajectories using road network information, addressing scenarios where the adversary has varying degrees of access to the black-box representation model. The first attack assumes unrestricted access to the model, allowing the adversary to construct a large-scale dataset and train a neural network to predict the road sequence of the trajectories. The second attack considers limited access, where the adversary computes distance coordinates between selected trajectory landmarks and road segments to infer different parts of the trajectory. Our experiments on a real-world dataset demonstrate that the reconstructed trajectories outperform baseline methods, achieving substantially lower reconstruction errors and more accurate alignment with the original trajectories, highlighting the significant vulnerability of trajectory embeddings to privacy breaches. These findings underscore the need for robust privacy-preserving mechanisms in spatio-temporal data analysis.
Haochen Han, Shuaiyu Yang, Jiaxin Ding 0001, Luoyi Fu, Xinbing Wang, Chenghu Zhou
SIGSPATIAL/GIS3
2023 Maximizing the Spread of Effective Information in Social Networks
abstract
Influence maximization through social networks has aroused tremendous interests nowadays. However, people’s various expressions or feelings about a same idea often cause ambiguity via word of mouth. Consequently, the problem of how to maximize the spread of “effective information” still remains largely open. In this paper, we consider a practical setting where ideas can deviate from their original version to invalid forms during message passing, and make the first attempt to seek a union of users that maximizes the spread of effective influence, which is formulated as an Influence Maximization with Information Variation (IMIV) problem. To this end, we model the information as a vector, and quantify the difference of two arbitrary vectors as a distance by a matching function. We further establish a process where such distance increases with the propagation and ensure the recipient whose vector distance is less than a threshold can be effectively influenced. Due to the NP-hardness of IMIV, we greedily select users that can approximately maximize the estimation of effective propagation. Especially, for networks of small scales, we derive a condition under which all the users can be effectively influenced. Our models and theoretical findings are further consolidated through extensive experiments on real-world datasets.
Haonan Zhang 0004, Luoyi Fu, Jiaxin Ding 0001, Feilong Tang 0001, Xinbing Wang, Guihai Chen, Chenghu Zhou
IEEE Trans. Knowl. Data Eng.3
2022 Analyzing sensitive information leakage in trajectory embedding models
abstract
With the proliferation of the mobile networks and location-based services, huge volume of user trajectories are collected to analyze the similarity among users and further unveil human mobility patterns for downstream tasks, such as point-of-interest recommendation and tourism planning. In recent works, trajectory embedding methods have been studied as efficient ways of trajectory similarity computation and effective inputs for downstream tasks, which embed trajectories into latent vector spaces equipped with the Euclidean distance to approximate the trajectory similarity and capture the characteristics of human mobility patterns. However, we demonstrate that such embedding, though hiding the locations, can leak the sensitive information of the trajectories, combined with auxiliary data. In this work, we propose trajectory embedding attack schemes to analyze the sensitive information leakage of the embedding vectors. In the experiment, we demonstrate that the passing areas, visited ROIs, and exact shapes of the trajectories are vulnerable under attacks on embedding vectors by the adversary with auxiliary information.
Jiaxin Ding 0001, Shichuan Xi, Kailong Wu, Xinbing Wang, Chenghu Zhou
SIGSPATIAL/GIS1
2022 TSNE: trajectory similarity network embedding
abstract
Trajectory representation learning studies the problem of embedding trajectories into low-dimensional vectors, while preserving mutual similarity for the convenience of downstream tasks, such as nearest neighbor search, clustering, classification, etc. In this work, we propose the Trajectory Similarity Network Embedding (TSNE) which exploits representation learning on the k-nearest neighbor partial similarity graph to generate trajectory embeddings, that preserve different similarity efficiently. In theory, we prove that TSNE is equivalent to factorizing the similarity graph, while in practice, TSNE achieves better performance. In the experiment, we show that TSNE outperforms the state-of-the-art baselines, including matrix factorization approaches and RNN based models in terms of similarity preserving and dimension reduction.
Jiaxin Ding 0001, Bowen Zhang 0005, Xinbing Wang, Chenghu Zhou
SIGSPATIAL/GIS1
2021 CLARA: A Constrained Reinforcement Learning Based Resource Allocation Framework for Network Slicing
abstract
As mobile networks proliferate, we are experiencing a strong diversification of services, which requires greater flexibility from the existing network. Network slicing is proposed as a promising solution for resource utilization in 5G and future networks to address this dire need. In network slicing, dynamic resource orchestration and network slice management are crucial for maximizing resource utilization. Unfortunately, this process is too complex for traditional approaches to be effective due to a lack of accurate models and dynamic hidden structures. We formulate the problem as a Constrained Markov Decision Process (CMDP) without knowing models and hidden structures. Additionally, we propose to solve the problem using CLARA, a Constrained reinforcement LeArning based Resource Allocation algorithm. In particular, we analyze cumulative and instantaneous constraints using adaptive interior-point policy optimization and projection layer, respectively. Evaluations show that CLARA clearly outperforms baselines in resource allocation with service demand guarantees.
Yongshuai Liu, Jiaxin Ding 0001, Zhi-Li Zhang, Xin Liu 0002
IEEE BigData2
2017 Fighting Statistical Re-Identification in Human Trajectory Publication
abstract
The maturing of mobile devices and systems provides an unprecedented opportunity to collect a large amount of real world human motion data at all scales. While the rich knowledge contained in these data sets is valuable in many fields, various types of personally sensitive information can be easily learned from such trajectory data. The ones that are of most concerns are frequent locations, frequent co-locations and trajectory re-identification through spatio-temporal data points. In this work we analyze privacy protection and data utility when trajectory IDs are randomly mixed during co-location events for data collection or publication. We demonstrate through both analyses and simulations that the global geometric shape of each individual trajectory is sufficiently altered such that re-identification via frequent locations, co-location pairs or spatial temporal data points is not possible with high probability. Meanwhile, a decent number of local geometric features of the trajectory data set are still preserved, including the density distribution and local traffic flow.
Jiaxin Ding 0001, Chien-Chun Ni, Jie Gao 0001
SIGSPATIAL/GIS1
2015 Understanding and modelling information dissemination patterns in vehicle-to-vehicle networks
abstract
Advances in wireless communication technology have enabled information exchange opportunities between moving vehicles within proximity. Potentially through such physical contacts a piece of information can diffuse to the entire network. While there has been extensive research on information diffusion in social networks, we do not know much about the spatial patterns in vehicle motion and how such patterns can support information dissemination. To this end, in this paper, we provide a systematic study of three large-scale data sets of taxi GPS traces from three big cities. The study shows the following properties universal of the three data sets: 1) the small world property, that information can be disseminated to almost the entire set of participants, within a very small number of hops; 2) certain physical contacts can be extremely effective in exchanging messages and such effectiveness shows a power law distribution; 3) the lack of hubs, no vehicle behaves as major hubs; removing top 20% nodes that have the highest number of physical contacts does not affect the effectiveness of information dissemination. 4) the information dissemination exhibits strong spatial temporal correlation. Finally, to explain the observations in particular the small world property, we develop mathematical models of the taxi movement patterns such that on graph topologies exhibiting properties of real-world road networks a number of observations can be rigorously proved.
Jiaxin Ding 0001, Jie Gao 0001, Hui Xiong 0001
SIGSPATIAL/GIS1
2015 Decentralized human trajectories tracking using hodge decomposition in sensor networks
abstract
With the recent development of localization and tracking systems for both indoor and outdoor settings, we consider the problem of analyzing and representing the huge amount of natural trajectories from human movements that we expect to gather in the near future. In this paper we argue the topological representation, which records how a target moves around the natural obstacles in the underlying environment, can be sufficiently descriptive for many applications and efficient enough for both storing, comparing and classifying these natural human trajectories. Technically, the representation uses the homotopy type of the trajectory. By using harmonic one-forms and Hodge decomposition, we pre-process the sensor network with a purely decentralized algorithm such that the homology class of a trajectory can be obtained by a simple integration along the trajectory. This supports real-time classification of trajectories up to the homology accuracy with minimum communication cost. We test the effectiveness of our approach by showing how to classify randomly generated trajectories in a multi-level arts museum layout as well as how to distinguish real world taxi trajectories in a large city.
Xiaotian Yin, Chien-Chun Ni, Jiaxin Ding 0001, Dengpan Zhou, Jie Gao 0001, Xianfeng Gu
SIGSPATIAL/GIS3