VLDB 2026 Research / reviewers in the wild / expert
Dongcheng Zou
dblp:342/8197
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-0634-3720ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Structural Information Guided Hierarchical Reconstruction for Graph Anomaly DetectionabstractAnomalies in graphs involve attributes and structures and may occur at different levels (e.g., node or community). Existing GNN-based detection methods often merely focus on anomalies of single nodes or neighborhoods, making it hard to cope with complex and organized networks. Towards this, we propose SI-HGAD, a novel Graph Anomaly Detection (GAD) approach that utilizes hierarchical information to detect anomalies. Powered by structural information, SI-HGAD can mine an optimal graph abstraction while enabling hierarchical substructural modeling. Also, we design a Graph Transformer to mine multi-range structural and attribute patterns for nodes. The decoders reconstruct both the node attributes and the multi-level subgraphs in a bottom-up manner. Extensive experiments demonstrate the superiority of SI-HGAD. Dongcheng Zou, Hao Peng 0001 |
CIKM | 1 |
| 2024 | MultiSPANS: A Multi-range Spatial-Temporal Transformer Network for Traffic Forecast via Structural Entropy OptimizationabstractTraffic forecasting is a complex multivariate time-series regression task of paramount importance for traffic management and planning. However, existing approaches often struggle to model complex multi-range dependencies using local spatiotemporal features and road network hierarchical knowledge. To address this, we propose MultiSPANS. First, considering that an individual recording point cannot reflect critical spatiotemporal local patterns, we design multi-filter convolution modules for generating informative ST-token embeddings to facilitate attention computation. Then, based on ST-token and spatial-temporal position encoding, we employ the Transformers to capture long-range temporal and spatial dependencies. Furthermore, we introduce structural entropy theory to optimize the spatial attention mechanism. Specifically, The structural entropy minimization algorithm is used to generate optimal road network hierarchies, i.e., encoding trees. Based on this, we propose a relative structural entropy-based position encoding and a multi-head attention masking scheme based on multi-layer encoding trees. Extensive experiments demonstrate the superiority of the presented framework over several state-of-the-art methods in real-world traffic datasets, and the longer historical windows are effectively utilized. The code is available at https://github.com/SELGroup/MultiSPANS. Dongcheng Zou, Senzhang Wang, Xuefeng Li 0003, Hao Peng 0001, Yuandong Wang 0002, Kehua Sheng, Bo Zhang 0106 |
WSDM | 1 |
| 2024 | CoSENT: Consistent Sentence Embedding via Similarity RankingabstractLearning the representation of sentences is fundamental work in the field of Natural Language Processing. Although BERT-like transformers have achieved new SOTAs for sentence embedding in many tasks, they have been proven difficult to capture semantic similarity without proper fine-tuning. A common idea to measure Semantic Textual Similarity (STS) is considering the distance between two text embeddings defined by the dot product or cosine function. However, the semantic embedding spaces induced by pretrained transformers are generally non-smooth and tend to deviate from a normal distribution, which makes traditional distance metrics imprecise. In this paper, we first empirically explain the failure of cosine similarity in semantic textual similarity measuring, and present CoSENT, a novelConsistentSENTence embedding framework. Concretely, a supervised objective function is designed to optimize the Siamese BERT network by exploiting ranked similarity labels of sample pairs. The loss function utilizes uniform cosine similarity-based optimization for both the training and prediction phases, improving the consistency of the learned semantic space. Additionally, the unified objective function can be adaptively applied to different datasets with various types of annotations and different comparison schemes of the STS tasks only by using sortable labels. Empirical evaluations on 14 common textual similarity benchmarks demonstrate that the proposed CoSENT excels in performance and reduces training time cost. Hao Peng 0001, Dongcheng Zou, Zhiwei Liu 0001, Jianxin Li 0002, Kay Liu, Jia Wu 0001, Jianlin Su, Philip S. Yu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | SE-GSL: A General and Effective Graph Structure Learning Framework through Structural Entropy OptimizationabstractGraph Neural Networks (GNNs) are de facto solutions to structural data learning. However, it is susceptible to low-quality and unreliable structure, which has been a norm rather than an exception in real-world graphs. Existing graph structure learning (GSL) frameworks still lack robustness and interpretability. This paper proposes a general GSL framework, SE-GSL, through structural entropy and the graph hierarchy abstracted in the encoding tree. Particularly, we exploit the one-dimensional structural entropy to maximize embedded information content when auxiliary neighbourhood attributes is fused to enhance the original graph. A new scheme of constructing optimal encoding trees are proposed to minimize the uncertainty and noises in the graph whilst assuring proper community partition in hierarchical abstraction. We present a novel sample-based mechanism for restoring the graph structure via node structural entropy distribution. It increases the connectivity among nodes with larger uncertainty in lower-level communities. SE-GSL is compatible with various GNN models and enhances the robustness towards noisy and heterophily structures. Extensive experiments show significant improvements in the effectiveness and robustness of structure learning and node representation learning. Dongcheng Zou, Hao Peng 0001, Renyu Yang, Jianxin Li 0002, Jia Wu 0001, Philip S. Yu |
WWW | 1 |