EDBT 2026 Demo / reviewers in the wild / expert
Peng Zhang 0077
dblp:21/1048-77
· DBLP profile ↗
11ranked-venue papers
0as first author
8since 2021 · last 2023
0000-0001-5656-1083ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | GLM-130B: An Open Bilingual Pre-trained Model
Aohan Zeng, Xiao Liu 0036, Zhengxiao Du, Hanyu Lai, Ming Ding 0004, Zhuoyi Yang, Yifan Xu 0014, Wendi Zheng, Weng Lam Tam, Zixuan Ma, Jidong Zhai, Zhiyuan Liu 0001, Peng Zhang 0077, Yuxiao Dong, Jie Tang 0001 |
ICLR | 17 |
| 2023 | WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human PreferencesabstractWe present WebGLM, a web-enhanced question-answering system based on the General Language Model (GLM). Its goal is to augment a pre-trained large language model (LLM) with web search and retrieval capabilities while being efficient for real-world deployments. To achieve this, we develop WebGLM with strategies for the LLM-augmented retriever, bootstrapped generator, and human preference-aware scorer. Specifically, we identify and address the limitations of WebGPT (OpenAI), through which WebGLM is enabled with accuracy, efficiency, and cost-effectiveness advantages. In addition, we propose systematic criteria for evaluating web-enhanced QA systems. We conduct multi-dimensional human evaluation and quantitative ablation studies, which suggest the outperformance of the proposed WebGLM designs over existing systems. WebGLM with the 10-billion-parameter GLM (10B) is shown to perform better than the similar-sized WebGPT (13B) and even comparably to WebGPT (175B) in human evaluation. The code, demo, and data are at https://github.com/THUDM/WebGLM. Xiao Liu 0036, Hanyu Lai, Hao Yu 0030, Yifan Xu 0014, Aohan Zeng, Zhengxiao Du, Peng Zhang 0077, Yuxiao Dong, Jie Tang 0001 |
KDD | 7 |
| 2023 | CogDL: A Comprehensive Library for Graph Deep LearningabstractGraph neural networks (GNNs) have attracted tremendous attention from the graph learning community in recent years. It has been widely adopted in various real-world applications from diverse domains, such as social networks and biological graphs. The research and applications of graph deep learning present new challenges, including the sparse nature of graph data, complicated training of GNNs, and non-standard evaluation of graph tasks. To tackle the issues, we present CogDL1, a comprehensive library for graph deep learning that allows researchers and practitioners to conduct experiments, compare methods, and build applications with ease and efficiency. In CogDL, we propose a unified design for the training and evaluation of GNN models for various graph tasks, making it unique among existing graph learning libraries. By utilizing this unified trainer, CogDL can optimize the GNN training loop with several training techniques, such as mixed precision training. Moreover, we develop efficient sparse operators for CogDL, enabling it to become the most competitive graph library for efficiency. Another important CogDL feature is its focus on ease of use with the aim of facilitating open and reproducible research of graph learning. We leverage CogDL to report and maintain benchmark results on fundamental graph tasks, which can be reproduced and directly used by the community. Yukuo Cen, Yan Wang 0120, Yizhen Luo, Zhongming Yu, Xingcheng Yao, Aohan Zeng, Shiguang Guo, Yuxiao Dong, Yang Yang 0009, Peng Zhang 0077, Guohao Dai 0001, Yu Wang 0002, Chang Zhou 0005, Hongxia Yang, Jie Tang 0001 |
WWW | 13 |
| 2023 | GCCAD: Graph Contrastive Coding for Anomaly DetectionabstractGraph-based anomaly detection has been widely used for detecting malicious activities in real-world applications. Existing attempts to address this problem have thus far focused on structural feature engineering or learning in the binary classification regime. In this work, we propose to leverage graph contrastive learning and present the supervised GCCAD model for contrasting abnormal nodes with normal ones in terms of their distances to the global context (e.g., the average of all nodes). To handle scenarios with scarce labels, we further enable GCCAD as a self-supervised framework by designing a graph corrupting strategy for generating synthetic node labels. To achieve the contrastive objective, we design a graph neural network encoder that can infer and further remove suspicious links during message passing, as well as learn the global context of the input graph. We conduct extensive experiments on four public datasets, demonstrating that 1) GCCAD significantly and consistently outperforms various advanced baselines and 2) its self-supervised version without fine-tuning can achieve comparable performance with its fully supervised version. Bo Chen 0026, Jing Zhang 0001, Yuxiao Dong, Jian Song 0016, Peng Zhang 0077, Kaibo Xu, Evgeny Kharlamov, Jie Tang 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | OAG$_{\mathrm {know}}$ know : Self-Supervised Learning for Linking Knowledge GraphsabstractWe propose a self-supervised embedding learning frameworkSelfLinKGto link concepts in heterogeneous knowledge graphs. Without any labeled data, SelfLinKG can achieve competitive performance against its supervised counterpart, and significantly outperforms state-of-the-art unsupervised methods by 26%-50%. The essential components of SelfLinKG are local attention-based encoding and momentum contrastive learning. The former aims to learn the graph representation using an attention network, while the latter is to learn a self-supervised model across knowledge graphs using contrastive learning. SelfLinKG has been deployed to build the the new version, called OAG_know of Open Academic Graph (OAG). All data and codes are publicly available. Xiao Liu 0036, Li Mian, Yuxiao Dong, Fanjin Zhang, Jing Zhang 0001, Jie Tang 0001, Peng Zhang 0077, Jibing Gong, Kuansan Wang |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | CODE: Contrastive Pre-training with Adversarial Fine-Tuning for Zero-Shot Expert LinkingabstractExpert finding, a popular service provided by many online websites such as Expertise Finder, LinkedIn, and AMiner, is beneficial to seeking candidate qualifications, consultants, and collaborators. However, its quality is suffered from lack of ample sources of expert information. This paper employs AMiner as the basis with an aim at linking any external experts to the counterparts on AMiner. As it is infeasible to acquire sufficient linkages from arbitrary external sources, we explore the problem of zero-shot expert linking. In this paper, we propose CODE, which first pre-trains an expert linking model by contrastive learning on AMiner such that it can capture the representation and matching patterns of experts without supervised signals, then it is fine-tuned between AMinerand external sources to enhance the model’s transferability in an adversarial manner. For evaluation, we first design two intrinsic tasks, author identification and paper clustering, to validate the representation and matching capability endowed by contrastive learning. Then the final external expert linking performance on two genres of external sources also implies the superiority of adversarial fine-tuning method. Additionally, we show the online deployment of CODE, and continuously improve its online performance via active learning. Bo Chen 0026, Jing Zhang 0001, Xiaobin Tang, Lingfan Cai, Hong Chen 0001, Cuiping Li 0001, Peng Zhang 0077, Jie Tang 0001 |
AAAI | 8 |
| 2022 | OAG-BERT: Towards a Unified Backbone Language Model for Academic Knowledge ServicesabstractAcademic Knowledge Services have substantially facilitated the development of human science and technology, providing a plenitude of useful research tools. However, many applications highly depend on ad-hoc models and expensive human labeling to understand professional contents, hindering deployments in real world. To create a unified backbone language model for various knowledge-intensive academic knowledge mining challenges, based on the world's largest public academic graph Open Academic Graph (OAG), we pre-train an academic language model, namely OAG-BERT, to integrate massive heterogeneous entity knowledge beyond scientific corpora. We develop novel pre-training strategies along with zero-shot inference techniques. OAG-BERT's superior performance on 9 knowledge-intensive academic tasks (including 2 demo applications) demonstrates its qualification to serve as a foundation for academic knowledge services. Its zero-shot capability also offers great potential to mitigate the need of costly annotations. OAG-BERT has been deployed to multiple real-world applications, such as reviewer recommendations for NSFC (National Nature Science Foundation of China) and paper tagging in the AMiner system. All codes and pre-trained models are available via the CogDL. Xiao Liu 0036, Da Yin, Jingnan Zheng, Xingjian Zhang 0009, Peng Zhang 0077, Hongxia Yang, Yuxiao Dong, Jie Tang 0001 |
KDD | 5 |
| 2022 | Taming the Big Data Monster: Managing Petabytes of Data with Multi-Model DatabasesabstractWith the development of big data technology, the amount of business data that Internet companies need to handle has reached the petabyte level, which poses great pressure on the system processing capacity. For example, the peak order volume of Alibaba's Global Shopping Festival in 2020 reached 583,000 orders per second. Even worse, multi-model data are involved in real business. The inability to perform high-throughput, lowlatency transaction processing can result in a poor user experience that can lead to serious financial losses due to customer churn. Although numerous optimizations have been proposed, they can fail in the face of petabytes of data, or be significantly less effective. In this paper, we propose a novel and practical multi-model big data system that can manage petabytes of data. Particularly, we show three special designs for processing the petabytes of data. First, we perform partition to reduce the amount of unnecessary data to be scanned. Second, we adaptively adopt row storage mode for big tables that are frequently updated and column storage mode for tables that are frequently queried to improve the system efficiency. Third, we conduct compression to accelerate IO access speed. We analyze Alibaba's two real PB-level business scenarios, Double 11 and Zhixingtong, and generate workloads and benchmark accordingly to verify our system. Experiments show that our system can efficiently manage petabyte-scale data in real scenarios, providing high-performance querying of terabyte-scale datasets, and be suitable for various workloads. Feng Zhang 0007, Yinhao Hong, Yunpeng Chai, Wei Lu 0015, Hong Chen 0001, Xiaoyong Du 0001, Le Mi, Xilin Tang, Yanliang Zhou, Peng Zhang 0077, Fengyi Chen |
SBAC-PAD | 14 |
| 2019 | Domain Adaptive Question Answering over Knowledge Base
Yulai Yang, Lei Hou 0001, Hailong Jin, Peng Zhang 0077, Juan-Zi Li, Yi Huang 0017 |
NLPCC (2) | 4 |
| 2015 | NewsMiner: Multifaceted news analysis for event search
Lei Hou 0001, Juan-Zi Li, Zhichun Wang, Jie Tang 0001, Peng Zhang 0077, Ruibing Yang |
Knowl. Based Syst. | 5 |
| 2014 | On Modelling Non-linear Topical DependenciesabstractProbabilistic topic models such as Latent Dirichlet Allocation (LDA) discover latent topics from large corpora by exploiting words’ co-occurring relation. By observing the topical similarity between words, we find that some other relations, such as semantic or syntax relation between words, lead to strong dependence between their topics. In this paper, sentences are represented as dependency trees and a Global Topic Random Field (GTRF) is presented to model the non-linear dependencies between words. To infer our model, a new global factor is defined over all edges and the normalization factor of GRF is proven to be a constant. As a result, no independent assumption is needed when inferring our model. Based on it, we develop an efficient expectation-maximization (EM) procedure for parameter estimation. Experimental results on four data sets show that GTRF achieves much lower perplexity than LDA and linear dependency topic models and produces better topic coherence. Siqiang Wen, Juan-Zi Li, Peng Zhang 0077, Jie Tang 0001 |
ICML | 4 |