EDBT 2026 Demo / reviewers in the wild / expert
Haisong Zhang
dblp:220/2004
· DBLP profile ↗
24ranked-venue papers
2as first author
11since 2021 · last 2025
0009-0008-0567-3673ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A novel beamforming based on deconvolution for angular super-resolution
Haisong Zhang, Bingbing Qi |
Signal Process. | 1 |
| 2023 | On Social Network De-Anonymization With Communities: A Maximum A Posteriori PerspectiveabstractA crucial privacy-driven issue nowadays is re-identifying anonymized social networks by mapping them to correlated cross-domain auxiliary networks. Prior works are typically based on modeling social networks as random graphs representing users and their relations, and subsequently quantify the quality of mappings through varied cost functions. However, many cost functions are empirically proposed without sufficient theoretical support. For some other works probing the theoretical bound, it remains unknown how to algorithmically meet the demand of such quantifications, i.e., to minimize the cost functions. Besides, only few prior works have discussed the de-anonymization of social networks with communities. We address those concerns in a social network modeling parameterized by community structures that can be leveraged as side information for de-anonymization. Based on the Maximum A Posteriori (MAP) estimation, our first contribution is a series of MAP-based cost functions, which, when minimized, enjoy superiority to previous ones in finding the correct mapping with the highest probability. The feasibility of the cost functions is then for the first time algorithmically characterized. We prove the general multiplicative inapproximability and thus propose two heuristics, which, respectively, enjoy an$\epsilon$-additive approximation and a conditional optimality in carrying out successful user re-identification. Our theoretical findings are also empirically validated under classical synthetic and real-wrold social networks. Both theoretical and empirical observations manifest the importance of community in enhancing privacy inferencing. Jiapeng Zhang 0001, Shan Qu, Huquan Kang, Luoyi Fu, Haisong Zhang, Xinbing Wang, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Learning from Sibling Mentions with Scalable Graph Inference in Fine-Grained Entity TypingabstractYi Chen, Jiayang Cheng, Haiyun Jiang, Lemao Liu, Haisong Zhang, Shuming Shi, Ruifeng Xu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yi Chen 0019, Cheng Jiayang, Haiyun Jiang, Lemao Liu, Haisong Zhang, Shuming Shi 0001, Ruifeng Xu 0001 |
ACL (1) | 5 |
| 2022 | Make More Connections: Urban Traffic Flow Forecasting with Spatiotemporal Adaptive Gated Graph Convolution NetworkabstractUrban traffic flow forecasting is a critical issue in intelligent transportation systems. Due to the complexity and uncertainty of urban road conditions, how to capture the dynamic spatiotemporal correlation and make accurate predictions is very challenging. In most of existing works, urban road network is often modeled as a fixed graph based on local proximity. However, such modeling is not sufficient to describe the dynamics of the road network and capture the global contextual information. In this paper, we consider constructing the road network as a dynamic weighted graph through attention mechanism. Furthermore, we propose to seek both spatial neighbors and semantic neighbors to make more connections between road nodes. We propose a novel Spatiotemporal Adaptive Gated Graph Convolution Network ( STAG-GCN ) to predict traffic conditions for several time steps ahead. STAG-GCN mainly consists of two major components: (1) multivariate self-attention Temporal Convolution Network ( TCN ) is utilized to capture local and long-range temporal dependencies across recent, daily-periodic and weekly-periodic observations; (2) mix-hop AG-GCN extracts selective spatial and semantic dependencies within multi-layer stacking through adaptive graph gating mechanism and mix-hop propagation mechanism. The output of different components are weighted fused to generate the final prediction results. Extensive experiments on two real-world large scale urban traffic dataset have verified the effectiveness, and the multi-step forecasting performance of our proposed models outperforms the state-of-the-art baselines. Bin Lu 0005, Xiaoying Gan, Haiming Jin, Luoyi Fu, Xinbing Wang, Haisong Zhang |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2022 | Neighborhood Matters: Influence Maximization in Social Networks With Limited AccessabstractInfluence maximization (IM) aims at maximizing the spread of influence by offering discounts to influential users (called seeding). In many applications, due to user’s privacy concern, overwhelming network scale etc., it is hard to target any user in the network as one wishes. Instead, only a small subset of users is initially accessible. Such access limitation would significantly impair the influence spread, since IM often relies on seeding high degree users, which are particularly rare in such a small subset due to the power-law structure of social networks. In this paper, we attempt to solve the limited IM in real-world scenarios by the adaptive approach with seeding and diffusion uncertainty considered. Specifically, we consider fine-grained discounts and assume users accept the discount probabilistically. The diffusion process is depicted by the independent cascade model. To overcome the access limitation, we prove the set-wise friendship paradox (FP) phenomenon that neighbors have higher degree in expectation, and propose a two-stage seeding model with the FP embedded, where neighbors are seeded. On this basis, for comparison we formulate the non-adaptive case and adaptive case, both proven to be NP-hard. In the non-adaptive case, discounts are allocated to users all at once. We show the monotonicity of influence spread w.r.t. discount allocation and design a two-stage coordinate descent framework to decide the discount allocation. In the adaptive case, users are sequentially seeded based on observations of existing seeding and diffusion results. We prove the adaptive submodularity and submodularity of the influence spread function in two stages. Then, a series of adaptive greedy algorithms are proposed with constant approximation ratio. Extensive experiments on real-world datasets show that our adaptive algorithms achieve larger influence spread than non-adaptive and other adaptive algorithms (up to a maximum of 116 percent). Chen Feng 0007, Luoyi Fu, Bo Jiang 0003, Haisong Zhang, Xinbing Wang, Feilong Tang 0001, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | GAKG: A Multimodal Geoscience Academic Knowledge GraphabstractThe research of geoscience plays a strong role in helping people gain a better understanding of the Earth. To effectively represent the knowledge (KG) from enormous geoscience research papers, knowledge graphs can be a powerful means. In the face of enormous geoscience research papers, knowledge graphs can be a powerful means to manage the relationships of data and integrate knowledge extracted from them. However, the existing geoscience KGs mainly focus on the external connection between concepts, whereas the potential abundant information contained in the internal multimodal data of the paper is largely overlooked for more fine-grained knowledge mining. To this end, we propose GAKG, a large-scale multimodal academic KG based on 1.12 million papers published in various geoscience-related journals. In addition to the bibliometrics elements, we also extracted the internal illustrations, tables, and text information of the articles, and dig out the knowledge entities of the papers and the era and spatial attributes of the articles, coupling multimodal academic data and features. Specifically, GAKG realizes knowledge entity extraction under our proposed Human-In-the-Loop framework, the novelty of which is to combine the techniques of machine reading and information retrieval with manual annotation of geoscientists in the loop. Considering the fact that literature of geoscience often contains more abundant illustrations and time scale information compared with that of other disciplines, we extract all the geographical information and era from the geoscience papers' text and illustrations, mapping papers to the atlas and chronology. Based on GAKG, we build several knowledge discovery benchmarks for finding geoscience communities and predicting potential links. GAKG and its services have been made publicly available and user-friendly. Cheng Deng 0001, Yuting Jia, Hui Xu 0011, Luoyi Fu, Weinan Zhang 0001, Haisong Zhang, Xinbing Wang, Chenghu Zhou |
CIKM | 8 |
| 2021 | Fine-grained Entity Typing without Knowledge BaseabstractExisting work on Fine-grained Entity Typing (FET) typically trains automatic models on the datasets obtained by using Knowledge Bases (KB) as distant supervision.However, the reliance on KB means this training setting can be hampered by the lack of or the incompleteness of the KB.To alleviate this limitation, we propose a novel setting for training FET models: FET without accessing any knowledge base.Under this setting, we propose a two-step framework to train FET models.In the first step, we automatically create pseudo data with fine-grained labels from a large unlabeled dataset.Then a neural network model is trained based on the pseudo data, either in an unsupervised way or using self-training under the weak guidance from a coarse-grained Named Entity Recognition (NER) model.Experimental results show that our method achieves competitive performance with respect to the models trained on the original KB-supervised datasets.* The first two authors (Jing and Yibin) contributed equally to this work during the internships at Tencent AI Lab. Lemao Liu, Yangming Li, Haiyun Jiang, Haisong Zhang, Shuming Shi 0001 |
EMNLP (1) | 6 |
| 2021 | IGCN: Infected Graph Convolutional Network based Source IdentificationabstractSource identification has a wide range of applications in daily life, including locating the rumor source in online social networks and finding origins of a rolling blackout in smart grids. Despite great success over the past decade, most prior arts are proposed based an assumption that the underlying propagation model is known in advance. However, this assumption may be impracticable on real scenarios, since it is usually difficult to acquire the actual underlying propagation model. To avoid this limitation, in this paper, we propose the Infected Graph Convolutional Network (IGCN) layer by combining infection network with GCN (Graph Convolutional Network) layers to locate the rumor source without prior knowledge of underlying propagation model. For the first time, we define the problem of source identification as a special graph classification problem with source node as the label. By introducing the feature update method of GCN layer with the idea of attention, we build an IGCN model to adapt the infection networks such that the prediction accuracy on the source is improved under model independent scenarios. We conduct experiments on several real datasets and the results show the superiority of IGCN model to baseline algorlthms. Haisong Zhang, Luoyi Fu |
GLOBECOM | 3 |
| 2021 | Bag of Tricks for Chinese Named Entity RecognitionabstractNamed entity recognition (NER) is an important and challenging task in natural language processing. In this paper, we investigate thoroughly about the advances of Chinese NER in recent years. We explore the validity of a wide range of approaches in the literature of NLP that may benefit NER. We further employ the effective ones, such as data augmentation, adversarial learning, cross-sentence context and cost-sensitive learning to improve the performance of our BERT-based backbone model. Empirical results show that our model with this bag of tricks outperforms previous state-of-the-art on Weibo and achieves competitive performance on MSRA. Our code is publicly available11https://github.com/ccoay/bag-ner. Jingbo Peng, Luoyi Fu, Haisong Zhang |
IJCNN | 4 |
| 2021 | Getting Your Conversation on Track: Estimation of Residual Life for ConversationsabstractThis paper presents a predictive study on the progress of conversations. Specifically, we estimate the residual life for conversations, which is defined as the count of new turns to occur in a conversation thread. While most previous work focus on coarse-grained estimation that classifies the number of coming turns into two categories, we study fine-grained categorization for varying lengths of residual life. To this end, we propose a hierarchical neural model that jointly explores indicative representations from the content in turns and the structure of conversations in an end-to-end manner. Extensive experiments on both human-human and human-machine conversations demonstrate the superiority of our proposed model and its potential helpfulness in chatbot response selection. Jing Li 0049, Haisong Zhang |
SLT | 4 |
| 2021 | Conversational Semantic Role LabelingabstractSemantic role labeling (SRL) aims to extract the arguments for each predicate in an input sentence. Traditional SRL can fail to analyze dialogues because it only works on every single sentence, while ellipsis and anaphora frequently occur in dialogues. To address this problem, we propose the conversational SRL task, where an argument can be the dialogue participants, a phrase in the dialogue history or the current sentence. As the existing SRL datasets are in the sentence level, we manually annotate semantic roles for 3000 chit-chat dialogues (27198 sentences) to boost the research in this direction. Experiments show that while traditional SRL systems (even with the help of coreference resolution or rewriting) perform poorly for analyzing dialogues, modeling dialogue histories and participants greatly helps the performance, indicating that adapting SRL to conversations is very promising for universal dialogue understanding. Our initial study by applying CSRL to two mainstream conversational tasks, dialogue response generation and dialogue context rewriting, also confirms the usefulness of CSRL. Kun Xu 0005, Han Wu 0004, Linfeng Song, Haisong Zhang, Linqi Song, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | CASE: Context-Aware Semantic ExpansionabstractIn this paper, we define and study a new task called Context-Aware Semantic Expansion (CASE). Given a seed term in a sentential context, we aim to suggest other terms that well fit the context as the seed. CASE has many interesting applications such as query suggestion, computer-assisted writing, and word sense disambiguation, to name a few. Previous explorations, if any, only involve some similar tasks, and all require human annotations for evaluation. In this study, we demonstrate that annotations for this task can be harvested at scale from existing corpora, in a fully automatic manner. On a dataset of 1.8 million sentences thus derived, we propose a network architecture that encodes the context and seed term separately before suggesting alternative terms. The context encoder in this architecture can be easily extended by incorporating seed-aware attention. Our experiments demonstrate that competitive results are achieved with appropriate choices of context encoder and attention scoring function. Jialong Han, Aixin Sun, Haisong Zhang, Chenliang Li 0005, Shuming Shi 0001 |
AAAI | 3 |
| 2020 | Learning Sense Representation from Word Representation for Unsupervised Word Sense Disambiguation (Student Abstract)abstractUnsupervised WSD methods do not rely on annotated training datasets and can use WordNet. Since each ambiguous word in the WSD task exists in WordNet and each sense of the word has a gloss, we propose SGM and MGM to learn sense representations for words in WordNet using the glosses. In the WSD task, we calculate the similarity between each sense of the ambiguous word and its context to select the sense with the highest similarity. We evaluate our method on several benchmark WSD datasets and achieve better performance than the state-of-the-art unsupervised WSD systems. Zhenxin Fu, Moxin Li, Haisong Zhang, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 4 |
| 2020 | Draft and Edit: Automatic Storytelling Through Multi-Pass Hierarchical Conditional Variational AutoencoderabstractAutomatic Storytelling has consistently been a challenging area in the field of natural language processing. Despite considerable achievements have been made, the gap between automatically generated stories and human-written stories is still significant. Moreover, the limitations of existing automatic storytelling methods are obvious, e.g., the consistency of content, wording diversity. In this paper, we proposed a multi-pass hierarchical conditional variational autoencoder model to overcome the challenges and limitations in existing automatic storytelling models. While the conditional variational autoencoder (CVAE) model has been employed to generate diversified content, the hierarchical structure and multi-pass editing scheme allow the story to create more consistent content. We conduct extensive experiments on the ROCStories Dataset. The results verified the validity and effectiveness of our proposed model and yields substantial improvement over the existing state-of-the-art approaches. Meng-Hsuan Yu, Juntao Li 0005, Bo Tang 0016, Haisong Zhang, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 5 |
| 2020 | Rigid Formats Controlled Text GenerationabstractNeural text generation has made tremendous progress in various tasks.One common characteristic of most of the tasks is that the texts are not restricted to some rigid formats when generating.However, we may confront some special text paradigms such as Lyrics (assume the music score is given), Sonnet, SongCi (classical Chinese poetry of the Song dynasty), etc.The typical characteristics of these texts are in three folds: (1) They must comply fully with the rigid predefined formats.(2) They must obey some rhyming schemes.(3) Although they are restricted to some formats, the sentence integrity must be guaranteed.To the best of our knowledge, text generation based on the predefined rigid formats has not been well investigated.Therefore, we propose a simple and elegant framework named SongNet to tackle this problem.The backbone of the framework is a Transformer-based auto-regressive language model.Sets of symbols are tailor-designed to improve the modeling performance especially on format, rhyme, and sentence integrity.We improve the attention mechanism to impel the model to capture some future information on the format.A pre-training and fine-tuning framework is designed to further improve the generation quality.Extensive experiments conducted on two collected corpora demonstrate that our proposed framework generates significantly better results in terms of both automatic metrics and the human evaluation. 1 Piji Li, Haisong Zhang, Xiaojiang Liu, Shuming Shi 0001 |
ACL | 2 |
| 2020 | Hypernymy Detection for Low-Resource Languages via Meta LearningabstractHypernymy detection, a.k.a.lexical entailment, is a fundamental sub-task of many natural language understanding tasks.Previous explorations mostly focus on monolingual hypernymy detection on high-resource languages, e.g., English, but few investigate the lowresource scenarios.This paper addresses the problem of low-resource hypernymy detection by combining high-resource languages.We extensively compare three joint training paradigms and for the first time propose applying meta learning to relieve the low-resource issue.Experiments demonstrate the superiority of our method among the three settings, which substantially improves the performance of extremely low-resource languages by preventing over-fitting on small datasets.* Work done when C. Yu and J. Han were with Tencent AI Lab. Changlong Yu, Jialong Han, Haisong Zhang, Wilfred Ng |
ACL | 3 |
| 2020 | Spatiotemporal Adaptive Gated Graph Convolution Network for Urban Traffic Flow ForecastingabstractUrban traffic flow forecasting is a critical issue in intelligent transportation systems. It is quite challenging due to the complicated spatiotemporal dependency and essential uncertainty brought about by the dynamic urban traffic conditions. In most of existing methods, the spatial correlation is captured by utilizing graph neural networks (GNNs) throughout a fixed graph based on local spatial proximity. However, urban road conditions are complex and changeable, which leads to the interactions between roads should also be dynamic over time. In addition, the global contextual information of roads are also crucial for accurate forecasting. In this paper, we exploit spatiotemporal correlation of urban traffic flow and construct a dynamic weighted graph by seeking both spatial neighbors and semantic neighbors of road nodes. Multi-head self-attention temporal convolution network is utilized to capture local and long-range temporal dependencies across historical observations. Besides, we propose an adaptive graph gating mechanism to extract selective spatial dependencies within multi-layer stacking and correct information deviations caused by artificially defined spatial correlation. Extensive experiments on real world urban traffic dataset from Didi Chuxing GAIA Initiative have verified the effectiveness, and the multi-step forecasting performance of our proposed models outperforms the state-of-the-art baselines. The source code of our model is publicly available at https://github.com/RobinLu1209/STAG-GCN. Bin Lu 0005, Xiaoying Gan, Haiming Jin, Luoyi Fu, Haisong Zhang |
CIKM | 5 |
| 2020 | Continuity of Topic, Interaction, and Query: Learning to Quote in Online ConversationsabstractQuotations are crucial for successful explanations and persuasions in interpersonal communications.However, finding what to quote in a conversation is challenging for both humans and machines.This work studies automatic quotation generation in an online conversation and explores how language consistency affects whether a quotation fits the given context.Here, we capture the contextual consistency of a quotation in terms of latent topics, interactions with the dialogue history, and coherence to the query turn's existing content.Further, an encoder-decoder neural framework is employed to continue the context with a quotation via language generation.Experiment results on two large-scale datasets in English and Chinese demonstrate that our quotation generation model outperforms the state-of-the-art models.Further analysis shows that topic, interaction, and query consistency are all helpful to learn how to quote in online conversations. Lingzhi Wang 0001, Jing Li 0049, Xingshan Zeng, Haisong Zhang, Kam-Fai Wong |
EMNLP (1) | 4 |
| 2020 | Semantic Role Labeling Guided Multi-turn Dialogue ReWriterabstractFor multi-turn dialogue rewriting, the capacity of effectively modeling the linguistic knowledge in dialog context and getting rid of the noises is essential to improve its performance.Existing attentive models attend to all words without prior focus, which results in inaccurate concentration on some dispensable words.In this paper, we propose to use semantic role labeling (SRL), which highlights the core semantic information of who did what to whom, to provide additional guidance for the rewriter model.Experiments show that this information significantly improves a RoBERTa-based model that already outperforms previous stateof-the-art systems. Kun Xu 0005, Haochen Tan, Linfeng Song, Han Wu 0004, Haisong Zhang, Linqi Song, Dong Yu 0001 |
EMNLP (1) | 5 |
| 2019 | Coupling Global and Local Context for Unsupervised Aspect ExtractionabstractMing Liao, Jing Li, Haisong Zhang, Lingzhi Wang, Xixin Wu, Kam-Fai Wong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ming Liao, Jing Li 0049, Haisong Zhang, Lingzhi Wang 0001, Xixin Wu, Kam-Fai Wong |
EMNLP/IJCNLP (1) | 3 |
| 2018 | hyperdoc2vec: Distributed Representations of Hypertext DocumentsabstractHypertext documents, such as web pages and academic papers, are of great importance in delivering information in our daily life.Although being effective on plain documents, conventional text embedding methods suffer from information loss if directly adapted to hyper-documents.In this paper, we propose a general embedding approach for hyper-documents, namely, hyperdoc2vec, along with four criteria characterizing necessary information that hyper-document embedding models should preserve.Systematic comparisons are conducted between hyperdoc2vec and several competitors on two tasks, i.e., paper classification and citation recommendation, in the academic paper domain.Analyses and experiments both validate the superiority of hyperdoc2vec to other models w.r.t. the four criteria. Jialong Han, Yan Song 0003, Wayne Xin Zhao, Shuming Shi 0001, Haisong Zhang |
ACL (1) | 5 |
| 2018 | Generating Classical Chinese Poems via Conditional Variational Autoencoder and Adversarial TrainingabstractIt is a challenging task to automatically compose poems with not only fluent expressions but also aesthetic wording.Although much attention has been paid to this task and promising progress is made, there exist notable gaps between automatically generated ones with those created by humans, especially on the aspects of term novelty and thematic consistency.Towards filling the gap, in this paper, we propose a conditional variational autoencoder with adversarial training for classical Chinese poem generation, where the autoencoder part generates poems with novel terms and a discriminator is applied to adversarially learn their thematic consistency with their titles.Experimental results on a large poetry corpus confirm the validity and effectiveness of our model, where its automatic and human evaluation scores outperform existing models. Juntao Li 0005, Yan Song 0003, Haisong Zhang, Dongmin Chen, Shuming Shi 0001, Dongyan Zhao 0001, Rui Yan 0001 |
EMNLP | 3 |
| 2018 | A Hybrid Approach to Automatic Corpus Generation for Chinese Spelling CheckabstractChinese spelling check (CSC) is a challenging yet meaningful task, which not only serves as a preprocessing in many natural language processing (NLP) applications, but also facilitates reading and understanding of running texts in peoples' daily lives.However, to utilize datadriven approaches for CSC, there is one major limitation that annotated corpora are not enough in applying algorithms and building models.In this paper, we propose a novel approach of constructing CSC corpus with automatically generated spelling errors, which are either visually or phonologically resembled characters, corresponding to the OCRand ASR-based methods, respectively.Upon the constructed corpus, different models are trained and evaluated for CSC with respect to three standard test sets.Experimental results demonstrate the effectiveness of the corpus, therefore confirm the validity of our approach.* This work was conducted during Dingmin Wang's internship in Tencent AI Lab. SentenceCorrection Dingmin Wang, Yan Song 0003, Jing Li 0049, Jialong Han, Haisong Zhang |
EMNLP | 5 |
| 2018 | When Less Is More: Using Less Context Information to Generate Better Utterances in Group Conversations
Haisong Zhang, Zhangming Chan, Yan Song 0003, Dongyan Zhao 0001, Rui Yan 0001 |
NLPCC (1) | 1 |