EDBT 2026 Demo / reviewers in the wild / expert
Yafang Wang
dblp:68/6229
· DBLP profile ↗
26ranked-venue papers
5as first author
2since 2021 · last 2021
0000-0003-0158-6210ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 14 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Question answering and dialogue systems · 44% Deep learning architectures and training · 20% Reinforcement learning · 14% | |
| Databases, data mining, and information retrieval
3 papers |
Knowledge graphs · 39% Query processing and optimization · 29% Graph data management · 29% | |
| Software engineering, system software, and programming languages
3 papers |
Programming languages and type systems · 38% Services computing and microservices · 33% Operating systems · 29% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › language modeling › language model architecture
hierarchical language model |
0.5 | 1 | 2021 | R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language Modeling · ACL/IJCNLP (1) 2021 |
Machine learning › Deep learning architectures and training › transformer
recursive transformer |
0.5 | 1 | 2021 | R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language Modeling · ACL/IJCNLP (1) 2021 |
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue policy learning |
0.4 | 1 | 2020 | Two-stage Behavior Cloning for Spoken Dialogue System in Debt Collection · IJCAI 2020 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.4 | 1 | 2020 | Long Short-Term Sample Distillation · AAAI 2020 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.4 | 1 | 2020 | Two-stage Behavior Cloning for Spoken Dialogue System in Debt Collection · IJCAI 2020 |
Machine learning › Deep learning architectures and training
teacher-student framework |
0.4 | 1 | 2020 | Long Short-Term Sample Distillation · AAAI 2020 |
Natural language and speech › Question answering and dialogue systems › knowledge base question answering
complex question answering |
0.4 | 1 | 2019 | Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs · SIGIR 2019 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.4 | 1 | 2019 | CRSRL: Customer Routing System Using Reinforcement Learning · IJCAI 2019 |
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering
multi-document question answering |
0.4 | 1 | 2019 | Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs · SIGIR 2019 |
Graph data management › graph algorithms
group steiner tree |
0.4 | 1 | 2019 | Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs · SIGIR 2019 |
Query processing and optimization
similarity join |
0.4 | 1 | 2019 | Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge Graphs · SIGIR 2019 |
Operating systems › resource management
resource allocation |
0.4 | 1 | 2019 | CRSRL: Customer Routing System Using Reinforcement Learning · IJCAI 2019 |
Natural language and speech › Information extraction and text analysis › named entity processing
named entity recognition and disambiguation |
0.2 | 1 | 2013 | YaLi: a crowdsourcing plug-in for NERD · SIGIR 2013 |
Machine learning › Reinforcement learning › imitation learning › offline imitation learning
behavior cloning |
0.1 | 1 | 2020 | Two-stage Behavior Cloning for Spoken Dialogue System in Debt Collection · IJCAI 2020 |
Machine learning › Reinforcement learning
imitation learning |
0.1 | 1 | 2020 | Two-stage Behavior Cloning for Spoken Dialogue System in Debt Collection · IJCAI 2020 |
Knowledge graphs › knowledge graph construction
knowledge extraction |
0.1 | 1 | 2020 | ServiceGroup: A Human-Machine Cooperation Solution for Group Chat Customer Service · SIGIR 2020 |
Web and social media mining › web mining
web page understanding |
0.0 | 1 | 2013 | YaLi: a crowdsourcing plug-in for NERD · SIGIR 2013 |
Methods — techniques the papers use, named apart from their topics
human-machine cooperation · 1.3recursive transformer · 1.0differentiable tree · 1.0similarity join · 0.8group steiner tree · 0.8deep reinforcement learning · 0.8teacher-student training · 0.4multi-label classification · 0.4knowledge distillation · 0.4behavior cloning · 0.4implicit feedback · 0.2crowdsourcing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language ModelingabstractXiang Hu, Haitao Mi, Zujie Wen, Yafang Wang, Yi Su, Jing Zheng, Gerard de Melo. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Haitao Mi, Zujie Wen, Yafang Wang, Gerard de Melo |
ACL/IJCNLP (1) | 4 |
| 2021 | Incorporating Specific Knowledge into End-to-End Task-oriented Dialogue SystemsabstractExternal knowledge is vital to many natural language processing tasks. However, current end-to-end dialogue systems often struggle to interface knowledge bases(KBs) with response smoothly and effectively. In this paper, we convert the raw knowledge into relation knowledge and integrated knowledge and then incorporate them into end-to-end task-oriented dialogue systems. The relation knowledge extracted from knowledge triples is combined with dialogue history, aiming to enhance semantic inputs and support better language understanding. Integrated knowledge involves entities and relations by graph attention, assisting the model in generating informative responses. The experimental results on three public dialogue datasets show that our model improves over the previous state-of-the-art models in sentence fluency and informativeness. Qingyue Wang, Yanan Cao 0001, Junyan Jiang, Yafang Wang, Lingling Tong, Li Guo 0001 |
IJCNN | 4 |
| 2020 | Long Short-Term Sample DistillationabstractIn the past decade, there has been substantial progress at training increasingly deep neural networks. Recent advances within the teacher–student training paradigm have established that information about past training updates show promise as a source of guidance during subsequent training steps. Based on this notion, in this paper, we propose Long Short-Term Sample Distillation, a novel training policy that simultaneously leverages multiple phases of the previous training process to guide the later training updates to a neural network, while efficiently proceeding in just one single generation pass. With Long Short-Term Sample Distillation, the supervision signal for each sample is decomposed into two parts: a long-term signal and a short-term one. The long-term teacher draws on snapshots from several epochs ago in order to provide steadfast guidance and to guarantee teacher–student differences, while the short-term one yields more up-to-date cues with the goal of enabling higher-quality updates. Moreover, the teachers for each sample are unique, such that, overall, the model learns from a very diverse set of teachers. Comprehensive experimental results across a range of vision and NLP tasks demonstrate the effectiveness of this new training method. Zujie Wen, Zhongping Liang, Yafang Wang, Gerard de Melo, Zhe Li 0007, Liangzhuang Ma, Xiaolong Li 0005, Yuan Qi 0001 |
AAAI | 4 |
| 2020 | Data Augmentation for Multiclass Utterance Classification - A Systematic StudyabstractUtterance classification is a key component in many conversational systems. However, classifying real-world user utterances is challenging, as people may express their ideas and thoughts in manifold ways, and the amount of training data for some categories may be fairly limited, resulting in imbalanced data distributions. To alleviate these issues, we conduct a comprehensive survey regarding data augmentation approaches for text classification, including simple random resampling, word-level transformations, and neural text generation to cope with imbalanced data. Our experiments focus on multi-class datasets with a large number of data samples, which has not been systematically studied in previous work. The results show that the effectiveness of different data augmentation schemes depends on the nature of the dataset under consideration. Binxia Xu, Siyuan Qiu, Jie Zhang 0060, Yafang Wang, Gerard de Melo |
COLING | 4 |
| 2020 | DAN: Dual-View Representation Learning for Adapting Stance Classifiers to New DomainsabstractWe address the issue of having a limited number of annotations for stance classification in a new domain, by adapting out-of-domain classifiers with domain adaptation. Existing approaches often align different domains in a single, global feature space (or view), which may fail to fully capture the richness of the languages used for expressing stances, leading to reduced adaptability on stance data. In this paper, we identify two major types of stance expressions that are linguistically distinct, and we propose a tailored dual-view adaptation network (DAN) to adapt these expressions across domains. The proposed model first learns a separate view for domain transfer in each expression channel and then selects the best adapted parts of both views for optimal transfer. We find that the learned view features can be more easily aligned and more stance-discriminative in either or both views, leading to more transferable overall features after combining the views. Results from extensive experiments show that our method can enhance the state-of-the-art single-view methods in matching stance data across different domains, and that it consistently improves those methods on various adaptation tasks. Chang Xu 0002, Cécile Paris, Surya Nepal, Ross Sparks, Chong Long, Yafang Wang |
ECAI | 6 |
| 2020 | Two-stage Behavior Cloning for Spoken Dialogue System in Debt CollectionabstractWith the rapid growth of internet finance and the booming of financial lending, the intelligent calling for debt collection in FinTech companies has driven increasing attention. Nowadays, the widely used intelligent calling system is based on dialogue flow, namely configuring the interaction flow with the finite-state machine. In our scenario of debt collection, the completed dialogue flow contains more than one thousand interactive paths. All the dialogue procedures are artificially specified, with extremely high maintenance costs and error-prone. To solve this problem, we propose the behavior-cloning-based collection robot framework without any dialogue flow configuration, called two-stage behavior cloning (TSBC). In the first stage, we use multi-label classification model to obtain policies that may be able to cope with the current situation according to the dialogue state; in the second stage, we score several scripts under each obtained policy to select the script with the highest score as the reply for the current state. This framework makes full use of the massive manual collection records without labeling and fully absorbs artificial wisdom and experience. We have conducted extensive experiments in both single-round and multi-round scenarios and showed the effectiveness of the proposed system. The accuracy of a single round of dialogue can be improved by 5%, and the accuracy of multiple rounds of dialogue can be increased by 3.1%. Hengbin Cui, Chunxiang Jin, Yafang Wang, Xiaolong Li 0005, Renxin Mao |
IJCAI | 6 |
| 2020 | ServiceGroup: A Human-Machine Cooperation Solution for Group Chat Customer ServiceabstractWith the rapid growth of B2B (Business-to-Business), how to efficiently respond to various customer questions is becoming an important issue. In this scenario, customer questions always involve many aspects of the products, so there are usually multiple customer service agents to response respectively. To improve efficiency, we propose a human-machine cooperation solution called ServiceGroup, where relevant agents and customers are invited into the same group, and the system can provide a series of intelligent functions, including question notification, question recommendation and knowledge extraction. With the assistance of our developed ServiceGroup, the response rate within 15 minutes is improved twice. Until now, our ServiceGroup has already supported thousands of enterprises by means of millions of groups in instant messaging softwares. Hengbin Cui, Shaosheng Cao, Yafang Wang, Xiaolong Li 0005 |
SIGIR | 4 |
| 2020 | Optimization of real-time traffic network assignment based on IoT data using DBN and clustering model in smart city
Yurong Han, Yafang Wang, Bin Jiang 0003, Zhihan Lyu, Houbing Song |
Future Gener. Comput. Syst. | 3 |
| 2020 | Multi-document semantic relation extraction for news analytics
Yongpan Sheng, Zenglin Xu, Yafang Wang, Gerard de Melo |
World Wide Web | 3 |
| 2019 | CRSRL: Customer Routing System Using Reinforcement LearningabstractAllocating resources to customers in the customer service is a difficult problem, because designing an optimal strategy to achieve an optimal trade-off between available resources and customers' satisfaction is non-trivial. In this paper, we formalize the customer routing problem, and propose a novel framework based on deep reinforcement learning (RL) to address this problem. To make it more practical, a demo is provided to show and compare different models, which visualizes all decision process, and in particular, the system shows how the optimal strategy is reached. Besides, our demo system also ships with a variety of models that users can choose based on their needs. Chong Long, Zining Liu, Xiaolu Lu 0002, Zehong Hu, Yafang Wang |
IJCAI | 5 |
| 2019 | Answering Complex Questions by Joining Multi-Document Evidence with Quasi Knowledge GraphsabstractDirect answering of questions that involve multiple entities and relations is a challenge for text-based QA. This problem is most pronounced when answers can be found only by joining evidence from multiple documents. Curated knowledge graphs (KGs) may yield good answers, but are limited by their inherent incompleteness and potential staleness. This paper presents QUEST, a method that can answer complex questions directly from textual sources on-the-fly, by computing similarity joins over partial results from different documents. Our method is completely unsupervised, avoiding training-data bottlenecks and being able to cope with rapidly evolving ad hoc topics and formulation style in user questions. QUEST builds a noisy quasi KG with node and edge weights, consisting of dynamically retrieved entity names and relational phrases. It augments this graph with types and semantic alignments, and computes the best answers by an algorithm for Group Steiner Trees. We evaluate QUEST on benchmarks of complex questions, and show that it substantially outperforms state-of-the-art baselines. Xiaolu Lu 0002, Soumajit Pramanik, Rishiraj Saha Roy, Abdalghani Abujabal, Yafang Wang, Gerhard Weikum |
SIGIR | 5 |
| 2018 | Five Shades of Untruth: Finer-Grained Classification of Fake NewsabstractPrior work on algorithmic truth assessment on unreliable content, has mostly pursued binary classifiers - factual vs. fake - and disregarded the finer shades of untruth. On the other hand, manual analysis of questionable content has proposed a more fine-grained classification: distinguishing between hoaxes, irony and propaganda, or the six-way rating by the PolitiFact community. In this paper, we present a principled approach to capture these finer shades in automatically assessing and classifying news articles and claims. We systematically explore a variety of signals from both news and social media, and give an analysis of the underlying features. Yafang Wang, Gerard de Melo, Gerhard Weikum |
ASONAM | 2 |
| 2018 | Social Media vs. News Media: Analyzing Real-World Events from Different Perspectives
Yafang Wang, Zeyuan Cui, Shijun Liu, Gerard de Melo |
DEXA (2) | 3 |
| 2018 | Visualizing Multi-document Semantics via Open Domain Information Extraction
Yongpan Sheng, Zenglin Xu, Yafang Wang, Zhonghui You, Gerard de Melo |
ECML/PKDD (3) | 3 |
| 2018 | Sparse representation based stereoscopic image quality assessment accounting for perceptual cognitive process
Bin Jiang 0003, Yafang Wang, Wen Lu 0004, Qinggang Meng |
Inf. Sci. | 3 |
| 2018 | Classification of gait anomalies from kinect
Qiannan Li, Yafang Wang, Andrei Sharf, Ya Cao, Changhe Tu, Baoquan Chen, Shengyuan Yu |
Vis. Comput. | 2 |
| 2017 | Link prediction by exploiting network formation games in exchangeable graphsabstractIn social network analysis, we often need to predict new links, given some available evidence. This may, for instance, enable us to study user behavior and infer likely new interactions in the near future. Recently, a family of algorithms based on exchangeable graphs has proven effective for link prediction. The network is modeled as an exchangeable array, whose entries can flexibly be traced back to random function priors (e.g., block models, Gaussian Processes). Unfortunately, the burdensome computational complexity of these methods inhibit their application to even just moderate-scale networks. In this paper, we present a novel online training algorithm based on local Gaussian processes on subgraphs, which successfully overcomes this challenge. Moreover, we address the sparsity problem of links in social networks by presenting an improved algorithm based on network formation games. The network formation games we design also shed light on the ambiguity of missing links - not observed vs. non-existing. We evaluate our method against state-of-the-art algorithms on real-world datasets, demonstrating both the effectiveness and the efficiency of our method. Yafang Wang, Bin Liu 0022, Lirong He, Shijun Liu, Gerard de Melo, Zenglin Xu |
IJCNN | 2 |
| 2016 | Summary Generation for Temporal Extractions
Yafang Wang, Zhaochun Ren, Martin Theobald, Maximilian Dylla, Gerard de Melo |
DEXA (1) | 1 |
| 2016 | ShapeLearner: Towards Shape-Based Visual Knowledge HarvestingabstractThe deluge of images on the Web has led to a number of efforts to organize images semantically and mine visual knowledge. Despite enormous progress on categorizing entire images or bounding boxes, only few studies have targeted fine-grained image understanding at the level of specific shape contours. For instance, beyond recognizing that an image portrays a cat, we may wish to distinguish its legs, head, tail, and so on. To this end, we present ShapeLearner, a system that acquires such visual knowledge about object shapes and their parts in a semantic taxonomy, and then is able to exploit this hierarchy in order to analyze new kinds of objects that it has not observed before. ShapeLearner jointly learns this knowledge from sets of segmented images. The space of label and segmentation hypotheses is pruned and then evaluated using Integer Linear Programming. Experiments on a variety of shape classes show the accuracy and effectiveness of our method. Huayong Xu, Yafang Wang, Kang Feng, Gerard de Melo, Andrei Sharf, Baoquan Chen |
ECAI | 2 |
| 2016 | ShapeExplorer: Querying and Exploring Shapes using Visual Knowledge
Tong Ge, Yafang Wang, Gerard de Melo, Zengguang Hao, Andrei Sharf, Baoquan Chen |
EDBT | 2 |
| 2016 | Quality assessment metric of stereo images considering cyclopean integration and visual saliency
Yafang Wang, Baihua Li, Wen Lu 0004, Qinggang Meng, Zhihan Lyu, Dezong Zhao, Zhiqun Gao |
Inf. Sci. | 2 |
| 2013 | A heterogenous automatic feedback semi-supervised method for image rerankingabstractImage reranking, which aims at enhancing the quality of keyword-based image search with the help of image features, recently has become attractive in image search community. A major challenging in this task is that image's visual features do not always well reflect image's semantic meaning. Thus, reranking methods only depending on visual features cannot guarantee to obtain good results. In addition, it is well known that the visual features of an image have strong/weak correlations with its surrounding text. Thus, it is expected that a model considering both visual features and its surrounding text can perform better than those only considering visual features. Motivated by this, in this paper, we propose the HAFSRerank--Heterogenous Automatic Feedback Semi-supervised Reranking method which makes use of both visual and textual features simultaneously during reranking. Specifically, in HAFSRerank, a multigraph is firstly constructed in which each node representing an image includes visual and textual features, and the parallel edges between them are weighted by intra-modal similarity and inter-modal similarity. A heterogenous complete graph is further derived from the multigraph. Then, an automatic feedback graph-based semi-supervised learning method is proposed to propagate the reranking scores on the complete graph, which can make use of the inter-modal similarity to update the weights of heterogenous graph automatically. Finally, the result of the semi-supervised learning is used to rerank the images. The experimental results show that HAFSRerank is superior or highly competitive to some state-of-the-art graph-based reranking methods. Moreover, the proposed reranking algorithm can be well interpreted by Bayesian theory, and does not require complex search models for special queries and any additional input from users. Xin-Chao Xu, Xin-Shun Xu, Yafang Wang, Xiaolin Wang 0003 |
CIKM | 3 |
| 2013 | YaLi: a crowdsourcing plug-in for NERDabstractWe demonstrate the YaLi browser plug-in which discovers named entities in Web pages and provides background knowledge about them. The plug-in is implemented with two purposes. From a user perspective, it enriches the browsing experience with entities, helping users with their information needs. From the research perspective, we aim to improve the methods that are used for named entity recognition and disambiguation (NERD) by leveraging the plug-in as an implicit crowdsourcing platform. YaLi tracks the system's errors and the users' corrections, and also gathers implicit training data for improving NERD accuracy. Yafang Wang, Lili Jiang 0002, Johannes Hoffart, Gerhard Weikum |
SIGIR | 1 |
| 2012 | PRAVDA-live: interactive knowledge harvestingabstractAcquiring high-quality (temporal) facts for knowledge bases is a labor-intensive process. Although there has been recent progress in the area of semi-supervised fact extraction, these approaches still have limitations, including a restricted corpus, a fixed set of relations to be extracted or a lack of assessment capabilities. In this paper we introduce PRAVDA-live, a framework that overcomes these limitations and supports the entire pipeline of interactive knowledge harvesting. To this end, our demo exhibits fact extraction from ad-hoc corpus creation, via relation specification, labeling and assessment all the way to ready-to-use RDF exports. Yafang Wang, Maximilian Dylla, Zhaochun Ren, Marc Spaniol, Gerhard Weikum |
CIKM | 1 |
| 2011 | Harvesting facts from textual web sources by constrained label propagationabstractThere have been major advances on automatically constructing large knowledge bases by extracting relational facts from Web and text sources. However, the world is dynamic: periodic events like sports competitions need to be interpreted with their respective timepoints, and facts such as coaching a sports team, holding political or business positions, and even marriages do not hold forever and should be augmented by their respective timespans. This paper addresses the problem of automatically harvesting temporal facts with such extended time-awareness. We employ pattern-based gathering techniques for fact candidates and construct a weighted pattern-candidate graph. Our key contribution is a system called PRAVDA based on a new kind of label propagation algorithm with a judiciously designed loss function, which iteratively processes the graph to label good temporal facts for a given set of target relations. Our experiments with online news and Wikipedia articles demonstrate the accuracy of this method. Yafang Wang, Bin Yang 0002, Lizhen Qu, Marc Spaniol, Gerhard Weikum |
CIKM | 1 |
| 2010 | Timely YAGO: harvesting, querying, and visualizing temporal knowledge from WikipediaabstractRecent progress in information extraction has shown how to automatically build large ontologies from high-quality sources like Wikipedia. But knowledge evolves over time; facts have associated validity intervals. Therefore, ontologies should include time as a first-class dimension. In this paper, we introduce Timely YAGO, which extends our previously built knowledge base YAGO with temporal aspects. This prototype system extracts temporal facts from Wikipedia infoboxes, categories, and lists in articles, and integrates these into the Timely YAGO knowledge base. We also support querying temporal facts, by temporal predicates in a SPARQL-style language. Visualization of query results is provided in order to better understand of the dynamic nature of knowledge. Yafang Wang, Mingjie Zhu, Lizhen Qu, Marc Spaniol, Gerhard Weikum |
EDBT | 1 |