Wenqi He

dblp:165/9599 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Information extraction and text analysis · 80% Trustworthy machine learning · 15% Knowledge representation and reasoning · 4%
Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
entity typing
0.522016
Label Noise Reduction in Entity Typing by Heterogeneous Partial-Label Embedding · KDD 2016
AFET: Automatic Fine-Grained Entity Typing by Hierarchical Partial-Label Embedding · EMNLP 2016
Natural language and speech › Information extraction and text analysis › relation extraction
joint entity and relation extraction
0.312017
CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases · WWW 2017
Natural language and speech › Information extraction and text analysis
relation extraction
0.312017
CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases · WWW 2017
Knowledge graphs
distant supervision
0.312017
CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases · WWW 2017
Knowledge graphs
knowledge graph construction
0.312017
CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases · WWW 2017
Natural language and speech › Information extraction and text analysis › entity typing
fine-grained entity typing
0.212016
AFET: Automatic Fine-Grained Entity Typing by Hierarchical Partial-Label Embedding · EMNLP 2016
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
noisy label correction
0.212016
Label Noise Reduction in Entity Typing by Heterogeneous Partial-Label Embedding · KDD 2016
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology › concept hierarchy
type hierarchy
0.112016
AFET: Automatic Fine-Grained Entity Typing by Hierarchical Partial-Label Embedding · EMNLP 2016

Methods — techniques the papers use, named apart from their topics

partial-label loss · 0.6joint embedding · 0.6margin-based loss · 0.5heterogeneous partial-label embedding · 0.5hierarchical partial-label embedding · 0.2distant supervision · 0.2
YearPublicationVenuePosition
2025 IIT: Accurate Decentralized Application Identification Through Mining Intra- and Inter-Flow Relationships
abstract
Identifying Decentralized Applications (DApps) from encrypted network traffic plays an important role in areas such as network management and threat detection. However, DApps deployed on the same platform use the same encryption settings, resulting in DApps generating encrypted traffic with great similarity. In addition, existing flow-based methods only consider each flow as an isolated individual and feed it sequentially into the neural network for feature extraction, ignoring other rich information introduced between flows, and therefore the relationship between different flows is not effectively utilized. In this study, we propose a novel encrypted traffic classification model IIT to heterogeneously mine the potential features of intra- and inter-flows, which contain two types of encoders based on the multi-head self-attention mechanism. By combining the complementary intra- and inter-flow perspectives, the entire process of information flow can be more completely understood and described. IIT provides a more complete perspective on network flows, with the intra-flow perspective focusing on information transfer between different packets within a flow, and the inter-flow perspective placing more emphasis on information interaction between different flows. We captured 44 classes of DApps in the real world and evaluated the IIT model on two datasets, including DApps and malicious traffic classification tasks. The results demonstrate that the IIT model achieves a classification accuracy of greater than 97% on the real-world dataset of 44 DApps, outperforming other state-of-the-art methods. In addition, the IIT model exhibits good generalization in the malicious traffic classification task.
Qianwei Meng, Qingjun Yuan, Weina Niu, Yongjuan Wang, Siqi Lu, Guangsong Li, Xiangbin Wang, Wenqi He
IEEE Trans. Netw. Serv. Manag.8
2023 A novel four image encryption approach in sparse domain based on biometric keys
Gaurav Verma 0001, Wenqi He
Multim. Tools Appl.2
2018 Open-Schema Event Profiling for Massive News Corpora
abstract
With the rapid growth of online information services, a sheer volume of news data becomes available. To help people quickly digest the explosive information, we define a new problem - schema-based news event profiling - profiling events reported in open-domain news corpora, with a set of slots and slot-value pairs for each event, where the set of slots forms the schema of an event type. Such profiling not only provides readers with concise views of events, but also facilitates various applications such as information retrieval, knowledge graph construction and question answering. It is however a quite challenging task. The first challenge is to find out events and event types because they are both initially unknown. The second difficulty is the lack of pre-defined event-type schemas. Lastly, even with the schemas extracted, to generate event profiles from them is still essential yet demanding.
Quan Yuan 0001, Xiang Ren 0001, Wenqi He, Chao Zhang 0014, Xinhe Geng, Lifu Huang, Heng Ji 0001, Chin-Yew Lin, Jiawei Han 0001
CIKM3
2017 CoType: Joint Extraction of Typed Entities and Relations with Knowledge Bases
abstract
Extracting entities and relations for types of interest from text is important for understanding massive text corpora. Traditionally, systems of entity relation extraction have relied on human-annotated corpora for training and adopted an incremental pipeline. Such systems require additional human expertise to be ported to a new domain, and are vulnerable to errors cascading down the pipeline. In this paper, we investigate joint extraction of typed entities and relations with labeled data heuristically obtained from knowledge bases (i.e., distant supervision). As our algorithm for type labeling via distant supervision is context-agnostic, noisy training data poses unique challenges for the task. We propose a novel domain-independent framework, called CoType, that runs a data-driven text segmentation algorithm to extract entity mentions, and jointly embeds entity mentions, relation mentions, text features and type labels into two low-dimensional spaces (for entity and relation mentions respectively), where, in each space, objects whose types are close will also have similar representations. CoType, then using these learned embeddings, estimates the types of test (unlinkable) mentions. We formulate a joint optimization problem to learn embeddings from text corpora and knowledge bases, adopting a novel partial-label loss function for noisy labeled data and introducing an object "translation" function to capture the cross-constraints of entities and relations on each other. Experiments on three public datasets demonstrate the effectiveness of CoType across different domains (e.g., news, biomedical), with an average of 25% improvement in F1 score compared to the next best method.
Xiang Ren 0001, Zeqiu Wu, Wenqi He, Meng Qu, Clare R. Voss, Heng Ji 0001, Tarek F. Abdelzaher, Jiawei Han 0001
WWW3
2016 AFET: Automatic Fine-Grained Entity Typing by Hierarchical Partial-Label Embedding
abstract
Distant supervision has been widely used in current systems of fine-grained entity typing to automatically assign categories (entity types) to entity mentions.However, the types so obtained from knowledge bases are often incorrect for the entity mention's local context.This paper proposes a novel embedding method to separately model "clean" and "noisy" mentions, and incorporates the given type hierarchy to induce loss functions.We formulate a joint optimization problem to learn embeddings for mentions and typepaths, and develop an iterative algorithm to solve the problem.Experiments on three public datasets demonstrate the effectiveness and robustness of the proposed method, with an average 15% improvement in accuracy over the next best compared method 1 . * Equal contribution.1 Codes and datasets used in this paper can be downloaded at https://github.com/shanzhenren/AFET.
Xiang Ren 0001, Wenqi He, Meng Qu, Lifu Huang, Heng Ji 0001, Jiawei Han 0001
EMNLP2
2016 Label Noise Reduction in Entity Typing by Heterogeneous Partial-Label Embedding
abstract
Current systems of fine-grained entity typing use distant supervision in conjunction with existing knowledge bases to assign categories (type labels) to entity mentions. However, the type labels so obtained from knowledge bases are often noisy (i.e., incorrect for the entity mention's local context). We define a new task, Label Noise Reduction in Entity Typing (LNR), to be the automatic identification of correct type labels (type-paths) for training examples, given the set of candidate type labels obtained by distant supervision with a given type hierarchy. The unknown type labels for individual entity mentions and the semantic similarity between entity types pose unique challenges for solving the LNR task. We propose a general framework, called PLE, to jointly embed entity mentions, text features and entity types into the same low-dimensional space where, in that space, objects whose types are semantically close have similar representations. Then we estimate the type-path for each training example in a top-down manner using the learned embeddings. We formulate a global objective for learning the embeddings from text corpora and knowledge bases, which adopts a novel margin-based loss that is robust to noisy labels and faithfully models type correlation derived from knowledge bases. Our experiments on three public typing datasets demonstrate the effectiveness and robustness of PLE, with an average of 25% improvement in accuracy compared to next best method.
Xiang Ren 0001, Wenqi He, Meng Qu, Clare R. Voss, Heng Ji 0001, Jiawei Han 0001
KDD2
2015 A Deep Neural Network for Modeling Music
abstract
We propose a convolutional neural network architecture with k-max pooling layer for semantic modeling of music. The aim of a music model is to analyze and represent the semantic content of music for purposes of classification, discovery, or clustering. The k-max pooling layer is used in the network to make it possible to pool the k most active features, capturing the semantic-rich and time-varying information about music. Our network takes an input music as a sequence of audio words, where each audio word is associated with a distributed feature vector that can be fine-tuned by backpropagating errors during the training. The architecture allows us to take advantage of the better trained audio word embeddings and the deep structures to produce more robust music representations. Experiment results with two different music collections show that our neural networks achieved the best accuracy in music genre classification comparing with three state-of-art systems.
Pengjing Zhang, Xiaoqing Zheng, Siyan Li, Sheng Qian, Wenqi He, Shangtong Zhang
ICMR6