VLDB 2026 Research / reviewers in the wild / expert
Teng Zhang 0001
dblp:38/5156-1
· DBLP profile ↗
21ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-4731-1477ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-instance multi-label position-aware doubly graph convolutional networks
Zhi Li 0048, Teng Zhang 0001, Caiwu Jiang, Xuanhua Shi, Hai Jin 0001 |
Frontiers Comput. Sci. | 2 |
| 2025 | E2CNN: entity-type-enriched cascaded neural network for Chinese financial relation extractionabstractAbstract Knowledge Graphs (KGs) are pivotal for effectively organizing and managing structured information across various applications. Financial KGs have been successfully employed in advancing applications such as audit, anti-fraud, and anti-money laundering. Despite their success, the construction of Chinese financial KGs has seen limited research due to the complex semantics. A significant challenge is the overlap triples problem, where entities feature in multiple relations within a sentence, hampering extraction accuracy–more than 39% of the triples in Chinese datasets exhibit the overlap triples. To address this, we propose the Entity-type-Enriched Cascaded Neural Network (E 2 CNN), leveraging special tokens for entity boundaries and types. E 2 CNN ensures consistency in entity types and excludes specific relations, mitigating overlap triple problems and enhancing relation extraction. Besides, we introduce the available Chinese financial dataset F in C orpus .CN, annotated from annual reports of 2,000 companies, containing 48,389 entities and 23,368 triples. Experimental results on the DUIE dataset and F in C orpus .CN underscore E 2 CNN’s superiority over state-of-the-art models. Mengfan Li 0001, Xuanhua Shi, Chenqi Qiao, Yao Wan 0001, Teng Zhang 0001, Hai Jin 0001 |
Frontiers Comput. Sci. | 7 |
| 2025 | RuYi: Optimizing Burst Buffer Through Automated, Fine-Grained Process-to-BB MappingabstractCurrent supercomputers use an SSD-based storage layer called Burst Buffer (BB) to provide I/O-intensive applications with accelerated storage access. However, efficiently utilizing this limited and expensive storage remains a critical issue, creating an urgent need for implementing Quality of Service (QoS) in BB. To address this, we propose RuYi, a QoS-aware method to provide applications with bandwidth guarantees in the BB file system. RuYi tackles two main issues. First, it quantitatively profiles available bandwidth resources in BB to ensure reliable QoS, a crucial aspect seldom studied in the literature. Second, RuYi offers fine-grained process-level QoS via an innovative process-to-BB mapping, maximizing resource utilization—something not achievable with conventional coarse-grained compute-to-BB mapping. We evaluated RuYi on a subsystem of the leading exascale supercomputer Sunway, consisting of 4,000 compute nodes and 200 BB nodes. The experimental results demonstrate that RuYi achieves an impressive end-to-end bandwidth control accuracy of 97%, while improving BB utilization by up to 116% compared to conventional coarse-grained compute-to-BB mapping. Yusheng Hua, Xuanhua Shi, Ligang He, Teng Zhang 0001, Hai Jin 0001, Yong Chen 0001 |
IEEE Trans. Computers | 5 |
| 2023 | Incremental and Decremental Optimal Margin Distribution LearningabstractIncremental and decremental learning (IDL) deals with the tasks where new data arrives sequentially as a stream or old data turns unavailable continually due to the privacy protection. Existing IDL methods mainly focus on support vector machine and its variants with linear-type loss. There are few studies about the quadratic-type loss, whose Lagrange multipliers are unbounded and much more difficult to track. In this paper, we take the latest statistical learning framework optimal margin distribution machine (ODM) which involves a quadratic-type loss due to the optimization of margin variance, for example, and equip it with the ability to handle IDL tasks. Our proposed ID-ODM can avoid updating the Lagrange multipliers in an infinite range by determining their optimal values beforehand so as to enjoy much more efficiency. Moreover, ID-ODM is also applicable when multiple instances come and leave simultaneously. Extensive empirical studies show that ID-ODM can achieve 9.1x speedup on average with almost no generalization lost compared to retraining ODM on new data set from scratch. Li-Jun Chen, Teng Zhang 0001, Xuanhua Shi, Hai Jin 0001 |
IJCAI | 2 |
| 2023 | Scalable Optimal Margin Distribution MachineabstractOptimal margin Distribution Machine (ODM) is a newly proposed statistical learning framework rooting in the novel margin theory, which demonstrates better generalization performance than the traditional large margin based counterparts. Nonetheless, it suffers from the ubiquitous scalability problem regarding both computation time and memory as other kernel methods. This paper proposes a scalable ODM, which can achieve nearly ten times speedup compared to the original ODM training method. For nonlinear kernels, we propose a novel distribution-aware partition method to make the local ODM trained on each partition be close and converge faster to the global one. When linear kernel is applied, we extend a communication efficient SVRG method to accelerate the training further. Extensive empirical studies validate that our proposed method is highly computational efficient and almost never worsen the generalization. Nan Cao 0002, Teng Zhang 0001, Xuanhua Shi, Hai Jin 0001 |
IJCAI | 3 |
| 2022 | Posistive-Unlabeled Learning via Optimal Transport and Margin DistributionabstractPositive-unlabeled (PU) learning deals with the circumstances where only a portion of positive instances are labeled, while the rest and all negative instances are unlabeled, and due to this confusion, the class prior can not be directly available. Existing PU learning methods usually estimate the class prior by training a nontraditional probabilistic classifier, which is prone to give an overestimation. Moreover, these methods learn the decision boundary by optimizing the minimum margin, which is not suitable in PU learning due to its sensitivity to label noise. In this paper, we enhance PU learning methods from the above two aspects. More specifically, we first explicitly learn a transformation from unlabeled data to positive data by entropy regularized optimal transport to achieve a much more precise estimation for class prior. Then we switch to optimizing the margin distribution, rather than the minimum margin, to obtain a label noise insensitive classifier. Extensive empirical studies on both synthetic and real-world data sets demonstrate the superiority of our proposed method. Nan Cao 0002, Teng Zhang 0001, Xuanhua Shi, Hai Jin 0001 |
IJCAI | 2 |
| 2022 | On the Optimization of Margin DistributionabstractMargin has played an important role on the design and analysis of learning algorithms during the past years, mostly working with the maximization of the minimum margin. Recent years have witnessed the increasing empirical studies on the optimization of margin distribution according to different statistics such as medium margin, average margin, margin variance, etc., whereas there is a relative paucity of theoretical understanding. In this work, we take one step on this direction by providing a new generalization error bound, which is heavily relevant to margin distribution by incorporating ingredients such as average margin and semi-variance, a new margin statistics for the characterization of margin distribution. Inspired by the theoretical findings, we propose the MSVMAv, an efficient approach to achieve better performance by optimizing margin distribution in terms of its empirical average margin and semi-variance. We finally conduct extensive experiments to show the superiority of the proposed MSVMAv approach. Meng-Zhang Qian, Zheng Ai, Teng Zhang 0001, Wei Gao 0008 |
IJCAI | 3 |
| 2022 | Enhance Temporal Knowledge Graph Completion via Time-Aware Attention Graph Convolutional Network
HaoHui Wei, Hong Huang 0001, Teng Zhang 0001, Xuanhua Shi, Hai Jin 0001 |
ECML/PKDD (2) | 3 |
| 2022 | Entropy Weight Allocation: Positive-unlabeled Learning via Optimal TransportabstractPositive-unlabeled learning (PU learning) aims to deal with the problem that only a fraction of positive instances are known. Due to the absence of negative instances, ordinary learning models cannot be directly applied. Existing PU learning methods either explicitly choose some unlabeled instances as negative instances in advance or reformulate the task as a weighted learning problem. Since working in such an ad-hoc fashion, these methods often suffer a bad performance and only have limited usage. This paper proposes a novel instance-dependent weighting method entropy weight allocation (EWA) for PU learning by optimal transport (OT). More specifically, we allocate each unlabeled instance an elaborate weight indicating the possibility that it is an underlying negative instance. Then any ordinary weighted learning models can be used to obtain a PU classifier. By concatenating EWA with four celebrated classification models, we show that EWA is a broad-spectrum weighting method that can boost almost all the mainstream machine learning models for PU learning. Wen Gu, Teng Zhang 0001, Hai Jin 0001 |
SDM | 2 |
| 2022 | Detecting Arbitrage on Ethereum Through Feature Fusion and Positive-Unlabeled LearningabstractDue to the lack of supervision in the decentralized exchanges (DEXs), arbitrageurs can utilize information and take advantage of price gap to make profits over such platforms such as Ethereum blockchain. DEX arbitrage poses possibilities and opportunities for defrauding and can seriously impair the operation of the Ethereum ecosystem. It motivates this work to explore and characterize the unique features of arbitrage which differ from other frauds such as money laundering and Ponzi games for better detection. This work makes the first attempt for detecting arbitrage on Ethereum through feature fusion and positive-unlabeled learning (PU learning). We first conduct an in-depth analysis and exploit two-fold arbitrage features by fusion including: 1) statistical features that explicitly represent the node activity levels according to expert knowledge; and 2) structural features that implicitly encode the transactions information by graph machine learning. We then apply PU learning to generate negative instances for compensating the imbalanced arbitrage datasets. We evaluate our proposed method through extensive experiments over a real-world dataset and demonstrate that it can achieve 90% accuracy in detecting arbitrage activities on Ethereum. Hai Jin 0001, Jiang Xiao 0001, Teng Zhang 0001, Xiaohai Dai, Bo Li 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2021 | Partial Multi-Label Optimal Margin Distribution MachineabstractPartial multi-label learning deals with the circumstance in which the ground-truth labels are not directly available but hidden in a candidate label set. Due to the presence of other irrelevant labels, vanilla multi-label learning methods are prone to be misled and fail to generalize well on unseen data, thus how to enable them to get rid of the noisy labels turns to be the core problem of partial multi-label learning. In this paper, we propose the Partial Multi-Label Optimal margin Distribution Machine (PML-ODM), which distinguishs the noisy labels through explicitly optimizing the distribution of ranking margin, and exhibits better generalization performance than minimum margin based counterparts. In addition, we propose a novel feature prototype representation to further enhance the disambiguation ability, and the non-linear kernels can also be applied to promote the generalization performance for linearly inseparable data. Extensive experiments on real-world data sets validates the superiority of our proposed method. Nan Cao 0002, Teng Zhang 0001, Hai Jin 0001 |
IJCAI | 2 |
| 2021 | On the noise estimation statistics
Wei Gao 0008, Teng Zhang 0001, Bin-Bin Yang, Zhi-Hua Zhou |
Artif. Intell. | 2 |
| 2020 | Optimal Margin Distribution Learning in Dynamic EnvironmentsabstractRecently a promising research direction of statistical learning has been advocated, i.e., the optimal margin distribution learning with the central idea that instead of the minimal margin, the margin distribution is more crucial to the generalization performance. Although the superiority of this new learning paradigm has been verified under batch learning settings, it remains open for online learning settings, in particular, the dynamic environments in which the underlying decision function varies over time. In this paper, we propose the dynamic optimal margin distribution machine and theoretically analyze its regret. Although the obtained bound has the same order with the best known one, our method can significantly relax the restrictive assumption that the function variation should be given ahead of time, resulting in better applicability in practical scenarios. We also derive an excess risk bound for the special case when the underlying decision function only evolves several discrete changes rather than varying continuously. Extensive experiments on both synthetic and real data sets demonstrate the superiority of our method. Teng Zhang 0001, Hai Jin 0001 |
AAAI | 1 |
| 2020 | Optimal Margin Distribution Machine for Multi-Instance LearningabstractMulti-instance learning (MIL) is a celebrated learning framework where each example is represented as a bag of instances. An example is negative if it has no positive instances, and vice versa if at least one positive instance is contained. During the past decades, various MIL algorithms have been proposed, among which the large margin based methods is a very popular class. Recently, the studies on margin theory disclose that the margin distribution is of more importance to generalization ability than the minimal margin. Inspired by this observation, we propose the multi-instance optimal margin distribution machine, which can identify the key instances via explicitly optimizing the margin distribution. We also extend a stochastic accelerated mirror prox method to solve the formulated minimax problem. Extensive experiments show the superiority of the proposed method. Teng Zhang 0001, Hai Jin 0001 |
IJCAI | 1 |
| 2020 | Optimal Margin Distribution MachineabstractSupport Vector Machine (SVM) has always been one of the most successful learning algorithms, with the central idea of maximizing theminimum margin, i.e., the smallest distance from the instances to the classification boundary. However, recent theoretical results disclosed that maximizing the minimum margin does not necessarily lead to better generalization performance, and instead, themargin distributionhas been proven to be more crucial. Based on this idea, we propose the Optimal margin Distribution Machine (ODM), which can achieve a better generalization performance by optimizing the margin distribution explicitly. We characterize the margin distribution by the first- and second-order statistics, i.e., the margin mean and variance. The proposed method is a general learning approach which can be applied in any place where SVMs are used, and its superiority is verified both theoretically and empirically in this paper. Teng Zhang 0001, Zhi-Hua Zhou |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Coreset Stochastic Variance-Reduced Gradient with Application to Optimal Margin Distribution Machine
Zhi-Hao Tan, Teng Zhang 0001, Wei Wang 0028 |
AAAI | 2 |
| 2018 | Optimal Margin Distribution ClusteringabstractMaximum margin clustering (MMC), which borrows the large margin heuristic from support vector machine (SVM), has achieved more accurate results than traditional clustering methods. The intuition is that, for a good clustering, when labels are assigned to different clusters, SVM can achieve a large minimum margin on this data. Recent studies, however, disclosed that maximizing the minimum margin does not necessarily lead to better performance, and instead, it is crucial to optimize the margin distribution. In this paper, we propose a novel approach ODMC (Optimal margin Distribution Machine for Clustering), which tries to cluster the data and achieve optimal margin distribution simultaneously. Specifically, we characterize the margin distribution by the first- and second-order statistics, i.e., the margin mean and variance, and extend a stochastic mirror descent method to solve the resultant minimax problem. Moreover, we prove theoretically that ODMC has the same convergence rate with state-of-the-art cutting plane based algorithms but involves much less computation cost per iteration, so our method is much more scalable than existing approaches. Extensive experiments on UCI data sets show that ODMC is significantly better than compared methods, which verifies the superiority of optimal margin distribution learning. Teng Zhang 0001, Zhi-Hua Zhou |
AAAI | 1 |
| 2018 | Semi-Supervised Optimal Margin Distribution MachinesabstractSemi-supervised support vector machines is an extension of standard support vector machines with unlabeled instances, and the goal is to find a label assignment of the unlabeled instances, so that the decision boundary has the maximal \textit{minimum margin} on both the original labeled instances and unlabeled instances. Recent studies, however, disclosed that maximizing the minimum margin does not necessarily lead to better performance, and instead, it is crucial to optimize the \textit{margin distribution}. In this paper, we propose a novel approach SODM (Semi-supervised Optimal margin Distribution Machine), which tries to assign the label to unlabeled instances and achieve optimal margin distribution simultaneously. Specifically, we characterize the margin distribution by the first- and second-order statistics, i.e., the margin mean and variance, and extend a stochastic mirror prox method to solve the resultant minimax problem. Extensive experiments on UCI data sets show that SODM is significantly better than compared methods, which verifies the superiority of optimal margin distribution learning. Teng Zhang 0001, Zhi-Hua Zhou |
IJCAI | 1 |
| 2017 | Multi-Class Optimal Margin Distribution MachineabstractRecent studies disclose that maximizing the minimum margin like support vector machines does not necessarily lead to better generalization performances, and instead, it is crucial to optimize the margin distribution. Although it has been shown that for binary classification, characterizing the margin distribution by the first- and second-order statistics can achieve superior performance. It still remains open for multi-class classification, and due to the complexity of margin for multi-class classification, optimizing its distribution by mean and variance can also be difficult. In this paper, we propose mcODM (multi-class Optimal margin Distribution Machine), which can solve this problem efficiently. We also give a theoretical analysis for our method, which verifies the significance of margin distribution for multi-class classification. Empirical study further shows that mcODM always outperforms all four versions of multi-class SVMs on all experimental data sets. Teng Zhang 0001, Zhi-Hua Zhou |
ICML | 1 |
| 2017 | Efficient Label Contamination Attacks Against Black-Box Learning ModelsabstractLabel contamination attack (LCA) is an important type of data poisoning attack where an attacker manipulates the labels of training data to make the learned model beneficial to him. Existing work on LCA assumes that the attacker has full knowledge of the victim learning model, whereas the victim model is usually a black-box to the attacker. In this paper, we develop a Projected Gradient Ascent (PGA) algorithm to compute LCAs on a family of empirical risk minimizations and show that an attack on one victim model can also be effective on other victim models. This makes it possible that the attacker designs an attack against a substitute model and transfers it to a black-box victim model. Based on the observation of the transferability, we develop a defense algorithm to identify the data points that are most likely to be attacked. Empirical studies show that PGA significantly outperforms existing baselines and linear learning models are better substitute models than nonlinear ones. Mengchen Zhao, Bo An 0001, Wei Gao 0008, Teng Zhang 0001 |
IJCAI | 4 |
| 2014 | Large margin distribution machineabstractSupport vector machine (SVM) has been one of the most popular learning algorithms, with the central idea of maximizing the minimum margin, i.e., the smallest distance from the instances to the classification boundary. Recent theoretical results, however, disclosed that maximizing the minimum margin does not necessarily lead to better generalization performances, and instead, the margin distribution has been proven to be more crucial. In this paper, we propose the Large margin Distribution Machine (LDM), which tries to achieve a better generalization performance by optimizing the margin distribution. We characterize the margin distribution by the first- and second-order statistics, i.e., the margin mean and variance. The LDM is a general learning approach which can be used in any place where SVM can be applied, and its superiority is verified both theoretically and empirically in this paper. Teng Zhang 0001, Zhi-Hua Zhou |
KDD | 1 |