Yifu Tang

dblp:305/7862 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2027 PTIL: Partitioned Tree-Cover Interval Labeling for Restricted-Domain Reachability Queries
abstract
Reachability queries are a core operation in graph analytics. Existing labeling methods assume a global query space and construct uniform labels for all vertices. In practice, however, many systems issue queries over semantically structured endpoint domains such as V_s × V_t, where only a subset of vertices serve as meaningful sources or targets. This mismatch makes global labeling redundant and inefficient. We propose Partitioned Tree-Cover Interval Labeling (PTIL), a restricted-domain reachability indexing framework that extends classical Tree-Cover Labeling. PTIL timestamps only target vertices and stores compact disjoint intervals at sources, reducing each query to a single membership test while avoiding unnecessary global labels. By partitioning targets into groups, PTIL provides an explicit and predictable space–time trade-off: finer grouping yields lower latency at the cost of larger index size. PTIL also supports dynamic updates to endpoint domains. We further develop two instantiations: PTIL-g, a grouped scheme for scalable domain-aligned indexing, and PTIL-p, which additionally leverages query distributions for further prioritization when available. Experiments on real and synthetic graphs show that PTIL-g achieves order-of-magnitude lower query latency than state-of-the-art baselines, while PTIL-p provides additional gains under skewed workloads, with near-linear and predictable space–time behavior.
Huangleshuai He, Zhengyi Yang 0001, Yifu Tang, Dong Wen 0001, Tianming Zhang
EDBT3
2025 A Survey on Multi-View Knowledge Graph: Generation, Fusion, Applications and Future Directions
abstract
Knowledge Graphs (KGs) have revolutionized structured knowledge representation, yet their capacity to model real-world complexity and heterogeneity remains fundamentally constrained. The emerging paradigm of Multi-View Knowledge Graphs (MVKGs) addresses this gap through multi-view learning, but existing research lacks systematic integration. This survey provides the first systematic consolidation of MVKG methodologies, with four pivotal contributions: 1) The first unified taxonomy of view generation paradigms that rigorously categorizes view into four types: structure, semantic, representation, and knowledge & modality; 2) A novel methodological typology for view fusion that systematically classifies techniques by fusion targets (feature, decision, and hybrid); 3) Task-centric application mapping that bridges theoretical MVKG constructs to node/link/graph-level downstream tasks; 4) A forward-looking roadmap identifying underexplored challenges. By unifying fragmented methodologies and formalizing MVKG design principles, this survey serves as a roadmap for advancing KG versatility in complex AI-driven scenarios. In doing so, it paves the way for more efficient knowledge integration, enhanced decision-making, and cross-domain learning in real-world applications.
Xiaohui Tao 0001, Taotao Cai, Yifu Tang, Haoran Xie 0001, Lin Li 0001, Jianxin Li 0001, Qing Li 0001
IJCAI4
2025 Finding Time-Proximity Communities in Temporal Heterogeneous Information Networks
Yifu Tang, Chengfei Liu, Lu Chen 0008, Rui Zhou 0001, Jianxin Li 0001
Proc. VLDB Endow.1
2025 Tabular-textual question answering: From parallel program generation to large language models
abstract
Abstract Hybrid tabular-textual question answering (HTQA) involves integrating multiple data sources, traditionally managed through LSTM-based step-by-step reasoning. However, such sequential approaches are prone to exposure bias and cumulative errors, limiting their effectiveness. This paper first introduces an innovative parallel program generation method, ConcurGen, aiming to transform this paradigm by simultaneously formulating comprehensive program constructs that seamlessly blend operations and values. This approach not only rectifies the inherent pitfalls of sequential methodologies but also infuses efficiency into the process. Through our further research, we found that some HTQA scenarios extend beyond traditional question-answering, often involving open-ended questions that demand dynamic, context-aware response generation. Therefore, we introduce a second framework that leverages large language models (LLMs) to effectively answer both traditional and open-ended questions. Our method demonstrates substantial improvements over existing models such as FinQANet and MT2Net on benchmarks including ConvFinQA and MultiHiertt, achieving new state-of-the-art performance across multiple evaluation metrics. In addition to its accuracy, it delivers a nearly 21x speedup in program generation, significantly enhancing inference efficiency. Unlike traditional models, our system maintains robust performance as the complexity of numerical reasoning increases, highlighting its adaptability in challenging scenarios. Furthermore, supplementary experiments on the LLM-based framework show that it provides enriched answer justifications while achieving similar performance to ConcurGen on standard benchmarks.
Xushuo Tang, Liuyi Chen, Wenke Yang 0001, Zhengyi Yang 0001, Mingchen Ju, Yifu Tang
World Wide Web (WWW)8
2024 Stochastic Multi-Armed Bandits with Strongly Reward-Dependent Delays
abstract
There has been increasing interest in applying multi-armed bandits to adaptive designs in clinical trials. However, most literature assumes that a previous patient’s survival response of a treatment is known before the next patient is treated, which is unrealistic. The inability to account for response delays is cited frequently as one of the problems in using adaptive designs in clinical trials. More critically, the “delays” in observing the survival response are the same as the rewards rather than being external stochastic noise. We formalize this problem as a novel stochastic multi-armed bandit (MAB) problem with reward-dependent delays, where the delay at each round depends on the reward generated on the same round. For general reward/delay distributions with finite expectation, our proposed censored-UCB algorithm achieves near-optimal regret in terms of both problem-dependent and problem-independent bounds. With bounded or sub-Gaussian reward distributions, the upper bounds are optimal with a matching lower bound. Our theoretical results and the algorithms’ effectiveness are validated by empirical experiments.
Yifu Tang, Yingfei Wang, Zeyu Zheng 0002
AISTATS1
2024 Reliability-Driven Local Community Search in Dynamic Networks
abstract
Community search over large dynamic graph has become an important research problem in modern complex networks, such as the online social network, collaboration network and biological networks. Network data in the time-varied environment has motivated several recent studies to identify the evolution of the communities. However, these studies mostly match communities of different snapshot or utilize the aggregation of the disjoint structural information and ignores the cohesion continuity. To fill this research gap, in this work, we propose a novel$(\theta ,k)$-core reliable community (CRC) and define the reliable community search problem which jointly considers member engagement, connection strength and cohesion continuity of the community in the dynamic network. We propose an online search algorithm based on eligible edge filtering and we further construct the Weighted Core Forest-Index (WCF-index) and develop efficient index-based querying algorithm with strong pruning properties. We also propose top-$l$reliable community search problem that couples query based distance to reduce the free rider effect in local community search and support flexible multiple query vertices. Extensive experiments are conducted to show the efficiency and effectiveness of the proposed algorithms.
Yifu Tang, Jianxin Li 0001, Nur Al Hasan Haldar, Ziyu Guan, Jiajie Xu 0001, Chengfei Liu
IEEE Trans. Knowl. Data Eng.1
2024 CPS Attack Detection under Limited Local Information in Cyber Security: An Ensemble Multi-Node Multi-Class Classification Approach
abstract
Cybersecurity breaches are common anomalies for distributed cyber-physical systems (CPS). However, the cyber security breach classification is still a difficult problem, even using cutting-edge artificial intelligence (AI) approaches. In this article, we study a multi-class classification problem in cyber security for attack detection. A challenging multi-node data-censoring case is considered. In such a case, data within each data center/node cannot be shared while the local data is incomplete. Particularly, local nodes contain only a part of the multiple classes. In order to train a global multi-class classifier without sharing the raw data across all nodes, we design a multi-node multi-class classification ensemble approach which is the main result of our study. By gathering the estimated parameters of the binary classifiers and data densities from each local node, the missing information for each local node is completed to build the global multi-class classifier. Numerical experiments are given to validate the effectiveness of the proposed approach under the multi-node data-censoring case. Under such a case, we even show the out-performance of the proposed approach over the full-data approach.
Yifu Tang, Haimeng Zhao, Xieheng Wang, Fangyu Li 0002, Jingyi Zhang 0004
ACM Trans. Sens. Networks2
2023 Course map learning with graph convolutional network based on AuCM
abstract
Abstract Concept map provides a concise structured representation of knowledge in the educational scenario. It consists of various concepts connected by prerequisite dependencies. With the abundance of educational resources available through MOOCs, encyclopedias, and electronic textbooks, extracting prerequisite dependencies and building concept maps becomes feasible. However, publicly accessible taxonomies or learning object information that can help identify prerequisites are rare. To address this, we have constructed a comprehensive dataset called the Australian Course Map data (AuCM), specifically tailored for training concept maps in the IT/CS field. The dataset comprises course descriptions from 14 different Australian universities. To identify prerequisite relationships between course concepts, we have employed an embedding-based approach that combines the Graph Convolutional Network (GCN) with pairwise features of concepts. We have evaluated the performance of our model with non-neural classifiers and neural networks for extracting these prerequisite relations.
Jianing Xia, Yifu Tang, Shuiqiao Yang
World Wide Web (WWW)3
2022 AuCM: Course Map Data Analytics for Australian IT Programs in Higher Education
Jianing Xia, Yifu Tang, Taige Zhao, Jianxin Li 0001
ADMA (1)2
2022 Reliable Community Search in Dynamic Networks
abstract
Searching for local communities is an important research problem that supports advanced data analysis in various complex networks, such as social networks, collaboration networks, cellular networks, etc. The evolution of such networks over time has motivated several recent studies to identify local communities in dynamic networks. However, these studies only utilize the aggregation of disjoint structural information to measure the quality and ignore the reliability of the communities in a continuous time interval. To fill this research gap, we propose a novel (θ, k )- core reliable community (CRC) model in the weighted dynamic networks, and define the problem of most reliable community search that couples the desirable properties of connection strength, cohesive structure continuity, and the maximal member engagement. To solve this problem, we first develop a novel edge filtering based online CRC search algorithm that can effectively filter out the trivial edge information from the networks while searching for a reliable community. Further, we propose an index structure, Weighted Core Forest-Index (WCF-index), and devise an index-based dynamic programming CRC search algorithm, that can prune a large number of insignificant intermediate results and support efficient query processing. Finally, we conduct extensive experiments systematically to demonstrate the efficiency and effectiveness of our proposed algorithms on eight real datasets under various experimental settings.
Yifu Tang, Jianxin Li 0001, Nur Al Hasan Haldar, Ziyu Guan, Jiajie Xu 0001, Chengfei Liu
Proc. VLDB Endow.1
2021 JKT: A joint graph convolutional network based Deep Knowledge Tracing
Jianxin Li 0001, Yifu Tang, Taige Zhao, Yunliang Chen 0002, Ziyu Guan
Inf. Sci.3