EDBT 2026 Demo / reviewers in the wild / expert
Yifeng Luo
dblp:01/5820
· DBLP profile ↗
23ranked-venue papers
6as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Message Injection Attack on Rumor Detection under the Black-Box Evasion Setting Using Large Language ModelabstractRecent analyses have disclosed that existing rumor detection techniques, despite playing a pivotal role in countering the dissemination of misinformation on social media, are vulnerable to both white-box and surrogate-based black-box adversarial attacks. However, such attacks depend heavily on unrealistic assumptions, e.g., modifiable user data and white-box access to the rumor detection models, or appropriate selections of surrogate models, which are impractical in the real world. Thus, existing analyses fail to uncover the robustness of rumor detectors in practice. In this work, we take a further step towards the investigation about the robustness of existing rumor detection solutions. Specifically, we focus on the state-of-the-art rumor detectors, which leverage graph neural network based models to predict whether a post is rumor based on the Message Propagation Tree (MPT), a conversation tree with the post as its root and the replies to the post as the descendants of the root. We propose a novel black-box attack method, HMIA-LLM, against these rumor detectors, which uses the Large Language Model to generate malicious messages and inject them into the targeted MPTs. Our extensive evaluation conducted across three rumor detection datasets, four target rumor detectors, and three baselines for comparison demonstrates the effectiveness of our proposed attack method in compromising the performance of the state-of-the-art rumor detectors. Yifeng Luo, Yupeng Li 0001, Dacheng Wen, Liang Lan |
WWW | 1 |
| 2024 | Predicting stock market trends with self-supervised learning
Zelin Ying, Dawei Cheng, Cen Chen 0001, Xiang Li 0067, Peng Zhu 0002, Yifeng Luo |
Neurocomputing | 6 |
| 2023 | Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document UnderstandingabstractHaoli Bai, Zhiguang Liu, Xiaojun Meng, Li Wentao, Shuang Liu, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Haoli Bai, Xiaojun Meng, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang 0004, Lu Hou 0002, Jiansheng Wei, Xin Jiang 0002, Qun Liu 0001 |
ACL (1) | 6 |
| 2023 | Adversarial Multi-task Learning for Efficient Chinese Named Entity RecognitionabstractNamed entity recognition (NER) is a fundamental task for information extraction applications. NER is challenging because of semantic ambiguities in academic literature, especially for non-Latin languages. Besides word semantic information, recognizing Chinese named entities needs to consider word boundary information, as words contained in Chinese texts are not separated with spaces. Leveraging word boundary information could help to determine entity boundaries and thus improve entity recognition performance. In this article, we propose to combine word boundary information and semantic information for named entity recognition based on multi-task adversarial learning. Specifically, we learn commonly shared boundary information of entities from multiple kinds of tasks, including Chinese word segmentation (CWS), part-of-speech (POS) tagging, and entity recognition, with adversarial learning. We learn task-specific semantic information of words from these tasks and combine the learned boundary information with the semantic information to improve entity recognition with multi-task learning. We then propose a compression method based on improved clustering to accelerate the proposed model. We conduct extensive experiments on four public benchmark datasets and two private datasets, compared with state-of-the-art baseline models, and the experimental results demonstrate that our model achieves considerable performance improvements on various evaluation datasets. Peng Zhu 0002, Dawei Cheng, Fangzhou Yang, Yifeng Luo |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2022 | TaskSum: Task-Driven Extractive Text Summarization for Long News Documents Based on Reinforcement Learning
Moming Tang, Dawei Cheng, Cen Chen 0001, Yifeng Luo, Weining Qian |
DASFAA (3) | 5 |
| 2022 | Multi-scale Time Based Stock Appreciation Ranking Prediction via Price Co-movement Discrimination
Ruyao Xu, Dawei Cheng, Cen Chen 0001, Siqiang Luo, Yifeng Luo, Weining Qian |
DASFAA (3) | 5 |
| 2022 | SI-News: Integrating social information for news recommendation with attention-based graph convolutional network
Peng Zhu 0002, Dawei Cheng, Siqiang Luo, Fangzhou Yang, Yifeng Luo, Weining Qian, Aoying Zhou |
Neurocomputing | 5 |
| 2022 | Improving Chinese Named Entity Recognition by Large-Scale Syntactic Dependency GraphabstractNamed entity recognition (NER) isa preliminary task in natural language processing (NLP). Recognizing Chinese named entities from unstructured texts is challenging due to the lack of word boundaries. Even if performing Chinese Word Segmentation (CWS) could help to determine word boundaries, it is still difficult to determine which words should be clustered together for entity identification, since entities are often composed of multiple-segmented words. As dependency relationships between segmented words could help to determine entity boundaries, it is crucial to employ information related to syntactic dependency relationships to improve NER performance. In this paper, we propose a novel NER model to learn information about syntactic dependency graphs with graph neural networks, and merge learned information into the classic Bidirectional Long Short-Term Memory (BiLSTM) - Conditional Random Field (CRF) NER scheme. In addition, we extract various kinds of task-specific hidden information from multiple CWS and part-of-speech (POS) tagging tasks, to further improve the NER model. We finally leverage multiple self-attention components to integrate multiple kinds of extracted information for named entity identification. Experimental results on three public benchmark datasets show that our model outperforms the state-of-the-art baselines in most scenarios. Peng Zhu 0002, Dawei Cheng, Fangzhou Yang, Yifeng Luo, Dingjiang Huang, Weining Qian, Aoying Zhou |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Leveraging enterprise knowledge graph to infer web events' influences via self-supervised learning
Peng Zhu 0002, Dawei Cheng, Siqiang Luo, Ruyao Xu, Yifeng Luo |
J. Web Semant. | 6 |
| 2021 | Leveraging Domain Information to Classify Financial Documents via Unsupervised Graph Momentum ContrastabstractFinancial documents often contain rich domain information, such as named entities, which could be used to indicate the documents' classification categories. Existing classification methods either ignore such contained financial domain information, achieving less optimal performances, or train document representations in supervised ways, with expensive data labeling costs. In this paper, we propose to leverage domain information to improve classification performance for financial documents, via a graph representation learning model, namely G-MoCo, based on unsupervised graph momentum contrast. With G-MoCo, we could extract latent features from massive unlabeled raw data, and then further use the learned representations for document classification. Compared with the state-of-the-art baselines, representations learned by our method could improve performances by significant margins on a financial document dataset and three non-financial public graph datasets. Xueni Luo, Dawei Cheng, Haorui Ma, Mengzhen Fan, Yifeng Luo |
CIKM | 6 |
| 2021 | ZH-NER: Chinese Named Entity Recognition with Adversarial Multi-task Learning and Self-Attentions
Peng Zhu 0002, Dawei Cheng, Fangzhou Yang, Yifeng Luo, Weining Qian, Aoying Zhou |
DASFAA (2) | 4 |
| 2020 | Fusing Global Domain Information and Local Semantic Information to Classify Financial DocumentsabstractMany institutions are devoted to providing investment advising services to stock investors to help them make sound investment decisions. Industry analysts at these institutions need to analyze huge amounts of financial news documents, and yield investment advising reports to the service subscribers. Automatic document classification is required to organize collected financial news documents into pre-defined fine-grained categories, before the document analysis tasks. It is challenging to implement accurate fine-grained classification over massive financial documents, because documents from close fine-grained categories are highly semantically similar, while existing classification methods may fail to differentiate the subtle differences for documents from close fine-grained categories. In this paper, we implement a document classification framework, named GraphSEAT, to classify financial documents for a leading financial information service provider in China. Specifically, we build a heterogeneous graph to model the global structure of our targeting financial documents, where documents and financial named entities are deemed as nodes, and a document is connected to a contained named entity with an edge, and we then train a graph convolutional network (GCN) with attention mechanisms, to learn an embedding representation containing domain information for a document. We also extract semantic information from a document's word sequence with a neural sequence encoder, and finally form an overall embedding representation for a document and make the prediction, via fusing the two learned representations of the document with attention mechanisms. We perform extensive experiments on our real-world financial news dataset and three public datasets, to evaluate the performance of the document classification framework, and the experimental results demonstrate that GraphSEAT outperforms all compared eight baseline models, especially on our dataset. Mengzhen Fan, Dawei Cheng, Fangzhou Yang, Siqiang Luo, Yifeng Luo, Weining Qian, Aoying Zhou |
CIKM | 5 |
| 2020 | F-HMTC: Detecting Financial Events for Investment Decisions Based on Neural Hierarchical Multi-Label Text ClassificationabstractThe share prices of listed companies in the stock trading market are prone to be influenced by various events. Performing event detection could help people to timely identify investment risks and opportunities accompanying these events. The financial events inherently present hierarchical structures, which could be represented as tree-structured schemes in real-life applications, and detecting events could be modeled as a hierarchical multi-label text classification problem, where an event is designated to a tree node with a sequence of hierarchical event category labels. Conventional hierarchical multi-label text classification methods usually ignore the hierarchical relationships existing in the event classification scheme, and treat the hierarchical labels associated with an event as uniform labels, where correct or wrong label predictions are assigned with equal rewards or penalties. In this paper, we propose a neural hierarchical multi-label text classification method, namely F-HMTC, for a financial application scenario with massive event category labels. F-HMTC learns the latent features based on bidirectional encoder representations from transformers, and directly maps them to hierarchical labels with a delicate hierarchy-based loss layer. We conduct extensive experiments on a private financial dataset with elaborately-annotated labels, and F-HMTC consistently outperforms state-of-art baselines by substantial margins. We will release both the source codes and dataset on the first author's repository. Dawei Cheng, Fangzhou Yang, Yifeng Luo, Weining Qian, Aoying Zhou |
IJCAI | 4 |
| 2020 | A privacy-enhancing scheme against contextual knowledge-based attacks in location-based services
Jiaxun Hua, Yibin Shen, Xiuxia Tian, Yifeng Luo, Cheqing Jin |
Frontiers Comput. Sci. | 5 |
| 2018 | Towards efficiently supporting database as a service with QoS guarantees
Yifeng Luo, Junshi Guo, Jiaye Zhu, Jihong Guan, Shuigeng Zhou |
J. Syst. Softw. | 1 |
| 2017 | Supporting Cost-Efficient Multi-tenant Database Services with Service Level Objectives (SLOs)
Yifeng Luo, Junshi Guo, Jiaye Zhu, Jihong Guan, Shuigeng Zhou |
DASFAA (1) | 1 |
| 2017 | A Graph Matching Based Method for Dynamic Passenger-Centered Ridesharing
Yifeng Luo, Shuigeng Zhou, Jihong Guan |
DEXA (1) | 2 |
| 2017 | JeCache: Just-Enough Data Caching with Just-in-Time Prefetching for Big Data ApplicationsabstractBig data clusters introduce an intermediate cache layer between the computing frameworks and the underlying distributed file systems, to enable upper-level applications or end users to efficiently access big datasets in cache and effectively share them among different computing frameworks. As caches are shared by multiple applications or end users, directly applying existing on-demand caching strategies will result in intense conflicts, when big datasets are cached as a whole. Meanwhile, big data applications usually involve massive numbers of file scans, cached-in data blocks may have little chance of being accessed before they are cached out to make way for other on-demand data blocks. Thus, it is unwise to cache data blocks long before they are actually accessed. In this paper, we propose a novel just-enough big data caching scheme for just-in-time block prefetching to improve the cache effectiveness of big data clusters. With just-in-time block prefetching, a block is cached in just before the task begins to process the block, rather than being cached in along with other blocks of the same dataset being processed. We monitor block accesses to measure the average processing time of data blocks, and then estimate the minimal number of blocks that should be kept in cache for a big dataset, so that the speed of data processing matches with that of data prefetching, and each upper-level task can obtain its input blocks from cache just in time. Our experimental results show that the proposed cache method can restrain over-requirement of cache resources in big data applications, and provides the same performance improvement as when all data blocks are cached. Yifeng Luo, Shuigeng Zhou |
ICDCS | 1 |
| 2016 | Wireless backhaul resource allocation and user-centric clustering in ultra-dense wireless networksabstractWireless backhaul is a promising technology to lower the cost for connecting a large number of densely deployed access points in the emerging fifth generation wireless system. In this study, the authors consider the joint optimisation of resource allocation in wireless backhaul links and user‐centric clustering in the access links. The objective is to maximise the weighted sum rate of all users under the backhaul resource constraints. To solve this intractable problem, they first adopt the ℓ 1 ‐norm approximation approach to convert the non‐convex backhaul constraints into the solvable forms. They then reformulate the objective function as a set of second‐order cone constraints based on the sequential parametric convex approximation method. An iterative algorithm is proposed to solve the transformed problem based on its special property. Simulation results show that the proposed algorithm outperforms other existing schemes under different network settings. Cunqing Hua, Yifeng Luo, Huibo Liu |
IET Commun. | 2 |
| 2015 | LAYER: A cost-efficient mechanism to support multi-tenant database as a service in cloud
Yifeng Luo, Shuigeng Zhou, Jihong Guan |
J. Syst. Softw. | 1 |
| 2014 | Distributed Spatial Keyword Querying on Road NetworksabstractSpatial-keyword queries on road networks are receiving in-creasing attention with the prominence of location-based services. There is a growing need to handle queries on road networks in distributed environments because a large net-work is typically distributed over multiple machines and it will improve query throughput. However, all the existing work on spatial keyword queries is based on a centralized setting. In this paper, we develop a distributed solution to answering spatial keyword queries on road networks. Exam-ple queries include “find locations near a supermarket and a hospital, ” and “find Chinese restaurants within 500 meters from my current location. ” We define an operation for an-swering such queries and reduce the problem of answering a query into computing a function of such operations. We pro-pose a new distributed index that enables each machine to independently evaluate the operation on its network frag-ment in a distributed setting. We theoretically prove the space optimality of the proposed index technique. We con-duct experiments with a distributed setting. Experimen-tal results demonstrate the promising performance of our method. Siqiang Luo, Yifeng Luo, Shuigeng Zhou, Gao Cong, Jihong Guan |
EDBT | 2 |
| 2013 | A RAMCloud Storage System based on HDFS: Architecture, implementation and evaluation
Yifeng Luo, Siqiang Luo, Jihong Guan, Shuigeng Zhou |
J. Syst. Softw. | 1 |
| 2012 | DISKs: A System for Distributed Spatial Group Keyword Search on Road NetworksabstractQuery (e.g., shortest path) on road networks has been extensively studied. Although most of the existing query processing approaches are designed for centralized environments, there is a growing need to handle queries on road networks in distributed environments due to the increasing query workload and the challenge of querying large networks. In this demonstration, we showcase a distributed system calledDISKs(DIstributedSpatialKeywordsearch) that is capable of efficiently supporting spatial group keyword search (S-GKS) on road networks. Given a group of keywordsXand a distancer, an SGKS returns locations on a road network, such that for each returned locationp, there exists a set of nodes (on the road network), which are located within a network distancerfrompand collectively containsX. We will demonstrate the innovative modules, performance and interactive user interfaces of DISKs. Siqiang Luo, Yifeng Luo, Shuigeng Zhou, Gao Cong, Jihong Guan |
Proc. VLDB Endow. | 2 |