VLDB 2026 Research / reviewers in the wild / expert
Zhijia Chen
dblp:20/5211
· DBLP profile ↗
16ranked-venue papers
8as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Demand-Oriented Route Recommendation for Shared Mobility Services
Zhijia Chen, Chen Zhang 0013, Peng Cheng 0003, Libin Zheng 0001, Jian Yin 0001 |
DASFAA (5) | 1 |
| 2025 | ComCrawler: General Crawling Solution for Aticle Comments
Zhijia Chen, Weiyi Meng, Eduard C. Dragut |
EDBT | 1 |
| 2024 | SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsabstractScientific information extraction (SciIE) is critical for converting unstructured knowledge from scholarly articles into structured data (entities and relations).Several datasets have been proposed for training and validating SciIE models.However, due to the high complexity and cost of annotating scientific texts, those datasets restrict their annotations to specific parts of paper, such as abstracts, resulting in the loss of diverse entity mentions and relations in context.In this paper, we release a new entity and relation extraction dataset for entities related to datasets, methods, and tasks in scientific articles.Our dataset contains 106 manually annotated full-text scientific publications with over 24k entities and 12k relations.To capture the intricate use and interactions among entities in full texts, our dataset contains a finegrained tag set for relations.Additionally, we provide an out-of-distribution test set to offer a more realistic evaluation.We conduct comprehensive experiments, including state-of-the-art supervised models and our proposed LLM baselines, and highlight the challenges presented by our dataset, encouraging the development of innovative models to further the field of SciIE.1 Zhijia Chen, Huitong Pan, Cornelia Caragea, Longin Jan Latecki, Eduard C. Dragut |
EMNLP | 2 |
| 2024 | Longer Pick-Up for Less Pay: Towards Discount-Based Mobility ServicesabstractWith the rapid development of mobile Internet technology, on-demand car-hailing services have become essential for people's daily commuting. Order dispatch is a critical problem in on-demand car-hailing services. However, in most existing works, the service provider is asked to set a unified pick-up distance to prevent long waiting time for requesters. Indeed, different requesters have different tolerance for pick-up distance, and some requesters may accept longer pick-ups if offered discounts for payment. Regarding this fact, we formulate discount-based order dispatch as a coupling of two subproblems, discount determination and order dispatch, aiming to dispatch more orders and thereby more platform profits. We propose customized methods to solve the problems for shared and non-shared mobility services, respectively. We also conduct extensive experiments to evaluate the effectiveness and efficiency of our proposed methods on a real dataset, which shows that our methods can achieve 170% improvements in non-shard services and 43% improvements in ridesharing services on average in terms of attained profit compared to the widely adopted baselines. Wanyi Xie, Zhijia Chen, Chen Zhang 0013, Libin Zheng 0001, Peng Cheng 0003, Jian Yin 0001, Xuemin Lin 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | COIN - an Inexpensive and Strong Baseline for Predicting Out of Vocabulary Word EmbeddingsabstractSocial media is the ultimate challenge for many natural language processing tools. The constant emergence of linguistic constructs challenge even the most sophisticated NLP tools. Predicting word embeddings for out of vocabulary words is one of those challenges. Word embedding models only include terms that occur a sufficient number of times in their training corpora. Word embedding vector models are unable to directly provide any useful information about a word not in their vocabularies. We propose a fast method for predicting vectors for out of vocabulary terms that makes use of the surrounding terms of the unknown term and the hidden context layer of the word2vec model. We propose this method as a strong baseline in the sense that 1) while it does not surpass all state-of-the-art methods, it surpasses several techniques for vector prediction on benchmark tasks, 2) even when it underperforms, the margin is very small retaining competitive performance in downstream tasks, and 3) it is inexpensive to compute, requiring no additional training stage. We also show that our technique can be incorporated into existing methods to achieve a new state-of-the-art on the word vector prediction problem. Andrew T. Schneider, Lihong He 0001, Zhijia Chen, Arjun Mukherjee, Eduard C. Dragut |
COLING | 3 |
| 2022 | Web Record Extraction with InvariantsabstractWeb records are structured data on a Web page that embeds records retrieved from an underlying database according to some templates. Mining data records on the Web enables the integration of data from multiple Web sites for providing value-added services. Most existing works on Web record extraction make two key assumptions: (1) records are retrieved from databases with uniform schemas and (2) records are displayed in a linear structure on a Web page. These assumptions no longer hold on the modern Web. A Web page may present records of diverse entity types with different schemas and organize records hierarchically, in nested structures, to show richer relationships among records. In this paper, we revisit these assumptions and modify them to reflect Web pages on the modern Web. Based on the reformulated assumptions, we introduce the concept of invariant in Web data records and propose Miria ( Mi ning r ecord i nvari a nt), a bottom-up, recursive approach to construct the Web records from the invariants. The proposed approach is both effective and efficient, consistently outperforming the state-of-the-art Web record extraction methods on modern Web pages. Zhijia Chen, Weiyi Meng, Eduard C. Dragut |
Proc. VLDB Endow. | 1 |
| 2019 | Internet Routing and Non-monotonic Reasoning
Anduo Wang, Zhijia Chen |
LPNMR | 2 |
| 2009 | How Scalable Could P2P Live Media Streaming System Be with the Stringent Time Constraint?abstractThe peer-to-peer (P2P) live video streaming system has been demonstrated to have great potential in the public Internet; the large-scale deployment of such systems, however, critically relies on how effective they can deal with the high dynamics encountered, in particular during flash crowd. The rationale behind is that the scaling in P2P live video streaming systems is heavily determined by the timing requirement that streaming applications demand. In this paper, we present an analytical and experimental study on the inherent relationship between the time constraint and the system scale. We develop a generic model for P2P live video streaming that focuses on the peer joining process during flash crowd. We first illustrate that the simple notion of "demand vs. supply" model is insufficient in describing the system scale. By computing the peer start-up time distribution, we demonstrate that the scale is affected by several key factors, especially the peer uploading capacity and the initial system size. We further show the scale is essentially bounded by the timing requirement and the system's capability to accommodate flash crowd is subject to a maximum limit. Zhijia Chen, Bo Li 0001, Gabriel Yik Keung, Chuang Lin 0002, Yuanzhuo Wang |
ICC | 1 |
| 2009 | Towards a universal friendly peer-to-peer media streaming: metrics, analysis and explorationsabstractPeer-to-peer (P2P) paradigm has provided a disruptive market opportunity to define cost-effective multimedia streaming services, but at the same time network-oblivious P2P applications have been posing substantial technical and social challenges on network efficiency, operator economics and user performance. While end users are concerned with quality upgrade, Internet content provider (ICP) considers more on service scale and Internet service provider (ISP) focuses on operating cost. In taming P2P for a more friendly large-scale application, this paper provides a framework of evaluating the P2P media streaming application performance from perspectives of all entities involved, that is ISP, ICP and end users. Three-level performance metrics are defined, essential concerns of each party are theoretically quantified and bottlenecks in affecting quality service are identified. In handling tussles between P2P performance against ISP traffic, system scale against cost and user QoS against Security, the authors present proposals in defining an unprecedented friendly and cost-effective P2P streaming application to achieve the ideal philosophy of ‘more users=better performance+lower cost’. Based on the explorations in academy and industry, the authors envision that a large-scale streaming system will be built with a synergy of P2P and content distribution networks (CDN), and explore the feasibility of a general peer–server–peer (PSP) structure on the basis of our evaluation framework. With our analytical study and industrial deployment, this paper captures a certain essence of deploying large-scale P2P streaming application from a commercial and realistic point of view and suggests many avenues for addressing the emerging tensions between P2P application and network operators. Zhijia Chen, Chuang Lin 0002, Yang Chen 0001, Mark Feng |
IET Commun. | 1 |
| 2009 | An Improved Markov Model for IEEE 802.15.4 Slotted CSMA/CA Mechanism
Hao Wen 0014, Chuang Lin 0002, Zhijia Chen, Tao He 0008, Eryk Dutkiewicz |
J. Comput. Sci. Technol. | 3 |
| 2008 | Performance Analysis and Industrial Practice of Peer-Assisted Content Distribution Network for Large-Scale Live Video StreamingabstractRecently efficient and scalable live video streaming system over the Internet has become a hot topic. In order to improve the system performance metrics, such as startup delay, source-to-end delay, playback continuity and scalability, many previous works developed two successful cases of content distribution network (CDN) and peer-to-peer (P2P) Network for the design of large-scale live video streaming systems, but no single one has yet delivered both the scale and service quality. To combine the advantages of CDN and P2P network has been considered as a feasible orientation for large-scale video stream delivering. In this paper, we propose a peer-assisted content distribution network, i.e. PACDN, which borrows the mesh-based P2P ideas into the traditional CDN to enhance the performance and scalability. The basic features of PACDN include: 1) To meet the real time requirement of live video stream service, i.e. to ensure that the video stream could be continuously and stably delivered from the source to each edge server for offering good QoS to different regions clients, the placement edge servers and source streaming server(s) build a hierarchical multi-tree based and in-hierarchy peer-assisted overlay, which is optimized according to the knowledge of underlying physical topology. This scheme in the design is called "server side peer-assisted". 2) To enhance the system scalability and reduce the deployment cost, clients and edge servers construct a Client/Server based and P2P network assisted overlay with the increasing of viewers, which is called "client side peer-assisted" in this design. We compare the inner performance of PACDN with existing approaches based on comprehensive simulations and analysis. The results show that our proposed design outperforms previous systems in the service quality and scalability. PACDN has been implemented as an Internet live video streaming service and it was successfully deployed for broadcasting many important live programs in China in 2007. The industrial experiences prove that this design is scalable and reliable. We believe that the wide deployment of PACDN and its further development will soon benefit many more Internet users. Xuening Liu, Chuang Lin 0002, Zhijia Chen |
AINA | 5 |
| 2008 | Experimental Analysis of Super-Seeding in BitTorrentabstractWith the popularity of BitTorrent, improving its performance has been an active research area. Super-seeding, a special upload policy for initial seeds, improves the efficiency in producing multiple seeds and reduces the uploading cost of the initial seeders. However, the overall benefit of super seeding remains a question. In this paper, we conduct an experimental study over the performance of super-seeding scheme of BitTornado. We attempt to answer the following questions: whether and how much super-seeding saves uploading cost, whether the download time of all peers is decreased by super-seeding, and in which scenario super-seeding performs worse. With varying seed bandwidth and peer behavior, we analyze the overall download time and upload cost of super seeding scheme during random period tests over 250 widely distributed PlanetLab nodes. The results show that benefits of super-seeding depend highly on the upload bandwidth of the initial seeds and the behavior of individual peers. Our work not only provides reference for the potential adoption of super-seeding in BitTorrent, but also much insights for the balance of enhancing Quality of Experience (QoE) and saving cost for a large-scale BitTorrent-like P2P commercial application. Zhijia Chen, Yang Chen 0001, Chuang Lin 0002, Vaibhav Nivargi |
ICC | 1 |
| 2008 | TrustStream: A Secure and Scalable Architecture for Large-Scale Internet Media StreamingabstractTo effectively address the explosive growth of multimedia applications over the Internet, a large-scale media streaming system has to fully take into account the issues of security, quality of service (QoS), scalability, and heterogeneity. However, current streaming solutions do not address all these challenges simultaneously. To address this limitation, this paper proposes a secure and high-performance streaming system called TrustStream, which combines the best features of scalable coding, content distribution network (CDN) and peer-to-peer (P2P) networks to achieve unprecedented security, scalability, heterogeneity, and certain QoS simultaneously under a unified architecture. In this architecture, raw video is encoded into two layers, namely, the base layer, which contains the most critical media content and is transmitted through a CDN-featured single-source multi-receiver (S-M) P2P network to guarantee a minimal level of quality, and the enhancement layer, which is transmitted in a pure multisource multi-receiver (M-M) P2P framework to achieve maximum scalability and bandwidth utilization. Heterogeneity is therefore addressed by delivering only the layers that a receiver is able to manage. Security is provided by combining our key distribution mechanism and key-embedding scheme under our proposed S-M P2P topology. We have implemented TrustStream system over the Internet. Deployed by ChinaCache, the largest CDN provider in China, TrustStream has broadcasted several popular live video programs over the Internet. The experimental results demonstrate the advantages and effectiveness of our architecture and system. Chuang Lin 0002, Qian Zhang 0001, Zhijia Chen, Dapeng Oliver Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2007 | 3D-wavelet based Secure and Scalable Media Streaming in a Centralcontrolled P2P FrameworkabstractTo meet with the ever-increasing needs of large-scale multimedia applications, a streaming media system has to be both secure and scalable. Conventional P2P technology used in streaming media could solve the problem of scalability and bandwidth bottleneck of traditional CIS architecture but still have problems in handling security for losing central manageability and robustness. Therefore, to handle security and scalability issues as a whole, in this paper we present a novel secure and scalable streaming media scheme in a central-controlled P2P framework. By firstly adopting 3D-wavelet coding in P2P streaming, we encode the raw data into different layers and separate security management from data transmission by transmitting the layer with most priority in CIS network to guarantee quality and conduct security management while transmitting the lower priority content layers in the pure P2P network to promote scalability .In our implementation, we specify the 3D-wavelet coding for P2P streaming, and design our handshaking and streaming process, load-balancing gossip-based management protocol for P2P peers. Our experimental results demonstrate our scalable framework exceed CIS streaming and meanwhile achieve better security with accepted encoding/decoding overheads over pure P2P media streaming. Zhijia Chen, Chuang Lin 0002, Lu Ai |
AINA | 1 |
| 2007 | Towards a Trustworthy and Controllable Peer-Server-Peer Media Streaming: An Analytical Study and An Industrial PerspectiveabstractPeer-to-peer technology gives novel opportunities to define a cost-effective multimedia streaming application, but at the same time, it brings a set of technical challenges due to its dynamic and heterogeneous nature. To guarantee QoS and facilitate management in large scale high-performance media streaming, we extend the current P2P networking towards a novel Peer-Server-Peer (PSP) architecture for media streaming, in which carefully-deployed servers form a trustworthy and controllable overlay network to stream P2P cluster peers. An analytical model is presented to calculate the quality of experience (QoE) and mapping QoS parameters to prove the effectiveness of Peer-Server-Peer streaming. Joint with the efforts in industry, we explore the feasibility of PSP streaming in the historical context of "demand economy" for media streaming and "best effort" Internet. The value of this paper lies not in its analytical study of this promising PSP concept with practical implementation but also its insight industrial perspective to attract further application. Zhijia Chen, Chuang Lin 0002, Xuening Liu, Yang Chen 0001 |
GLOBECOM | 1 |
| 2007 | PMTA: Potential-Based Multicast Tree Algorithm with Connectivity Restricted HostsabstractA large number of overlay protocols have been developed, almost all of which assume each host has two-way communication capability. However, this does not hold as the deployment of firewalls and Network Address Translators (NAT) is widespread in the current Internet, which is a challenge to the design and implementation of overlay models and protocols. In this paper, we present Potential-based Multicast Tree Algorithm (PMTA) to enhance the multicast tree construction in presence of connectivity restricted hosts. We evaluate PMTA and previous multicast tree protocols based on real Internet end-to-end delay datasets. According to evaluation results, PMTA outperforms those protocols in terms of all metrics. PMTA reduces ARDP by 26%, and it also results in 23%-54% reduction in average overlay latencies. As the results suggest, PMTA can build efficient and effective multicast tree and is suitable for Internet multicast applications in the presence of connectivity restricted hosts. Xiaohui Shi, Yang Chen 0001, Guohan Lu, Beixing Deng, Xing Li 0001, Zhijia Chen |
GLOBECOM | 6 |