EDBT 2026 Demo / reviewers in the wild / expert
Chengwen Wu
dblp:22/5285
· DBLP profile ↗
5ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0002-7233-7062ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Graph data management › graph processing
graph processing systems |
0.4 | 1 | 2019 | Redio: Accelerating Disk-Based Graph Processing by Reducing Disk I/Os · IEEE Trans. Computers 2019 |
Storage systems › out-of-core computation
disk-based graph processing |
0.4 | 1 | 2019 | Redio: Accelerating Disk-Based Graph Processing by Reducing Disk I/Os · IEEE Trans. Computers 2019 |
Storage systems › i/o optimization
disk i/o reduction |
0.4 | 1 | 2019 | Redio: Accelerating Disk-Based Graph Processing by Reducing Disk I/Os · IEEE Trans. Computers 2019 |
Storage systems › flash and SSD
solid-state drive |
0.1 | 1 | 2019 | Redio: Accelerating Disk-Based Graph Processing by Reducing Disk I/Os · IEEE Trans. Computers 2019 |
Methods — techniques the papers use, named apart from their topics
selective scheduling · 0.8indexed bitmap · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Boafft: Distributed Deduplication for Big Data Storage in the CloudabstractAs data progressively grows within data centers, the cloud storage systems continuously facechallenges in saving storage capacity and providing capabilities necessary to move big data within an acceptable time frame. In this paper, we present the Boafft, a cloud storage system with distributed deduplication. The Boafft achieves scalable throughput and capacity usingmultiple data servers to deduplicate data in parallel, with a minimal loss of deduplication ratio. Firstly, the Boafft uses an efficient data routing algorithm based on data similarity that reduces the network overhead by quickly identifying the storage location. Secondly, the Boafft maintains an in-memory similarity indexing in each data server that helps avoid a large number of random disk reads and writes, which in turn accelerates local data deduplication. Thirdly, the Boafft constructs hot fingerprint cache in each data server based on access frequency, so as to improve the data deduplication ratio. Our comparative analysis with EMC's stateful routing algorithm reveals that the Boafft can provide a comparatively high deduplication ratio with a low network bandwidth overhead. Moreover, the Boafft makes better usage of the storage space, with higher read/write bandwidth and good load balance. Shengmei Luo, Guangyan Zhang, Chengwen Wu, Samee Ullah Khan, Keqin Li 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2019 | Redio: Accelerating Disk-Based Graph Processing by Reducing Disk I/OsabstractDisk-based graph systems store part or all of graph data on external devices like hard drives or SSDs, achieving scalability without excessive hardware. However, massive expensive disk I/Os remain the major performance bottleneck of disk-based graph processing. In this paper, we propose Redio, a new approach to accelerating disk-based graph processing by reducing disk I/Os. First, Redio observes that it is feasible to accommodate all vertex states in main memory and this can eliminate almost all vertex-related disk I/Os. Second, Redio introduces a dynamic selective scheduling scheme to identify inactive edges in each iteration and skip them when and only when such skipping can bring performance benefit. To improve its effectiveness, Redioin corporates a compact edge storage to improve data locality and an indexed bitmap to minimize its memory and computation overheads. We have implemented a single-node prototype for Redio under the edge-centric computation model. Extensive experiments show that Redio consistently outperforms well-known edge-centric disk-based systems in all experiments, delivering an average speedup of$4.33\times$on HDDs and$5.33\times$on SSDs over the fastest among them (i.e., GridGraph). Experimental results also show that Redio delivers an average speedup of$3.13\times$on HDDs and$1.28\times$on SSDs over the fastest among representative vertex-centric disk-based systems (i.e., FlashGraph). Chengwen Wu, Guangyan Zhang, Yang Wang 0009, Xinyang Jiang |
IEEE Trans. Computers | 1 |
| 2017 | Building a fault tolerant framework with deadline guarantee in big data stream computing environments
Dawei Sun 0001, Guangyan Zhang, Chengwen Wu, Keqin Li 0001 |
J. Comput. Syst. Sci. | 3 |
| 2016 | Rethinking Computer Architectures and Software Systems for Phase-Change MemoryabstractWith dramatic growth of data and rapid enhancement of computing powers, data accesses become the bottleneck restricting overall performance of a computer system. Emerging phase-change memory (PCM) is byte-addressable like DRAM, persistent like hard disks and Flash SSD, and about four orders of magnitude faster than hard disks or Flash SSDs for typical file system I/Os. The maturity of PCM from research to production provides a new opportunity for improving the I/O performance of a system. However, PCM also has some weaknesses, for example, long write latency, limited write endurance, and high active energy. Existing processor cache systems, main memory systems, and online storage systems are unable to leverage the advantages of PCM, and/or to mitigate PCM’s drawbacks. The reason behind this incompetence is that they are designed and optimized for SRAM, DRAM memory, and hard drives, respectively, instead of PCM memory. There have been some efforts concentrating on rethinking computer architectures and software systems for PCM. This article presents a detailed survey and review of the areas of computer architecture and software systems that are oriented to PCM devices. First, we identify key technical challenges that need to be addressed before this memory technology can be leveraged, in the form of processor cache, main memory, and online storage, to build high-performance computer systems. Second, we examine various designs of computer architectures and software systems that are PCM aware. Finally, we obtain several helpful observations and propose a few suggestions on how to leverage PCM to optimize the performance of a computer system. Chengwen Wu, Guangyan Zhang, Keqin Li 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2006 | An Index Scheme for XML Documents Based on Relationship JoinsabstractXML is rapidly emerging as a standard for information storage, representation and exchange on the Web. How to rapidly search and query XML documents efficiently has received many attentions in resent research. However, current querying schemes of XML documents typically involve in both node content and tree structural information, which may limit efficiency when facing the application that the tree structural information is more complicated than the tree node itself. In this paper, we propose the node relationships joins algorithms that utilize available indexes mainly on tree structural information. The relationships join algorithms work perfectly especially for searching paths that are very long or whose lengths are unknown. Experimental results from our prototype system implementation highlight the correctness and efficiency of our solution Chengwen Wu, Jinxiang Dong, Gang Chen 0001, Lihua Yu |
CSCWD | 1 |