VLDB 2026 Research / reviewers in the wild / expert
Haipeng Cheng
dblp:67/4193
· DBLP profile ↗
3ranked-venue papers
2as first author
1since 2021 · last 2024
0009-0007-8723-4510ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
1 paper |
Internet architecture and protocols · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Internet architecture and protocols › packet processing
packet classification |
0.1 | 1 | 2008 | Scalable packet classification using interpreting: a cross-platform multi-core solution · PPoPP 2008 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.0 | 1 | 2008 | Scalable packet classification using interpreting: a cross-platform multi-core solution · PPoPP 2008 |
Methods — techniques the papers use, named apart from their topics
interpreter-based execution · 0.2data partitioning · 0.2CISC-style instruction encoding · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Are Dense Retrieval Models Few-Shot Learners?
Haipeng Cheng, Si Sun, Benchang Zheng |
PRICAI (2) | 1 |
| 2009 | Practice of parallelizing network applications on multi-core architecturesabstractThe industry wide shift to multi-core architectures arouses great interests in parallelizing sequential applications. However, it is very difficult to parallelize fine-grained applications for multi-core architectures due to insufficient hardware support of fast communication and synchronization. Fortunately, network applications can be decomposed into pipelined structures that are amenable to streaming based parallel processing. To realize the potential of pipelining on multi-core architectures, it requires reevaluating the basic tradeoffs in parallel processing, including the ones between load balance and data locality and between general lock mechanisms and special lock-free data structures. This paper presents the practice of building a high-performance multi-core based network processing platform in which connection-affinity and lock-free design principles are applied effectively for better data locality and faster core-to-core synchronization and communication.We parallelize a complete Layer 2 to Layer 7 (L2-L7) network processing system on an Intel Core 2 Quad processor, including a TCP/IP stack based on Libnids (L2-L4) and a port-independent protocol identification engine by deep packet inspection (L7+). Furthermore, we develop a compiling method to transform sequential network applications to parallel ones to enable those applications to run on multi-core architectures. Our experience suggests that (1) fine-grained pipelining can be a good software solution for parallelizing network applications on multi-core architectures if connection-affinity and lock-free are used as the first design principles; (2) a delicate partitioning scheme is required to map pipelined structures onto specific multi-core architecture; (3) an automatic parallelization approach can work if domain knowledge is considered in the parallelizing process. Our multi-core based network processing platform can deliver not only 6Gbps processing speed for large packet sizes but also more challenging 2Gbps speed for smaller packets. Junchang Wang, Haipeng Cheng, Bei Hua, Xinan Tang |
ICS | 2 |
| 2008 | Scalable packet classification using interpreting: a cross-platform multi-core solutionabstractPacket classification is an enabling technology to support advanced Internet services. It is still a challenge for a software solution to achieve 10Gbps (line-rate) classification speed. This paper presents a classification algorithm that can be efficiently implemented on a multi-core architecture with or without cache. The algorithm embraces the holistic notion of exploiting application characteristics, considering the capabilities of the CPU and the memory hierarchy, and performing appropriate data partitioning. The classification algorithm adopts two stages: searching on a reduction tree and searching on a list of ranges. This decision is made based on a classification heuristic: the size of the range list is limited after the first stage search. Optimizations are then designed to speed up the two-stage execution. To exploit the speed gap (1) between the CPU and external memory; (2) between internal memory (cache) and external memory, an interpreter is used to trade the CPU idle cycles with demanding memory access requirements. By applying the CISC style of instruction encoding to compress the range expressions, it not only significantly reduces the total memory requirement but also makes effective use of the internal memory (cache) bandwidth. We show that compressing data structures is an effective optimization across the multi-core architectures. Haipeng Cheng, Bei Hua, Xinan Tang |
PPoPP | 1 |