EDBT 2026 Demo / reviewers in the wild / expert
Jun Chai
dblp:64/7601
· DBLP profile ↗
7ranked-venue papers
2as first author
1since 2021 · last 2023
0000-0002-7133-3481ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 100% |
Topics — the 1 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU computing › GPU video coding
GPU-accelerated video encoding |
0.1 | 1 | 2011 | High-efficient software parallel CAVLC encoder based on programmable stream processor · ACM Multimedia 2011 |
Methods — techniques the papers use, named apart from their topics
programmable stream processor · 0.2block-based parallel processing · 0.2CUDA · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | ST-Bikes: Predicting Travel-Behaviors of Sharing-Bikes Exploiting Urban Big DataabstractWith the development of the modern smart city, sharing-bikes require behaviors prediction for grid-level areas which is essential for intelligent transportation systems. A model which can predict bike sharing demand behaviours accurately can allocate sharing-bikes in advance to satisfy travel demands alongside saving energy, reducing traffic, cutting down waste for those sharing-bikes companies putting excessive sharing-bikes in unsaturated demand areas. In this paper, we abandon the traditional time series prediction method and use a more efficient deep learning method to solve the traffic forecasting problem. Moreover, instead of considering spatial relation and temporal relation relatively, we produced a deep multi-view spatial-temporal network to combine them into one prediction model framework. In the experimental section, we investigate in the experiment on enormous amount of real sharing-bikes application use data in the core region of Beijing to test the performance of the model framework with a 1 km$\times $1 km grid-level scale and compare it with other existing machine learning approaches and prediction models. And the 4G/5G/6G communication technology facilitate the real-time control of the space-time locations of sharing bikes dynamically. Thus, it provides the basis for high-frequency analysis of space-time patterns, especially supported by the 6G large-scale application in the future. Jun Chai, Hongwei Fan, Le Zhang 0004, Bing Guo 0003, Yawen Xu |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | Communication-hiding programming for clusters with multi-coprocessor nodesabstractSummary Future exascale systems are expected to adopt compute nodes that incorporate many accelerators. To shed some light on the upcoming software challenge, this paper investigates the particular topic of programming clusters that have multiple Xeon Phi coprocessors in each compute node. A new offload approach is considered for intra‐node communication, which combines Intel's APIs of coprocessor offload infrastructure (COI) and symmetric communication interface (SCIF) for achieving low latency. While the conventional pragma‐based offload approach allows simpler programming, the COI‐SCIF approach has three advantages in (1) lower overhead associated with launching offloaded code, (2) higher data transfer bandwidths, and (3) more advanced asynchrony between computation and data movement. The low‐level COI‐SCIF approach is also shown to have benefits over the MPI‐OpenMP counterpart, which belongs to the symmetric usage mode. Moreover, a hybird programming strategy based on COI‐SCIF is presented for joining the computational force of all CPUs and coprocessors, while realizing communication hiding. All the programming approaches are tested by a real‐world 3D application, for which the COI‐SCIF‐based approach shows a performance advantage on Tianhe‐2. Copyright © 2015 John Wiley & Sons, Ltd. Xinnan Dong, Mei Wen, Jun Chai, Xing Cai, Mandan Zhao, Chunyuan Zhang |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Parallel performance modeling of irregular applications in cell-centered finite volume methods over unstructured tetrahedral meshes
Johannes Langguth, Nan Wu 0003, Jun Chai, Xing Cai |
J. Parallel Distributed Comput. | 3 |
| 2014 | Utilizing Multiple Xeon Phi Coprocessors on One Compute Node
Xinnan Dong, Jun Chai, Mei Wen, Nan Wu 0003, Xing Cai, Chunyuan Zhang, Zhaoyun Chen |
ICA3PP (2) | 2 |
| 2013 | Resource-efficient utilization of CPU/GPU-based heterogeneous supercomputers for Bayesian phylogenetic inference
Jun Chai, Huayou Su, Mei Wen, Xing Cai, Nan Wu 0003, Chunyuan Zhang |
J. Supercomput. | 1 |
| 2012 | Parallelization Design of Irregular Algorithms of Video Processing on GPUsabstractIn this paper, we present the parallelization design consideration for irregular algorithms of video processing on GPUs. Enrich parallelism can be exploited by scheduling the processing order or making a tradeoff between performance and parallelism for irregular algorithms (such as CAVLC and deblocking filter). We implement a component-oriented CAVLC encoder and a direction-oriented deblocking filter on GPUs. The experiment results show that, compared with the implementation on CPU, the optimized parallel methods achieve high performance in term of speedup ratio from 63 to 44, relatively for deblocking filter and CAVLC. It shows that the rich parallelism is one of the most important factors to gain high performance for irregular algorithms based on GPUs. In addition, it seems that for some irregular kernels, the number of SM of GPU is more important to the performance than the computation capability. Huayou Su, Jun Chai, Mei Wen, Ju Ren 0002, Chunyuan Zhang |
ICME | 2 |
| 2011 | High-efficient software parallel CAVLC encoder based on programmable stream processorabstractThis article presents an efficient software parallel CAVLC encoder based on programmable stream processors (Storm- SP16 and GPU). For static processor Storm SP16, a block-based 16 ways parallel CAVLC is presented with streaming processing. A component-oriented CAVLC encoder is proposed aiming at dynamic stream processor GPU. Experiments results show that, compared to the CPU version, more than 70 times of speedup can be obtained for the CAVLC based on Storm and over 50 times for GPU-based component-oriented CAVLC encoder. The throughput of the presented CAVLC encoder is more than 10 times higher over that of published software CAVLC encoders on DSP and multi-core platforms. Huayou Su, Chunyuan Zhang, Jun Chai, Mei Wen, Nan Wu 0003, Ju Ren 0002 |
ACM Multimedia | 3 |