VLDB 2026 Research / reviewers in the wild / expert
Yujie Cai
dblp:147/0048
· DBLP profile ↗
8ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Lossless Compression Algorithm with Hardware Implementation for Dynamic Vision SensorabstractNowadays, with the increase resolution of Dynamic Vision Sensor (DVS), efficient compression algorithm for event stream is needed urgently. Conventional DVS system encodes event data in address event representation (AER) for output while ignores the data redundancy imposed by the correlation of events. To address this challenge, this paper first analyzes the spatiotemporal characteristics of event stream and the impact of readout circuits. Based on the analysis, the context-based encoding strategies for spatial address, timestamp and polarity of events are proposed respectively with the consideration of data flow in DVS hardware. Besides, the hardware architecture with high parallelism is presented to implement the compression algorithm, which achieves high throughput at an affordable cost. The hardware is implemented in the 55nm process as part of a 512x512 resolution DVS. The experimental results demonstrate that our methods achieves higher average compression ratio compared to conventional and DVS-specific coding algorithms. Zewei Ding, Shangmei Wang, Yujie Cai, Xiaoyang Zeng, Wenhong Li, Mingyu Wang 0001 |
ISCAS | 3 |
| 2024 | FSS: algorithm and neural network accelerator for style transfer
Yi Ling, Yujie Cai, Zhaojie Li, Wenhong Li, Xiaoyang Zeng |
Sci. China Inf. Sci. | 3 |
| 2023 | Decomposing Synthesized Strategies for Reactive Multi-agent Reinforcement Learning
Chenyang Zhu 0001, Yujie Cai, Fang Wang 0010 |
TASE | 3 |
| 2022 | Efficient Reinforcement Learning with Generalized-Reactivity SpecificationsabstractReinforcement learning has been used to solve sequential decision-making problems in intelligent systems. However, current RL approaches suffer from slow convergence and reward sparsity, and its reward mechanism is challenging to deal with complex task specifications. As one of the software engineering practices, temporal logic can describe nonMarkovian task specifications, the synthesized strategy of which could be used as a priori knowledge to train the agents to interact with the environment efficiently. This paper considers the intelligent agent reacts to the environment with a high-level reactive temporal logic specification called Generalized Reactivity of rank 1 (GR(1)). We first use the synthesized strategy of GR(1) to construct the Markov Decision Process with a potentialbased reward machine, which integrates the environment with high-level reactive temporal specifications. Then we developed a topological-sort-based reward shaping approach to calculate the potential functions of the reward machine, based on which we used Q-learning to train the agents. Experiments on multitask learning show that the proposed approach outperforms the state-of-art algorithms in learning rate and optimal rewards. Also, compared with the value-iteration-based reward shaping approaches, our topological-sort-based reward shaping approach could handle the cases where the synthesized strategies are in the form of directed cyclic graphs. Chenyang Zhu 0001, Yujie Cai, Can Hu, Jia Bi |
APSEC | 2 |
| 2022 | A 3.1 Gbin/s advanced entropy coding hardware design for AVS3abstractAVS3 is a newly proposed video coding standard by the Audio Video coding Standard Workgroup, demonstrating higher compression efficiency than the High Efficiency Video Coding standard. Advanced entropy coding is one of the performance bottlenecks of the AVS3 standard video encoder due to the strong data dependency in its arithmetic coding process. A novel arithmetic encoding hardware structure is presented in this paper, and as we know, this is the first paper on AVS3 AEC hardware implementation. Firstly, we select and apply the typical optimization schemes adopted in HEVC context-based adaptive binary arithmetic coding designs. Secondly, Utilizing the unique characteristics of AVS3 AEC, we propose mathematical reordering, variable-clock-cycle range updating and variable-clock-cycle context modeling methods to optimize the critical path. Our design can encode 2.6457 bins per clock cycle, and the corresponding throughput is 3131 Mbin/s in Globalfoundries 28nm process. Compared with the basic anchor structure, it has obtained a performance improvement of 319% and can meet the 8k@l20fps ultra-high-definition video encoding requirements. Yujie Cai, Xiaoyang Zeng, Yibo Fan, Peng Zhang 0007, Guoqing Xiang, Haibing Yin |
ISCAS | 1 |
| 2022 | A Fast CABAC Hardware Design for Accelerating the Rate Estimation in HEVCabstractThe latest High Efficiency Video Coding standard achieves twice the coding efficiency of the H264 standard through a complex rate-distortion optimization (RDO). The coded bit-streams are produced with context adaptive binary arithmetic coding (CABAC). CABAC itself is a very time-consuming process that includes binarization, context modeling, interval subdivision, renormalization, outstanding bit handling, and context updating. The aim of this research is to speed up the CABAC process through several simplifications. First, we approximate three parts of the CABAC, i.e., interval subdivision, renormalization, and outstanding bit handling, with a piecewise-linear function that is very friendly to hardware implementation. In order to achieve better hardware parallelism, we also improve the coding process at the sub-block level. The context of syntax elements in a sub-block is redistributed to skip the complex calculation of context indexing. We perform context updating at the granularity of sub-blocks so that the data dependency of the context updating is removed completely, and the original serial encoding process is changed to a parallel encoding process. At the same time, we make another simplification for the context modeling ofcu_skip_flag. Based on these simplifications, we build a parallel hardware architecture for the rate estimation of the RDO process. This architecture completes the bit estimation of a$32\times 32$coding tree unit (CTU) in 220.8 nano-seconds, whereas the Bjøntegaard Delta rate increases by only 2.225%. We believe that the proposed architecture can meet the requirements of 8K@120 fps ultra-high-definition videos. This is the first study to simplify the hardware design of rate estimation by changing the context allocation and updating rules. Yujie Cai, Yibo Fan, Leilei Huang, Xiaoyang Zeng, Haibing Yin, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Dynamic Task Scheduler for Real Time Requirement in Cloud Computing System
Yujie Cai, Ming-e Jing, Yibo Fan, Xiaoyang Zeng |
ICA3PP (4) | 3 |
| 2014 | Biochemical systems identification by a random drift particle swarm optimization approachabstractBACKGROUND: Finding an efficient method to solve the parameter estimation problem (inverse problem) for nonlinear biochemical dynamical systems could help promote the functional understanding at the system level for signalling pathways. The problem is stated as a data-driven nonlinear regression problem, which is converted into a nonlinear programming problem with many nonlinear differential and algebraic constraints. Due to the typical ill conditioning and multimodality nature of the problem, it is in general difficult for gradient-based local optimization methods to obtain satisfactory solutions. To surmount this limitation, many stochastic optimization methods have been employed to find the global solution of the problem. RESULTS: This paper presents an effective search strategy for a particle swarm optimization (PSO) algorithm that enhances the ability of the algorithm for estimating the parameters of complex dynamic biochemical pathways. The proposed algorithm is a new variant of random drift particle swarm optimization (RDPSO), which is used to solve the above mentioned inverse problem and compared with other well known stochastic optimization methods. Two case studies on estimating the parameters of two nonlinear biochemical dynamic models have been taken as benchmarks, under both the noise-free and noisy simulation data scenarios. CONCLUSIONS: The experimental results show that the novel variant of RDPSO algorithm is able to successfully solve the problem and obtain solutions of better quality than other global optimization methods used for finding the solution to the inverse problems in this study. Jun Sun 0008, Vasile Palade, Yujie Cai, Wei Fang 0001, Xiaojun Wu 0001 |
BMC Bioinform. | 3 |