Weijing Shi

dblp:158/7437 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0003-0575-2271ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 first-author
YearPublicationVenuePosition
2026 Reinforced Rate Control for Neural Video Compression via Inter-Frame Rate-Distortion Awareness
abstract
Neural video compression (NVC) has demonstrated superior compression efficiency, yet effective rate control remains a significant challenge due to complex temporal dependencies. Existing rate control schemes typically leverage frame content to capture distortion interactions, overlooking inter-frame rate dependencies arising from shifts in per-frame coding parameters. This often leads to suboptimal bitrate allocation and cascading parameter decisions. To address this, we propose a reinforcement‑learning (RL)‑based rate control framework that formulates the task as a frame‑by‑frame sequential decision process. At each frame, an RL agent observes a spatiotemporal state and selects coding parameters to optimize a long‑term reward that reflects rate‑distortion (R-D) performance and bitrate adherence. Unlike prior methods, our approach jointly determines bitrate allocation and coding configuration in a single step, independent of group‑of‑pictures (GOP) structure. Extensive experiments across diverse NVC architectures show that our method reduces the average relative bitrate error to 1.20 percent and achieves up to 13.45 percent bitrate savings at typical GOP sizes, outperforming existing approaches. In addition, our framework demonstrates improved robustness to content variation and bandwidth fluctuations with lower encoding/decoding overhead, making it highly suitable for practical deployment.
Wuyang Cong, Junqi Shi, Lizhong Wang, Weijing Shi, Ming Lu 0003, Hao Chen 0036, Zhan Ma 0001
AAAI4
2025 Integrating Adaptive Sampling for Optimal Learned Video Compression
abstract
We propose a novel adaptive prediction network that dynamically determines the optimal sampling factor and Lagrangian multiplier for encoding each frame, guided by sequential information. By exploiting spatio-temporal redundancy through adaptive sampling, our method reduces bitrate consumption while preserving the reconstruction quality with adjusted rate-distortion coefficients. Experimental results demonstrate significant performance gains over representative learned video compression models across various datasets, with reduced encoding and decoding latency.
Wuyang Cong, Yuzhuo Kong, Ming Lu 0003, Lizhong Wang, Weijing Shi, Zhan Ma 0001
ICASSP5
2025 SPCODEC: Split and Prediction for Neural Speech Codec
Liang Wen, Lizhong Wang, Yuxing Zheng, Weijing Shi, Kwangpyo Choi
INTERSPEECH4
2024 FT-CSR: Cascaded Frequency-Time Method for Coded Speech Restoration
abstract
Lossy speech codecs often introduce coding distortions such as coding noise and constrained bandwidth, which can affect the quality of the decoded speech. This paper proposes a method called FT-CSR, which is used for coded speech restoration. FT-CSR reduces coding noise and recovers missing frequencies sequentially using a cascaded frequency-time domain model. In experiments using the Opus codec, FT-CSR was found to be effective across bitrates ranging from 8 to 16 kbps and outperformed the baseline on both objective and subjective measurements. FT-CSR achieves a MOS-POLQA score of 3.6 or higher and improves MOS-POLQA by more than 0.23 when compared to decoded speech. The results of the subjective test show that FT-CSR can improve MOS by over 0.85 for decoded speech.
Liang Wen, Lizhong Wang, Yuxing Zheng, Weijing Shi, Kwangpyo Choi
ICME4
2023 Distortion-Aware Convolutional Neural Network-Based Interpolation Filter for AVS3
abstract
Motion compensation is a key technology in video coding for removing the temporal redundancy between video frames. Considering the incompatibility between traditional interpolation filters and diversified video content, the inter prediction method still has considerable room for improvement. This paper proposed a distortion-aware convolutional neural network-based interpolation filter (DA-NNIF) to further improve the interpolation prediction accuracy of sub-pixels with one model. Distortion parameters are introduced into the proposed network to reflect the quantization noise of reference frames. The experimental result shows that the proposed method achieves on average 1.47 % BD-rate reduction on Y component for ClassB, ClassC and ClassD sequences under the random access configuration of AVS3.
Liang Wen, Lizhong Wang, Yinji Piao, Weijing Shi, Kwangpyo Choi
ICASSP5
2021 Opportunities and Challenges for Flagman Recognition in Autonomous Vehicles
abstract
Autonomous vehicles promise significant advances in transportation safety, efficiency and comfort. However, achieving the goal of full autonomy is impeded by the need to address several operational challenges encountered in practice. Gesture recognition of flagmen on roads is one such set of challenges. An autonomous vehicle needs to make safe decisions and facilitate forward progress in the presence of road construction workers and flagmen. However, human gestures under diverse environmental conditions are very varied and represent significant complexity. In this work, we present (i) a taxonomy of challenges for organizing traffic gestures, (ii) a sizeable flagman gesture dataset, and (iii) extensive experiments on practical algorithms for gesture recognition. We categorize traffic gestures according to their semantics, flagman appearances and the environmental context. We then collect a dataset covering a range of common flagman gestures with and without props such as signs and flags. Finally, we develop a recognition algorithm using different feature representations of the human pose and perform extensive ablation experiments on each component.
Weijing Shi, Ragunathan Rajkumar, Eran Kishon
IV1
2020 Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud
abstract
In this paper, we propose a graph neural network to detect objects from a LiDAR point cloud. Towards this end, we encode the point cloud efficiently in a fixed radius near-neighbors graph. We design a graph neural network, named Point-GNN, to predict the category and shape of the object that each vertex in the graph belongs to. In Point-GNN, we propose an auto-registration mechanism to reduce translation variance, and also design a box merging and scoring operation to combine detections from multiple vertices accurately. Our experiments on the KITTI benchmark show the proposed approach achieves leading accuracy using the point cloud alone and can even surpass fusion-based algorithms. Our results demonstrate the potential of using the graph neural network as a new approach for 3D object detection. The code is available at https://github.com/WeijingShi/Point-GNN.
Weijing Shi, Ragunathan Rajkumar
CVPR1
2017 Algorithm and hardware implementation for visual perception system in autonomous vehicle: A survey
Weijing Shi, Mohamed Baker Alawieh, Xin Li 0001, Huafeng Yu
Integr.1
2017 An FPGA-Based Hardware Accelerator for Traffic Sign Detection
abstract
Traffic sign detection plays an important role in a number of practical applications, such as intelligent driver assistance and roadway inventory management. In order to process the large amount of data from either real-time videos or large off-line databases, a high-throughput traffic sign detection system is required. In this paper, we propose an FPGA-based hardware accelerator for traffic sign detection based on cascade classifiers. To maximize the throughput and power efficiency, we propose several novel ideas, including: 1) rearranged numerical operations; 2) shared image storage; 3) adaptive workload distribution; and 4) fast image block integration. The proposed design is evaluated on a Xilinx ZC706 board. When processing high-definition (1080p) video, it achieves the throughput of 126 frames/s and the energy efficiency of 0.041 J/frame.
Weijing Shi, Xin Li 0001, Zhiyi Yu, Gary Overett
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Efficient statistical validation of machine learning systems for autonomous driving
abstract
Today's automotive industry is making a bold move to equip vehicles with intelligent driver assistance features. A modern automobile is now equipped with a powerful computing platform to run multiple machine learning algorithms for environment perception (e.g., pedestrian detection) and motion control (e.g., vehicle stabilization). These machine learning systems must be highly robust with extremely small failure rate in order to ensure safe and reliable driving. In this paper, we propose a novel Subset Sampling (SUS) algorithm to efficiently validate a machine learning system. In particular, a Markov Chain Monte Carlo algorithm based on graph mapping is developed to accurately estimate the rare failure rate with a minimal amount of test data, thereby minimizing the validation cost. Our numerical experiments show that SUS achieves 15.2× runtime speed-up over the conventional brute-force Monte Carlo method.
Weijing Shi, Mohamed Baker Alawieh, Xin Li 0001, Huafeng Yu, Nikos Aréchiga, Nobuyuki Tomatsu
ICCAD1