VLDB 2026 Research / reviewers in the wild / expert
Minhao Tang
dblp:157/8719
· DBLP profile ↗
19ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0002-5421-1116ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 40 |
| 2025 | Enhanced Frame Context Initialization for Video Coding Beyond AV1abstractEntropy coding is an integral part of all modern hybrid block-based video codecs. The context adaptive binary arithmetic coding (CABAC) is a normative part of ITU-T/ISO/IEC video coding standards (H.264/AVC, HEVC and VVC), while the Alliance for Open Media (AOMedia) video standard (AV1) utilizes a context adaptive multi-symbol version of it for entropy coding. Recent explorations of new coding tools beyond AV1 capabilities have led to the improvement of the context adaptive multi-symbol arithmetic entropy coder. In this study, improvements to the context initialization process of the entropy coder are discussed in detail and results reported. The improvements include a) optimal selection of reference frame pairs and b) context modeling from selected reference frame pairs for tile context initialization. Two variants of the proposed method, Variant1 and Variant2 are implemented on top of the reference codebase (research-v8.0.0). Experimental results with CTCv7 show that for random access (RA) configuration, average BDRATE (PSNR-YUV) gains of -0.13% and -0.18% are achievable for Variant1 and Variant2 respectively. Meanwhile, for low delay (LD) configuration the reported gains are -0.47% and -0.50% respectively for Variant1 and Variant2. Madhu Peringassery Krishnan, Wei Kuang, Minhao Tang, Shan Liu 0001 |
ICIP | 6 |
| 2025 | Hardware Friendly Multi-Hypothesis Cross Component PredictionabstractThis paper introduces a novel Multi-Hypothesis Cross Component Prediction (MHCCP) method to enhance coding efficiency in image and video compression on top of AOMedia Video Model (AVM). Inspired by prior cross-component coding techniques, this work firstly presents a new cross component prediction method and then introduces a hardware-friendly design for practical deployment. The proposed MHCCP method initially generates multiple hypothesis predictions, and a linear combination of these hypothesis predictions is used to estimate the final chroma intensity. To determine the coefficients of the linear model in MHCCP, both the encoder and decoder use the Gaussian elimination method, which involves significant computational complexity. To address hardware implementation challenges, such as the high complexity for parameter derivation and limited above line buffer at the superblock boundary, the proposed method incorporates several optimizations, including decoupled vertical-horizontal prediction modes to capture diverse texture patterns, a padding mechanism to overcome line buffer constraints, and a sub-sampling strategy to reduce computational complexity. Experimental results on Common Test Condition (CTC) v7 demonstrate consistent coding gains: -0.57% (YUV-PSNR), -0.42% (Y-PSNR), -1.92% (U-PSNR), and -2.09% (V-PSNR) under all-intra settings on anchor research-v8.0.0. Notably, classes A1 and A2 achieve significant gains of -0.88% and -0.61%, respectively, highlighting the efficacy of the proposed approach. Madhu Peringassery Krishnan, Shan Liu 0001, Minhao Tang |
ICIP | 6 |
| 2025 | Extension of Semi-Decoupled Partitioning in Inter FramesabstractThe Alliance for Open Media (AOMedia) has been exploring new coding tools to enhance AV1 capabilities. Semi-Decoupled Partitioning (SDP), originally designed for intra frames in research-v2.0.0, improves coding by decoupling luma and chroma block partitioning. This study extends SDP to inter frames by introducing intra region coding, where the root node is explicitly signaled in the bitstream. Within the intra region, luma components of the intra-coded blocks can be further split, while chroma components remain unsplit. The experiments are implemented on the 8thanchor, research-v8.0.0, of AVM reference software with CTCv7, and experimental results show that the proposed method can achieve 0.12%, 2.39%, 2.58% coding gain for Y, U, and V component separately with random access configuration and 5% encoding time increase and almost no decoding time increase. Madhu Peringassery Krishnan, Shan Liu 0001, Jayasingam Adhuran, Minhao Tang, Jianle Chen, Urvang Joshi, Mohammed Golam Sarwer, Debargha Mukerjee |
ICIP | 5 |
| 2025 | Lossless Coding Improvement beyond AV1abstractLossless compression plays an important role in the storage and transmission of data with stringent quality requirements. There is a substantial demand for enhancing the lossless compression performance of the current AV1 codec. In this paper, two novel techniques, named Residual Block Refinement (RBR) mode and Multi-Residual Blocks (MRB) mode, are introduced to improve the lossless coding performance beyond AV1. For the RBR mode, the main idea is to perform a lossless block refinement within the residual block to further reduce redundancy. For the MRB mode, the first partial residual block utilizes a traditional transform and quantization process to generate a lossy representation of the original residual samples with efficient energy compaction. The second partial residual block is further coded to achieve a perfect representation of the difference between the original residual block and the reconstructed first residual block. The experimental results reveal that an average coding performance -2.76 %, -1.04 %, and -1.18 % are achieved on top of the AOMedia Video Model (AVM) v6.0.0 in terms of Bitrate savings for allIntra (AI), Random Access (RA), and Low Delay (LD) configurations, respectively. Madhu Peringassery Krishnan, Shan Liu 0001, Minhao Tang |
VCIP | 5 |
| 2019 | A Conditional Bayesian Block Structure Inference Model for Optimized AV1 EncodingabstractAV1, a next-generation open-source and royalty-free video coding standard, achieves high compression performance at high computational cost. To meet the requirements of HD and UHD video applications, extensive optimizations in both the algorithm and implementation of AV1 are required. In this paper, we analyze the similarities between the block structure decisions after rate-distortion (RD) optimized AV1 and HEVC encodings of the same input. Taking advantage of such similarities, we propose a conditional Bayesian inference model to perform early termination in block partition determination of AV1 based on HEVC encoding outputs. An estimation algorithm is designed to iteratively calculate the prior probability for Bayesian inference. Experiment results show that our proposed algorithm could realize an average time saving of 35.7% and negligible BD-rate loss (0.61%), with the pre-encoding time taken into consideration. Bichuan Guo, Minhao Tang, Yuxing Han 0001, Jiangtao Wen |
ICME | 3 |
| 2019 | Hadamard Transform-Based Optimized HEVC Video CodingabstractThe High Efficiency Video Coding (HEVC/H.265) standard achieves great improvement in compression efficiency over the widely used H.264/AVC standard at a cost of much higher complexity. When encoding videos using HEVC, the selection of the quantization parameter (QP) can significantly affect the coding efficiency. Typical algorithms for adaptive quantization employ fixed bitrate budgeting or fixed QP offsets for different frames and different blocks, without considering detailed input video characteristics, some of which might be captured using computer vision methods. The problem of determining adaptive settings of HEVC coding parameters has not been satisfactorily solved. In this paper, we proposed a Hadamard (“HAD”) energy-based optimized HEVC video encoder, in which HAD energy is used to measure the amount of residual information to be encoded in a block so as to determine the QP value for each block for a better coding efficiency. HAD energy is also used to expedite the time-consuming mode decision process and for a precise scene change detection to avoid flicker artifacts. Experiment using the widely used open source HEVC encoder x265-v1.8 showed that the proposed algorithm was able to achieve an average of 14% saving in Bjøntegaard-delta-rate and an average of 20% saving in encoding time as compared the “medium” preset of x265, while the proposed algorithm also produced an improvement of 3.3% in coding efficiency for the HEVC reference software HM-16.6. Minhao Tang, Jiangtao Wen, Yuxing Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | A Universal Optical Flow Based Real-Time Low-Latency Omnidirectional Stereo Video SystemabstractOmnidirectional stereoscopic video (ODSV) is a key element of creating an immersive experience for virtual reality that has attracted extensive interest while presenting many technical challenges. Two such key challenges are real-time, low-latency high-quality seamless video stitching from multiple cameras, and faithful reconstruction of 3-D information. Even though various attempts have been made to achieve different combinations of real-time, low-latency, automation, and high output resolution in stereoscopic panoramic video communication, achieving these characteristics simultaneously remains a challenge to be tackled. In this paper, we present a universally applicable and practical end-to-end system based on a novel real-time optical flow algorithm to produce high-quality real-time ODSV with reconstructed 3-D depth information at low latency. Through a configurable process, various camera systems can be calibrated and seamlessly stitched together using the proposed system. The stitched 3-D panoramic video is encoded with a standard compliant video encoder that is optimized for panoramic video. Thanks to various optimizations introduced in this paper, the proposed system is capable of producing real-time ODSV of ultra High definition resolution with a glass to glass latency of 2.2 s using a desktop computer with a single Nvidia graphic card. Experiments show that the proposed system achieves an encoding performance superior to existing open-source HEVC implementations and an optical flow estimation performance better than the Facebook algorithm while running two orders of magnitudes faster. Minhao Tang, Jiangtao Wen, Jiawen Gu, Philip Junker, Bichuan Guo, Guansyun Jhao, Yuxing Han 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | Adaptive Intra Candidate Selection With Early Depth Decision for Fast Intra Prediction in HEVCabstractTo better exploit spatial correlations in a video frame, the High Efficiency Video Coding (HEVC) standard has adopted a great many more intra prediction modes than H.264/AVC. As a result, the complexity of rate-distortion-optimized (RDO) HEVC intra mode selection is very high. Many techniques have been proposed to expedite the intra mode selection process to achieve a better overall tradeoff between complexity and RD performance. In this paper, two novel techniques for adaptive intra mode candidate selection and bidirectional depth search algorithms are utilized to accelerate intra prediction. The proposed techniques show an average of 63% (up to 67%) time saving with only 1% BD-Rate increase, outperforming most of existing intra prediction algorithms. Jiawen Gu, Minhao Tang, Jiangtao Wen, Yuxing Han 0001 |
IEEE Signal Process. Lett. | 2 |
| 2018 | Accelerating HEVC Encoding Using Early-SplitabstractThe increase in coding efficiency and complexity of high efficiency video coding (HEVC) over H.264 is due to, among other factors, the time needed to find the optimal partition structure among the more flexible encoding modes for the coding units (CUs) and prediction units (PUs). Although many classification-based algorithms have been proposed to expedite the partition decision, the features that can be acquired from current HEVC encoding order are not sufficient to minimize the loss in coding efficiency. In this letter, we proposed an early-split (ES) order for HEVC CU-level encoding, where the encoder checks the split mode before the nonsquare PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm can save 48% of encoding time on average with only about 0.8% loss in coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
IEEE Signal Process. Lett. | 1 |
| 2017 | SATD Based Fast Intra Prediction for HEVCabstractSummary form only given. To better exploit spatial correlations in a video frame, the HEVC video coding standard has introduced many intra prediction modes and a recursive quadtree-based coding unit (CU) structure. As a result, the complexity of Rate-distortion optimized (RDO) HEVC intra mode selection is significantly higher. Many techniques have been proposed to expedite the intra mode selection process to achieve a good overall trade-off between complexity and RD performance. In this paper, we proposed a fast intra decision algorithm based on Hadamard Transform. The algorithm consist of three parts: SATD calculation reduction, adaptive intra candidate selection, and SATD based early termination. Experiments conducted using the HEVC common test conditions show an average of 56.4% (up to 64.1%) time saving with only 1.2% increase in Bjontegaard delta rate (BD-rate) using the proposed algorithm. Jiawen Gu, Minhao Tang, Jiangtao Wen |
DCC | 2 |
| 2017 | Early-Split Based Fast HEVC EncodingabstractThe High Efficiency Video Coding (HEVC) standard achieves 50% improvement incompression efficiency over the widely used H.264/AVC standard at a cost of much higher complexity. The increase in complexity is due to, among other factors, the time needed to findthe optimal partition structure among the more flexible possibilities for the coding units (CUs) and prediction units (PUs). Many classification based algorithms have been proposed to reduce this partition decision time, but the features that can be acquired from current HEVC encoding order may not be sufficient to control the loss in coding efficiency. In this paper, we proposed an Early-Split (ES) order for HEVC encoding, where the encoder checks the split mode before the non-square PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm achieved an average of 48% saving in encoding time with only 0.92% loss in the coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
DCC | 1 |
| 2017 | HEVC-based motion compensated joint temporal-spatial video denoisingabstractA novel HEVC-based efficient video denoising algorithm is proposed in this paper. It uses a spatial Gaussian filter for the chrominance components and then utilizes the HEVC motion estimation process to find the best temporal correspondence for low-pass filtering. Other HEVC tools such as quantization, the interpolation and the in-loop filters are also used. Experiments implementing the proposed algorithm in the open-source HEVC encoder ×265 showed a good denoising performance with a much lower computing complexity than the competitors. The performance was comparable to those highly sophisticated algorithms such as the VBM4D, which is 200 times slower. The proposed algorithm can be easily integrated into the real-world video processing systems due to its compatibility with the HEVC standard. Minhao Tang, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
ICASSP | 1 |
| 2017 | A novel satd based fast intra prediction for HEVCabstractTo better exploit spatial correlations in a video frame, the HEVC video coding standard has introduced many intra prediction modes and a recursive quadtree-based coding unit (CU) structure. As a result, the complexity of Rate-distortion optimized (RDO) HEVC intra mode selection is significantly higher. Many techniques have been proposed to expedite the intra mode selection process to achieve a good overall trade-off between complexity and RD performance. In this paper, we proposed a fast intra decision algorithm consisting of three parts: calculation reduction, adaptive intra candidate selection, and fast depth decision used early termination. Experiments conducted using the HEVC common test conditions show an average of 61.1% (up to 67.7%) time saving with only 1.03% increase in Bjontegaard delta rate (BD-rate) using the proposed algorithm, out-performing existing state-of-the-art algorithms. Jiawen Gu, Minhao Tang, Jiangtao Wen |
ICIP | 2 |
| 2017 | Optimized video coding for omnidirectional videosabstractThe ever widening application of virtual reality requires the ultra high resolution omnidirectional videos (OVs) to be transmitted over the wired and wireless Internet at low cost (i.e. bitrate). Various solutions have been proposed to intelligently reduce the bitrate, e.g. adapting the spatial resolution of the video for different directions of the panorama with regard to current direction that the viewer is looking at and the distribution of the probability of each direction to be watched according to the video content. Due to various reasons, spatial resolution adaptation may often cause perceivable quality degradation to user experience. In this paper, we proposed two adaptive encoding techniques to reduce the bitrate of OVs after compression. The first is a content adaptive temporal resolution adaptation scheme for OVs using cube map projection. The second is a quantization and rate-distortion optimization scheme for equirectangular projection. Experiments implementing the proposed algorithms in the open source HEVC encoder x265 show that the proposed algorithms can save 14.69% and 13.01% in BD-rate on average for the cube map and equirectangular projected OVs respectively, while no degradation to the visual experience was reported in subjective tests. The proposed algorithms are compatible with and therefore can collaborate with the currently used adaptive spatial resolution scheme for further bitrate reduction. Minhao Tang, Jiangtao Wen, Shiqiang Yang |
ICME | 1 |
| 2015 | R-(lambda) Model Based Improved Rate Control for HEVC with Pre-EncodingabstractIn this paper, we proposed a new rate control algorithm for High Efficiency Video Coding (HEVC), the latest video coding standard from the ITU/ISO. We use the information of pre-encoded 16x16 coding units (CUs) to estimate the characteristics of the largest coding unit (LCU). Based on the estimates, the proposed R -- λ model can be refined before the real encoding process. This is in contrast to rate control algorithms such as that in the HEVC reference software, where the model is updated based on a previously encoded picture. Experimental results show that the proposed rate control scheme can achieve accurate rate control with a BD-PSNR gain up to 5.37dB, compared to the state-of-the-art rate control algorithm in the HEVC test model (HM) 16.0. The largest PSNR improvement was over 6dB. Jiangtao Wen, Meiyuan Fang, Minhao Tang, Kuang Wu |
DCC | 3 |
| 2015 | An efficient HEVC to H.264/AVC transcoding systemabstractThe latest High Efficiency Video Coding (HEVC) achieves significant compression performance improvement over Advanced Video Coding (AVC). The high compression efficiency of HEVC together with the wide availabilities of H.264/AVC decoders necessitates transcoding between these two technologies. In this paper, we propose a novel algorithm for software-based HEVC to H.264/AVC transcoding. By utilizing the information extracted from the input HEVC stream, the transcoding process can be accelerated with relatively minor compression efficiency loss. Experiment results show that the proposed transcoding algorithm can save around 60% of the time cost of re-encoding process compared with x264, one of the most widely used H.264 encoders, with very small loss of compression performance. Minhao Tang, Jiangtao Wen |
ISCAS | 1 |
| 2015 | Efficient Software H.264/AVC to HEVC Transcoding on Distributed Multicore ProcessorsabstractThe latest High Efficiency Video Coding (HEVC) standard achieves a significant compression efficiency improvement over the H.264/Advanced Video Coding (AVC) standard, but with a much higher computational complexity. In this paper, we propose a novel framework for software-based H.264/AVC to HEVC transcoding, integrated with tools such as wavefront parallel processing that are useful for achieving higher levels of parallelism on multicore processors and distributed systems. By utilizing information extracted from the input H.264/AVC bitstream, the transcoding process can be greatly accelerated with a visual quality loss that is modest for many applications. Based on the HEVC HM 14.0 reference software and using standard HEVC test bitstreams, the proposed transcoder can achieve up to 60× speedup on a Quad Core 8-thread server over decoding-re-encoding based on FFMPEG and the HM software with a BD-rate loss of 15%-20%. By implementing a group of picture-level task distribution on a distributed system with nine processing units, the proposed software transcoder can achieve a speed for transcoding 720 p at 30 Hz in real time. Yucong Chen, Ziyu Wen, Jiangtao Wen, Minhao Tang, Pin Tao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Accelerating HEVC using heterogeneous platforms
Gabriel Cebrián-Márquez, José Luis Hernández-Losada, José Luis Martínez 0001, Pedro Cuenca 0001, Minhao Tang, Jiangtao Wen |
J. Supercomput. | 5 |