EDBT 2026 Demo / reviewers in the wild / expert
Jiro Katto
dblp:55/3226
· DBLP profile ↗
137ranked-venue papers
9as first author
60since 2021 · last 2026
0000-0002-1671-2614ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 77 · 8 first-author · 31 since 2021Computer networks · 29 · 1 first-author · 16 since 2021Systems, architecture and hardware · 10 · 8 since 2021Artificial intelligence and machine learning · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Variable-Rate Learned Image Compression Using Parameter Efficient Fine-Tuning and Trainable Quantization Step-Size
Ran Wang 0015, Heming Sun, Jiro Katto |
ISCAS | 4 |
| 2025 | Adaptive Quality Control Method for a Room-Scale 6DoF Point Cloud Streaming and its Evaluationabstract6Degree of Freedom (6DoF) point cloud streaming is a promising technology for immersive communication. To achieve immersive communication, it is necessary to improve technologies at each phase, such as 3D tiling, 3D tile selection, and bitrate allocation to 3D tiles, and there is a need to integrate them. Researchers have proposed many methods for efficient point cloud streaming in each phase. However, many related studies remain limited to the improvement of individual technologies. In addition, they have not evaluated the case of 6DoF point cloud streaming using spatially larger scale (room-scale) point clouds. Therefore, we propose an efficient quality control method combining the three essential methods: adaptive 3D tiling, occlusion considering 3D tile selection, and distance-driven adaptive quality optimization. Furthermore, we propose quality control strategies by not only coordinating three technologies but also altering their integration method (i.e., changing the order of applying three methods). The proposed method enables various quality controls, such as data size minimization or quality maximization, according to requirements. In performance evaluations, we conduct experiments on Unity using an actual-scale digital room and compare performances of the proposed method against baseline methods. The results confirm that the proposed method improves perceptual quality or efficiently reduces the data size. Yusuke Tagashira, Yumeka Chujo, Kenji Kanai, Jiro Katto |
CCNC | 4 |
| 2025 | FLAVC: Learned Video Compression with Feature Level AttentionabstractLearned Video Compression (LVC) aims to reduce redundancy in sequential data through deep learning approaches. Recent advances have significantly boosted LVC performance by shifting compression operations to the feature domain, often combining Motion Estimation and Motion Compensation modules (MEMC) with CNN-based context extraction. However, reliance on motions and convolution-driven context models limits generalizability and global perception. To address these issues, we propose a Feature-level Attention (FLA) module within a Transformer-based framework that explicitly perceives full-frame, thus bypassing confined motion signatures. FLA accomplishes global perception by converting high-level local patch embeddings into one-dimensional batch-wise vectors and replacing traditional attention weights to a global context matrix. Additionally, a dense overlapping patcher (DP) is introduced to retain local features before embedding projection. Furthermore, a Transformer-CNN mixed encoder is applied to alleviate the spatial feature bottleneck without increasing latent size. Experiments demonstrate excellent generalizability with universally efficient redundancy reduction in different scenarios. Extensive tests on four video compression datasets show that our method achieves state-of-the-art Rate-Distortion performance compared to existing LVC methods and traditional codecs. A down-scaled version of our model reduced computation overhead by a great margin while maintained strong performance. The code is available at https://github.com/Z-CV-code/FLAVC. Heming Sun, Jiro Katto |
CVPR | 3 |
| 2025 | Performance Analysis of 5G FR2 (mmWAVE) Downlink 256QAM on Commercial 5G NetworksabstractThe 5G New Radio (NR) standard introduces new frequency bands allocated in Frequency Range 2 (FR2) to support enhanced Mobile Broadband (eMBB) in congested environments and enables new use cases such as Ultra-Reliable Low Latency Communication (URLLC). The 3GPP introduced 256QAM support for FR2 frequency bands to further enhance downlink capacity. However, sustaining 256QAM on FR2 in practical environments is challenging due to strong path loss and susceptibility to distortion. While 256QAM can improve theoretical throughput by 33%, compared to 64QAM, and is widely adopted in FR1, its real-world impact when utilized in FR2 is questionable, given the significant path loss and distortions experienced in the FR2 range. Additionally, using higher modulation correlates to higher BLER, increased instability, and retransmission. Moreover, 256QAM also utilizes a different MCS table defining the modulation and code rate at different Channel Quality Indexes (CQI), affecting the UE's link adaptation behavior. This paper investigates the real-world performance of 256QAM utilization on FR2 bands in two countries, across three RAN manufacturers, and in both NSA (EN-DC) and SA (NR-DC) configurations, under various scenarios, including open-air plazas, city centers, footbridges, train station platforms, and stationary environments. The results show that 256QAM provides a reasonable throughput gain when stationary but marginal improvements when there is UE mobility while increasing the probability of NACK responses, increasing BLER, and the number of retransmissions. Finally, MATLAB simulations are run to validate the findings as well as explore the effect of the recently introduced 1024 QAM on FR2. Kasidis Arunruangsirilert, Pasapong Wongprasert, Jiro Katto |
ICC | 3 |
| 2025 | Evaluations of High Power User Equipment (HPUE) in Urban EnvironmentabstractWhile Time Division Duplexing (TDD) 5G New Radio (NR) networks offers higher downlink throughput due to the utilization of the middle frequency band, the uplink performance is negatively impacted due to higher path loss associated with higher frequencies, which degrade the users’ QoE in less optimal conditions. With the growing demand for high performance uplink throughput from novel applications such as Metaverse, Internet of Things (IoTs) and Smart City, 3GPP introduced High Power User Equipment (HPUE) on 5G TDD bands, allowing UEs to utilize more than 23 dBm of power for transmission to improve throughput, QoE, and reliability, especially at the cell edges. In this paper, the performance of HPUE is evaluated in the urban area on a commercial 5G network in terms of Uplink Throughput, Modulation Efficiency, Re-transmission Rate (ReTx Rate), and Power Consumption in both Standalone (SA) and Non-Standalone (NSA) modes. Through modem firmware modification, the performance is also compared across different power classes and antenna configurations. Kasidis Arunruangsirilert, Pasapong Wongprasert, Jiro Katto |
ICCCN | 3 |
| 2025 | Adaptive Data Transmission Management by Incorporating Sensing in mmWave V2V CommunicationabstractHigh-speed data transmission is enabled by millimeter wave communication. While the short wavelength may cause reliability issues, which is critical in some scenarios such as vehicle-to-vehicle communication. In this work, an adaptive data transmission management method is proposed by incorporating sensing results. With the moving information of the vehicles, the transmission strategy is adaptively adjusted to improve communication efficiency. The performance of the proposed method is evaluated in simulation experiment by extending the network simulator framework NS3. The results demonstrate that it is promising to utilize the sensing results for intelligent network management. Bo Wei 0001, Hang Song 0001, Jiro Katto |
ICCCN | 3 |
| 2025 | Evaluation of NVENC Split-Frame Encoding (SFE) for UHD Video Transcoding
Kasidis Arunruangsirilert, Jiro Katto |
PCS | 2 |
| 2025 | A Multi-Grid Implicit Neural Representation for Multi-View Videos
Qingyue Ling, Zhengxue Cheng, Donghui Feng 0003, Shen Wang 0013, Guo Lu, Heming Sun, Jiro Katto, Li Song 0001 |
PCS | 8 |
| 2025 | Evaluation of GPU Video Encoder for Low-Latency Real-Time 4K UHD EncodingabstractThe demand for high-quality, real-time video streaming has grown exponentially, with 4K Ultra High Definition (UHD) becoming the new standard for many applications such as live broadcasting, TV services, and interactive cloud gaming. This trend has driven the integration of dedicated hardware encoders into modern Graphics Processing Units (GPUs). Nowadays, these encoders support advanced codecs like HEVC and AV1 and feature specialized Low-Latency and Ultra Low-Latency tuning, targeting end-to-end latencies of <2 seconds and <500 ms, respectively. As the demand for such capabilities grows toward the 6G era, a clear understanding of their performance implications is essential. In this work, we evaluate the low-latency encoding modes on GPUs from NVIDIA, Intel, and AMD from both Rate-Distortion (RD) performance and latency perspectives. The results are then compared against both the normal-latency tuning of hardware encoders and leading software encoders. Results show hardware encoders achieve significantly lower E2E latency than software solutions with slightly better RD performance. While standard Low-Latency tuning yields a poor quality-latency trade-off, the Ultra Low-Latency mode reduces E2E latency to 83 ms (5 frames) without additional RD impact. Furthermore, hardware encoder latency is largely insensitive to quality presets, enabling high-quality, low-latency streams without compromise. Kasidis Arunruangsirilert, Jiro Katto |
VCIP | 2 |
| 2025 | FlashGMM: Fast Gaussian Mixture Entropy Model for Learned Image CompressionabstractHigh-performance learned image compression codecs require flexible probability models to fit latent representations. Gaussian Mixture Models (GMMs) were proposed to satisfy this demand, but suffer from a significant runtime performance bottleneck due to the large Cumulative Distribution Function (CDF) tables that must be built for rANS coding. This paper introduces a fast coding algorithm that entirely eliminates this bottleneck. By leveraging the CDF’s monotonic property, our decoder performs a dynamic binary search to find the correct symbol, eliminating the need for costly table construction and lookup. Aided by SIMD optimizations and numerical approximations, our approach accelerates the GMM entropy coding process by up to approximately 90x without compromising rate-distortion performance, significantly improving the practicality of GMM-based codecs. The implementation will be made publicly available at https://github.com/tokkiwa/FlashGMM. Shimon Murai, Fangzheng Lin, Jiro Katto |
VCIP | 3 |
| 2025 | Evaluation of 2D Video Interpolation and Extrapolation Methods for Real-Time V-PCC Error ConcealmentabstractPacket losses in streaming 3D point clouds with V-PCC over RTP significantly impact reconstruction quality, affecting users’ quality of experience. Despite this, practical error concealment methodologies are not yet widely researched, as previous 3D-based methods suffer from real-time performance and restrict coding modes to all-intra only. In this work, we evaluate various state-of-the-art 2D video interpolation and extrapolation works as candidates for the real-time V-PCC error concealment task in the video domain. We show that although their effectiveness varies with the regularity of generated patches or V-PCC video types, they can run in real-time, significantly faster than 3D methods, and can perform effectively on V-PCC attribute loss and temporally consistent patch videos. Eiko Nakajima, Fangzheng Lin, Kasidis Arunruangsirilert, Jiro Katto |
VCIP | 4 |
| 2025 | CGICM: CLIP-Guided Semantic Frequency Adaptation in Image Compression for MachinesabstractIn recent years, deep learning-based image compression techniques have advanced rapidly, surpassing traditional methods in terms of rate-distortion performance. However, in machine-oriented image compression, preserving high-level semantic information is of greater importance. Most existing methods employ only image-level prompts to guide frequency domain processing, leading to suboptimal preservation of semantic information for downstream machine vision tasks. To address this limitation, we propose a CLIP-guided semantic frequency domain adaptation module that extracts frequency features by applying both the fast Fourier transform and the wavelet transform. Guided by text-based semantics, the module further enhances the frequency components relevant to the target task, thereby improving machine perception performance. The proposed adapter is designed to be plug-and-play with existing learned image compression (LIC) models without requiring retraining of the full model. Experimental results demonstrate that our method outperforms state-of-the-art approaches in multiple machine vision tasks. Feng Liang 0001, Heming Sun, Jiro Katto |
VCIP | 4 |
| 2025 | Storage-and-Memory-Efficient Learned Image Compression With Quality-Aware Hyperprior PruningabstractABSTRACT Learned image compression (LIC) has become more and more important in recent years. The hyperprior‐module‐based LIC models, which use hyperprior module to predict the distribution of image features and improve entropy coder performance, have achieved remarkable rate‐distortion (RD) performance. However, the storage and memory costs of these LIC models are too high, resulting in higher difficulty to be applied to various devices, especially portable or edge devices. The storage and memory cost are directly linked to the parameter number. As a preliminary experiment, we manually assigned half channels for the hyperprior module in LIC models, reducing about 30% parameters in the model. The pruned models still kept similar RD performance to the original ones. This reveals that the hyperprior module in LIC models is highly redundant. In the meanwhile, LIC models with different reconstruction qualities require different amounts of parameters for the hyperprior module. Based on these phenomena, we propose a quality‐aware hyperprior pruning method that efficiently reduces the storage and memory cost of the hyperprior module and various context models. It consists of two parts. The first part is the pruning method itself, called enhanced ResRep on hyper path (ERHP). The second part is a quality‐aware threshold searching method, called pruning threshold searching (PTS), which prunes the hyperprior module based on the reconstruction qualities of LIC models. The experiments on various LIC models show that our methods reduce large volumes of storage cost (up to 74.6%) and memory cost (up to 41.5%), while keeping the performance the same before pruning. Ao Luo, Diego Fujii, Keisuke Nonaka, Heming Sun, Jiro Katto |
IET Image Process. | 5 |
| 2025 | MDLPCC: Misalignment-aware dynamic LiDAR point cloud compressionabstractLiDAR point cloud plays an important role in various real-world areas. It is usually generated as sequences by LiDAR on moving vehicles. Regarding the large data size of LiDAR point clouds, Dynamic Point Cloud Compression (DPCC) methods are developed to reduce transmission and storage data costs. However, most existing DPCC methods neglect the intrinsic misalignment in LiDAR point cloud sequences, limiting the rate–distortion (RD) performance. This paper proposes a Misalignment-aware Dynamic LiDAR Point Cloud Compression method (MDLPCC), which alleviates the misalignment problem in both macroscope and microscope. MDLPCC exploits a global transformation (GlobTrans) method to eliminate the macroscopic misalignment problem, which is the obvious gap between two continuous point cloud frames. MDLPCC also uses a spatial–temporal mixed structure to alleviate the microscopic misalignment, which still exists in the detailed parts of two point clouds after GlobTrans. The experiments on our MDLPCC show superior performance over existing point cloud compression methods. Ao Luo, Linxin Song, Keisuke Nonaka, Jinming Liu 0001, Kyohei Unno, Kohei Matsuzaki, Heming Sun, Jiro Katto |
J. Vis. Commun. Image Represent. | 8 |
| 2025 | Single model learned image compression utilizing multiple scaling factorsabstractImage compression is a critical task in multimedia. However, all learned-based single rate compression methods face challenges, such as prolonged training time due to the need for a dedicated model per bitrate and increased memory usage. Some variable rate methods require extra input, conditional networks, or still involve training multiple models. In this paper, we propose a unified approach using scaling factors to enable variable rate compression within a single model. The scaling factors consist of multi-gain units and quantization step size. The multi-gain units reduce redundancy in encoder and decoder representations, while the quantization step size controls quantization error. We also observe unevenness among slices in the Channel-Wise entropy model, and propose channel-wise quantization compensation by assigning specific step sizes to each slice. Our method supports continuous rate adaptation without retraining. Extensive experiments on CNN-based, Transformer-based, and CNN-Transformer mixed models demonstrate superior performance across a wide range of bitrates. • We enable variable-rate compression with a single model via multiple scaling factors. • We introduce channel-wise quantization compensation with slice-specific step sizes. • Our approach supports continuous rate adaptation without additional parameters. • Weighted training of larger Lagrange multipliers improves performance at all rates. • We validate our methods on CNN, Transformer, and hybrid CNN-Transformer models. Ran Wang 0015, Heming Sun, Jiro Katto |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Toward Multitask Perception for Remote Sensing Imagery via Compression and Prompt TuningabstractRecently, advancements in satellite technology have greatly increased the availability of high-resolution remote sensing images. Concurrently, learning-based image compression (LIC) has significantly improved the efficiency of transmitting and storing such images. As machine recognition tasks increasingly depend on transmitting visual data across devices, compressed images play a key role in both human and machine perception during downstream tasks. However, most LIC approaches are not optimized for machine recognition tasks. To address this limitation, we propose a remote sensing image compression network called RSIC, which integrates multi-task perception and supports downstream tasks such as object detection. Specifically, we introduce a wavelet-based frequency-spatial block (WFSB) that separates frequency components and processes them using Transformer and CNN blocks to effectively capture frequency-specific features. Within WFSB, the Prompting Swin-Transformer Block (PSTB) extracts spatial information while enabling prompt tuning. Additionally, after primary codec training, instance and task prompts are applied during the encoding and decoding stages, respectively, facilitating machine perception without full fine-tuning. Extensive experimental results show that our model achieves better rate-distortion performance for image compression on the AID test dataset, surpassing the traditional VVC codec and several recent LIC methods. Furthermore, our method demonstrates superior performance in terms of rate-accuracy for machine perception on the NWPU VHR-10 and HRSID remote sensing datasets. Feng Liang 0001, Haisheng Fu, Jiro Katto |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Q-LIC: Quantizing Learned Image Compression With Channel SplittingabstractLearned image compression (LIC) has reached a comparable coding gain with traditional hand-crafted methods such as VVC intra. However, the large network complexity prohibits the usage of LIC on resource-limited embedded systems. Network quantization is an efficient way to reduce the network burden. This paper presents a quantized LIC (QLIC) by channel splitting. First, we explore that the influence of quantization error to the reconstruction error is different for various channels. Second, we split the channels whose quantization has larger influence to the reconstruction error. After the splitting, the dynamic range of channels is reduced so that the quantization error can be reduced. Finally, we prune several channels to keep the number of overall channels as origin. By using the proposal, in the case of 8-bit quantization for weight and activation of both main and hyper path, we can reduce the BD-rate by 0.61%-4.74% compared with the previous QLIC. Besides, we can reach better coding gain compared with the state-of-the-art network quantization method when quantizing MS-SSIM models. Moreover, our proposal can be combined with other network quantization methods to further improve the coding gain. The moderate coding loss caused by the quantization validates the feasibility of the hardware implementation for QLIC in the future. Heming Sun, Lu Yu 0003, Jiro Katto |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | SCP: Spherical-Coordinate-Based Learned Point Cloud CompressionabstractIn recent years, the task of learned point cloud compression has gained prominence. An important type of point cloud, LiDAR point cloud, is generated by spinning LiDAR on vehicles. This process results in numerous circular shapes and azimuthal angle invariance features within the point clouds. However, these two features have been largely overlooked by previous methodologies. In this paper, we introduce a model-agnostic method called Spherical-Coordinate-based learned Point cloud compression (SCP), designed to fully leverage the features of circular shapes and azimuthal angle invariance. Additionally, we propose a multi-level Octree for SCP to mitigate the reconstruction error for distant areas within the Spherical-coordinate-based Octree. SCP exhibits excellent universality, making it applicable to various learned point cloud compression techniques. Experimental results demonstrate that SCP surpasses previous state-of-the-art methods by up to 29.14% in point-to-point PSNR BD-Rate. Ao Luo, Linxin Song, Keisuke Nonaka, Kyohei Unno, Heming Sun, Masayuki Goto, Jiro Katto |
AAAI | 7 |
| 2024 | Evaluation of Hardware-based Video Encoders on Modern GPUs for UHD Live-StreamingabstractMany GPUs have incorporated hardware-accelerated video encoders, which allow video encoding tasks to be offloaded from the main CPU and provide higher power efficiency. Over the years, many new video codecs such as H.265/HEVC, VP9, and AV1 were added to the latest GPU boards. Recently, the rise of live video content such as VTuber, game live-streaming, and live event broadcasts, drives the demand for high-efficiency hardware encoders in the GPUs to tackle these real-time video encoding tasks, especially at higher resolutions such as 4K/8K UHD. In this paper, RD performance, encoding speed, as well as power consumption of hardware encoders in several generations of NVIDIA, Intel GPUs as well as Qualcomm Snapdragon Mobile SoCs were evaluated and compared to the software counterparts, including the latest H.266/VVC codec, using several metrics including PSNR, SSIM, and machine-learning based VMAF. The results show that modern GPU hardware encoders can match the RD performance of software encoders in real-time encoding scenarios, and while encoding speed increased in newer hardware, there is mostly negligible RD performance improvement between hardware generations. Finally, the bitrate required for each hardware encoder to match YouTube transcoding quality was also calculated. Kasidis Arunruangsirilert, Jiro Katto |
ICCCN | 2 |
| 2024 | Real-Time Video Prediction With Fast Video Interpolation Model and Prediction TrainingabstractTransmission latency significantly affects users’ quality of experience in real-time interaction and actuation. As latency is principally inevitable, video prediction can be utilized to mitigate the latency and ultimately enable zero-latency transmission. However, most of the existing video prediction methods are computationally expensive and impractical for real-time applications. In this work, we therefore propose real-time video prediction towards the zero-latency interaction over networks, called IFRVP (Intermediate Feature Refinement Video Prediction). Firstly, we propose three training methods for video prediction that extend frame interpolation models, where we utilize a simple convolution-only frame interpolation network based on IFRNet. Secondly, we introduce ELAN-based residual blocks into the prediction models to improve both inference speed and accuracy. Our evaluations show that our proposed models perform efficiently and achieve the best trade-off between prediction accuracy and computational speed among the existing video prediction methods. A demonstration movie is also provided at http://bit.ly/IFRVPDemo. Shota Hirose, Kazuki Kotoyori, Kasidis Arunruangsirilert, Fangzheng Lin, Heming Sun, Jiro Katto |
ICIP | 6 |
| 2024 | Perceptual Quality Driven Point Cloud Compression for 6DoF 3D Point Cloud StreamingabstractA 6 Degree of Freedom (6DoF) 3D point cloud streaming system based on MPEG-Dynamic Adaptive Streaming over HTTP (DASH) is expected to improve immersiveness of communication. 3D point cloud requires a more extensive data size than the 2D and 360-degree video, and significant storage space for the DASH system is required. This paper proposes a perceptual-quality-driven DASH representation generation method to reduce the total data volume of required DASH representations. The evaluation results conclude that the proposed DASH recipe reduces significant server storage space (approximately 75% reduction) while keeping the perceptual quality of 6DoF 3D point cloud streaming compared to the previous bitrate-driven DASH recipe. Yumeka Chujo, Yusuke Tagashira, Yukiko Harada, Kenji Kanai, Jiro Katto |
ISM | 5 |
| 2024 | UplinkNet: Practical Commercial 5G Standalone (SA) Uplink Throughput PredictionabstractWhile 5G New Radio (NR) networks offer significant uplink throughput improvements, these gains are primarily realized when User Equipment (UE) connects to high-frequency millimeter wave (mmWave) bands. The growing demand for uplink-intensive applications, such as real-time UHD 4K/8K video streaming and Virtual Reality (VR)/Augmented Reality (AR) content, highlights the need for accurate uplink throughput prediction to optimize user Quality of Experience (QoE). In this paper, we introduce UplinkNet, a compact neural network designed to predict future uplink throughput using past throughput and RF parameters available through the Android API. With a model size limited to approximately 4,000 parameters, UplinkNet is suitable for IoT and low-power devices. The network was trained on real-world drive test data from commercial 5G Standalone (SA) networks in Tokyo, Japan, and Bangkok, Thailand, across various mobility conditions. To ensure practical implementation, the model uses only Android API data and was evaluated on unseen data against other models. Results show that UplinkNet achieves an average prediction accuracy of 98.9% and an RMSE of 5.22 Mbps, outperforming all other models while maintaining a compact size and low computational cost. Kasidis Arunruangsirilert, Jiro Katto |
VCIP | 2 |
| 2024 | Lightweight Stochastic Video Prediction via Hybrid WarpingabstractAccurate video prediction by deep neural networks, especially for dynamic regions, is a challenging task in computer vision for critical applications such as autonomous driving, remote working, and telemedicine. Due to inherent uncertainties, existing prediction models often struggle with the complexity of motion dynamics and occlusions. In this paper, we propose a novel stochastic long-term video prediction model that focuses on dynamic regions by employing a hybrid warping strategy. By integrating frames generated through forward and backward warpings, our approach effectively compensates for the weaknesses of each technique, improving the prediction accuracy and realism of moving regions in videos while also addressing uncertainty by making stochastic predictions that account for various motions. Furthermore, considering real-time predictions, we introduce a MobileNet-based lightweight architecture into our model. Our model, called SVPHW, achieves state-of-the-art performance on two benchmark datasets. Kazuki Kotoyori, Shota Hirose, Heming Sun, Jiro Katto |
VCIP | 4 |
| 2024 | LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image CompressionabstractSupported by powerful generative models, low-bitrate learned image compression (LIC) models utilizing perceptual metrics have become feasible. Some of the most advanced models achieve high compression rates and superior perceptual quality by using image captions as sub-information. This paper demonstrates that using a large multi-modal model (LMM), it is possible to generate captions and compress them within a single model. We also propose a novel semantic-perceptual-oriented fine-tuning method applicable to any LIC network, resulting in a 41.58% improvement in LPIPS BD-rate compared to existing methods. Our implementation and pre-trained weights are available at https://github.com/tokkiwa/ImageTextCoding. Shimon Murai, Heming Sun, Jiro Katto |
VCIP | 3 |
| 2024 | Variable Bitrate Models For Learned Image Compression with Multi-gain units and Weighted Probability AssignmentabstractWith the advancement of deep learning techniques, learned image compression (LIC) has surpassed traditional compression methods. However, these methods typically require training separate models to achieve optimal rate-distortion performance, leading to increased time and resource consumption. To tackle this challenge, we propose leveraging multi-gain and inverse multi-gain unit pairs to enable variable rate adaptation within a single model. Nevertheless, experiments have shown that rate-distortion performance may degrade at certain bitrates. Therefore, we introduce weighted probability assignment, where different selection probabilities are assigned during training based on lambda values, to increase the model’s training frequency under specific bitrate conditions. To validate our approach, extensive experiments were conducted on Transformer-based and CNN-based models. The experimental results validate the efficiency of our proposed method. Ran Wang 0015, Heming Sun, Jiro Katto |
VCIP | 4 |
| 2024 | Performance Evaluation of Uplink 256QAM on Commercial 5G New Radio (NR) NetworksabstractWhile Uplink 256QAM (UL-256QAM) has been introduced since 2016 as a part of 3GPP Release 14, the adoption was quite poor as many Radio Access Network (RAN) and User Equipment (UE) vendors didn't support this feature. With the introduction of 5G, the support of UL-256QAM has been greatly improved due to a big re-haul of RAN by Mobile Network Operators (MNOs). However, many RAN manufacturers charge MNOs for licenses to enable UL-256QAM per cell basis. This led to some MNOs hesitating to enable the feature on some of their gNodeB or cells to save cost. Since it's known that 256QAM modulation requires a very good channel condition to operate, but UE has a very limited transmission power budget. In this paper, 256QAM utilization, throughput and latency impact from enabling UL-256QAM will be evaluated on commercial 5G Standalone (SA) networks in two countries: Japan and Thailand on various frequency bands, mobility characteristics, and deployment schemes. By modifying the modem firmware, UL-256QAM can be turned off and compared to the conventional UL-64QAM. The results show that UL-256QAM utilization was less than 20% when deployed on a passive antenna network resulting in an average of 8.22% improvement in throughput. However, with Massive MIMO deployment, more than 50% utilization was possible on commercial networks. Furthermore, despite a small uplink throughput gain, enabling UL-256QAM can lower the latency when the link is fully loaded with an average improvement of 7.97 ms in TCP latency observed across various test cases with two TCP congestion control algorithms. Kasidis Arunruangsirilert, Pasapong Wongprasert, Jiro Katto |
WCNC | 3 |
| 2024 | Real-World Performance Evaluations of Low-Band 5G NR/4G LTE 4×4 MIMO on Commercial SmartphonesabstractAll 3GPP-compliant commercial 5G New Radio (NR)-capable UEs on the market are equipped with$4\times 4$MIMO support for Mid-Band frequencies$( > 1.7\text{GHz})$and above, enabling up to rank 4 MIMO transmission. This doubles the theoretical throughput compared to rank 2 MIMO and also improves reception performance. However, 4x4 MIMO support on low-band frequencies$(< 1\text{GHz})$is absent in every commercial UEs, with the exception of the Xperia 1 flagship smartphones manufactured by Sony Mobile and the Xiaomi 14 Pro as of January 2024. The reason most manufacturers omit$4\times 4$MIMO support for low-band frequencies is likely due to design challenges or relatively small performance gains in real-world usage due to the lack of 4T4R deployment on low-band by mobile network operators around the world. In Thailand, 4T4R deployment on the b28/n28 (APT) band is common on True-H and dtac networks, enabling 4x4 MIMO transmission on supported UEs. In this paper, the real-world 4x4 MIMO performance on the b28/n28 (APT) band will be investigated by evaluating the reliability test under different signal conditions and the maximum throughput test by evaluating the performance under optimal conditions, using the Sony Xperia 1 III and the Sony Xperia 1 IV smartphone. Devices from other manufacturers are also used in the experiment to investigate the performance with 2Rx antennas for comparison. Through firmware modifications, the Sony Xperia 1 III and IV can be configured to use only 2 Rx ports on low-band, enabling the collection of comparative 2 Rx performance data as a reference. Pasapong Wongprasert, Kasidis Arunruangsirilert, Jiro Katto |
WCNC | 3 |
| 2023 | Level of Detail-based 3D Space Point Cloud Streaming and its EvaluationabstractDeveloping an efficient quality control method for streaming a 3D space expressed by point clouds is mandatory. Based on this motivation, we develop a Level of Detail (LOD)-based quality optimization method for 3D space point cloud streaming. The proposed method adopts a 3D tile-based method and considers a distance between a user's viewpoint and a 3D tile (i.e., LOD) to maximize the user's perceptual quality. In the experiment, we capture an actual 3D space of our laboratory room using a LiDAR camera and conduct performance evaluations on the Unity platform. The results conclude that the proposed method achieves the highest perceptual quality. Yusuke Tagashira, Yumeka Chujo, Kenji Kanai, Chihiro Nakatsuka, Kyohei Unno, Jiro Katto |
CCNC | 6 |
| 2023 | Learned Image Compression with Mixed Transformer-CNN ArchitecturesabstractLearned image compression (LIC) methods have exhibited promising progress and superior rate-distortion performance compared with classical image compression standards. Most existing LIC methods are Convolutional Neural Networks-based (CNN-based) or Transformer-based, which have different advantages. Exploiting both advantages is a point worth exploring, which has two challenges: 1) how to effectively fuse the two methods? 2) how to achieve higher performance with a suitable complexity? In this paper, we propose an efficient parallel Transformer-CNN Mixture (TCM) block with a controllable complexity to incorporate the local modeling ability of CNN and the non-local modeling ability of transformers to improve the overall architecture of image compression models. Besides, inspired by the recent progress of entropy estimation models and attention modules, we propose a channel-wise entropy model with parameter-efficient swin-transformer-based attention (SWAtten) modules by using channel squeezing. Experimental results demonstrate our proposed method achieves state-of-the-art rate-distortion performances on three different resolution datasets (i.e., Kodak, Tecnick, CLIC Professional Validation) compared to existing LIC methods. The code is at https://github.com/jmliu206/LIC_TCM. Jinming Liu 0001, Heming Sun, Jiro Katto |
CVPR | 3 |
| 2023 | A Predictive Approach for Compensating Transmission Latency in Remote Robot Control for Improving Teleoperation EfficiencyabstractTransmission latency presents a significant challenge when operating remote equipment, such as a robotic arm. To address this, we developed a platform and conducted experiments to reduce transmission latency to near-zero levels. These experiments employed Long Short-Term Memory (LSTM) to anticipate future motion trends, leveraging both the controller's movement variables and Electromyography (EMG) data from the operator's arm muscles. Our findings indicate the potential to decrease transmission latency by approximately 500ms. Additionally, our research confirms a direct correlation between prediction accuracy and the brevity of prediction time, suggesting that shorter prediction times yield more accurate results when using EMG. In the context of video transmission for a remotely located robotic arm, we applied video prediction techniques using the Predictive Coding Network (PredNet) to counter network latency. Our results suggest that these predictive methods can effectively compensate for a latency period of 300ms, thereby highlighting their potential for reducing transmission latency in remote robotic operations. Yutaka Katsuyama, Toshio Sato, Zheng Wen 0001, Xin Qi 0002, Kazuhiko Tamesue, Wataru Kameyama, Yuichi Nakamura 0001, Takuro Sato, Jiro Katto |
GLOBECOM | 9 |
| 2023 | Multistage Spatial Context Models for Learned Image CompressionabstractRecent state-of-the-art Learned Image Compression methods feature spatial context models, achieving great rate-distortion improvements over hyperprior methods. However, the autoregressive context model requires serial decoding, limiting run-time performance. The Checkerboard context model allows parallel decoding at a cost of reduced RD performance. We present a series of multistage spatial context models allowing both fast decoding and better RD performance. We split the latent space into square patches and decode serially within each patch while different patches are decoded in parallel. The proposed method features a comparable decoding speed to Checkerboard while reaching the RD performance of Autoregressive and even also outperforming Autoregressive. Inside each patch, the decoding order must be carefully decided as a bad order negatively impacts performance; therefore, we also propose a decoding order optimization algorithm. Fangzheng Lin, Heming Sun, Jinming Liu 0001, Jiro Katto |
ICASSP | 4 |
| 2023 | Recoil: Parallel rANS Decoding with Decoder-Adaptive ScalabilityabstractEntropy coding is essential to data compression, image and video coding, etc. The Range variant of Asymmetric Numeral Systems (rANS) is a modern entropy coder, featuring superior speed and compression rate. As rANS is not designed for parallel execution, the conventional approach to parallel rANS partitions the input symbol sequence and encodes partitions with independent codecs, and more partitions bring extra overhead. This approach is found in state-of-the-art implementations such as DietGPU. It is unsuitable for content-delivery applications, as the parallelism is wasted if the decoder cannot decode all the partitions in parallel, but all the overhead is still transferred. Fangzheng Lin, Kasidis Arunruangsirilert, Heming Sun, Jiro Katto |
ICPP | 4 |
| 2023 | LSTM-Based GNSS Spoofing Detection for Drone Formation FlightsabstractIn the rapidly evolving logistics industry, drones are becoming indispensable for automated delivery operations. As drone traffic escalates, formation flying is being explored to enhance operational control and increase drone density, thereby reducing the space they occupy. Drones typically rely on Global Navigation Satellite System (GNSS) positioning information for autonomous flight. However, civilian-grade GNSS devices are susceptible to spoofing via Software Defined Radio (SDR), posing significant challenges. In this study, we introduce a novel approach to detect GNSS spoofing by leveraging the multiple GNSS information available from each drone during formation flight. Our investigations, involving two GNSS receivers spoofed by an SDR, reveal that spoofing results in a calculated distance between two receivers that is smaller than the actual value. Capitalizing on this characteristic, we designed simulations of formation flights involving two and five drones. We also developed a GNSS spoofing detection method using the Long Short-Term Memory (LSTM) network. The performance of our spoof detection method was evaluated using simulation data. The results demonstrate that using multiple GNSS data from drones in formation flight significantly enhances performance, achieving an F1 score of 0.96 or higher. This study underscores the potential of our proposed method in improving the security and reliability of drone operations. Zheng Wen 0001, Xin Qi 0002, Toshio Sato, Kazuhiko Tamesue, Yutaka Katsuyama, Kazue Sako, Jiro Katto, Takuro Sato |
IECON | 7 |
| 2023 | PTS-LIC: Pruning Threshold Searching for Lightweight Learned Image CompressionabstractLearned Image Compression (LIC), which uses neural networks to compress images, has experienced significant growth in recent years. The hyperprior-module-based LIC model has achieved higher performance than classical codecs. However, the LIC models are too heavy (in calculation and parameter amounts) to apply to edge devices. To solve this problem, some former papers focus on structural pruning for LIC models. However, they either cause noticeable performance decrement or neglect the appropriate pruning threshold for each LIC model. These problems keep their pruning results sub-optimal. This paper proposes a Pruning Threshold Searching on the hyperprior module for different-quality LIC models. Our method removes most parameters and calculations while keeping the performance the same as the models before pruning. We removed at least 49.8% of parameters and 28.5% of calculations for the Channel-Wise-Context-Model-based models and 29.1% of parameters for the Cheng-2020 models. Ao Luo, Heming Sun, Jinming Liu 0001, Fangzheng Lin, Jiro Katto |
VCIP | 5 |
| 2023 | Performance Evaluations of C-Band 5G NR FR1 (Sub-6 GHz) Uplink MIMO on Urban TrainabstractDue to the recent demand for huge Uplink through-put on Mobile networks driven by the rapid development of social media platforms, UHD 4K/8K video, and VR/AR contents, Uplink MIMO (UL-MIMO) has now been deployed on commercial 5G networks with reasonable availability of supported User Equipment (UE) for consumers. By utilizing up to 2 Tx antenna ports, UL-MIMO-capable UE promised to achieve up to two times the uplink throughput in ideal conditions, while providing improved uplink performance over UE with 1Tx in challenging conditions.In Japan, SoftBank, one of the carriers, introduced 5G Standalone (SA) services for the Fixed Wireless Access (FWA) application back in October 2021. Mobile services were commenced in May 2022, which provide UL-MIMO for supported UE on C-Band or Band n77 (3.7 GHz). In this paper, the uplink performance of UL-MIMO-capable UE will be compared against the conventional UL-1Tx UE on trains, which is the most popular method of transportation for the Japanese. The results show that UL-MIMO-capable UE delivers an average of 19.8% better throughput on moving trains with up to 33.5% in the more favorable signal conditions. A moderate relationship between downlink 5G NR SS-RSRP and uplink throughput also has been observed. Kasidis Arunruangsirilert, Pasapong Wongprasert, Jiro Katto |
WCNC | 3 |
| 2023 | Pensieve 5G: Implementation of RL-based ABR Algorithm for UHD 4K/8K Content Delivery on Commercial 5G SA/NR-DC NetworkabstractWhile the rollout of the fifth-generation mobile network (5G) is underway across the globe with the intention to deliver 4K/8K UHD videos, Augmented Reality (AR), and Virtual Reality (VR) content to the mass amounts of users, the coverage and throughput are still one of the most significant issues, especially in the rural areas, where only 5G in the low-frequency band are being deployed. This called for a highperformance adaptive bitrate (ABR) algorithm that can maximize the user quality of experience given 5G network characteristics and data rate of UHD contents.Recently, many of the newly proposed ABR techniques were machine-learning based. Among that, Pensieve is one of the state-of-the-art techniques, which utilized reinforcement-learning to generate an ABR algorithm based on observation of past decision performance. By incorporating the context of the 5G network and UHD content, Pensieve has been optimized into Pensieve 5G. New QoE metrics that more accurately represent the QoE of UHD video streaming on the different types of devices were proposed and used to evaluate Pensieve 5G against other ABR techniques including the original Pensieve. The results from the simulation based on the real 5G Standalone (SA) network throughput shows that Pensieve 5G outperforms both conventional algorithms and Pensieve with the average QoE improvement of 8.8% and 14.2%, respectively. Additionally, Pensieve 5G also performed well on the commercial 5G NR-NR Dual Connectivity (NR-DC) Network, despite the training being done solely using the data from the 5G Standalone (SA) network. Kasidis Arunruangsirilert, Bo Wei 0001, Hang Song 0001, Jiro Katto |
WCNC | 4 |
| 2023 | RSSI-CSI Measurement and Variation Mitigation With Commodity Wi-Fi DeviceabstractOwing to the plentiful information released by the commodity devices, Wi-Fi signals have been widely studied for various wireless sensing applications. In many works, both received signal strength indicator (RSSI) and the channel state information (CSI) are utilized as the key factors for precise sensing. However, the calculation and relationship between RSSI and CSI is not explained in detail. Furthermore, there are few works focusing on the measurement variation of the Wi-Fi signal which impacts the sensing results. In this article, the relationship between RSSI and CSI is studied in detail and the measurement variation of amplitude and phase information is investigated by extensive experiments. In the experiments, the transmitter and receiver are directly connected by power divider and RF cables and the signal transmission is quantitatively controlled by RF attenuators. By changing the intensity of attenuation, the measurement of RSSI and CSI is carried out under different conditions. From the results, it is found that in order to get a reliable measurement of the signal amplitude and phase by commodity Wi-Fi, the attenuation of the channels should not exceed 60 dB. Meanwhile, the difference between two channels should be lower than 10 dB. An active control mechanism is suggested to ensure the measurement stability. The findings and criteria of this work is promising to facilitate more precise sensing technologies with Wi-Fi signal. Bo Wei 0001, Hang Song 0001, Jiro Katto, Takamaro Kikkawa |
IEEE Internet Things J. | 3 |
| 2022 | Performance Evaluation of Low-Latency Live Streaming of MPEG-DASH UHD video over Commercial 5G NSA/SA Networkabstract5G Standalone (SA) is the goal of the 5G evolution, which aims to provide higher throughput and lower latency than the existing LTE network. One of the main applications of 5G is the real-time distribution of Ultra High-Definition (UHD) content with a resolution of 4K or 8K. In Q2/2021, Advanced Info Service (AIS), the biggest operator in Thailand, launched 5G SA, providing both 5G SA/NSA service nationwide in addition to the existing LTE network. While many parts of the world are still in process of rolling out the first phase of 5G in Non-Standalone (NSA) mode, 5G SA in Thailand already covers more than 76% of the population. In this paper, UHD video will be a real-time live streaming via MPEG-DASH over different mobile network technologies with minimal buffer size to provide the lowest latency. Then, performance such as the number of dropped segments, MAC throughput, and latency are evaluated in various situations such as stationary, moving in the urban area, moving at high speed, and also an ideal condition with maximum SINR. It has been found that 5G SA can deliver more than 95% of the UHD video segment successfully within the required time window in all situations, while 5G NSA produced mixed results depending on the condition of the LTE network. The result also reveals that the LTE network failed to deliver more than 20 % of the video segment within the deadline, which shows that 5G SA is absolutely necessary for low-latency UHD video streaming and 5G NSA may not be good enough for such task as it relies on the legacy control signal. Kasidis Arunruangsirilert, Bo Wei 0001, Hang Song 0001, Jiro Katto |
ICCCN | 4 |
| 2022 | Streaming-Capable High-Performance Architecture of Learned Image Compression CodecsabstractLearned image compression allows achieving state-of-the-art accuracy and compression ratios, but their relatively slow runtime performance limits their usage. While previous attempts on optimizing learned image codecs focused more on the neural model and entropy coding, we present an alternative method to improving the runtime performance of various learned image compression models. We introduce multi-threaded pipelining and an optimized memory model to enable GPU and CPU workloads’ asynchronous execution, fully taking advantage of computational resources. Our architecture alone already produces excellent performance without any change to the neural model itself. We also demonstrate that combining our architecture with previous tweaks to the neural models can further improve runtime performance. We show that our implementations excel in throughput and latency compared to the baseline and demonstrate the performance of our implementations by creating a real-time video streaming encoder-decoder sample application, with the encoder running on an embedded device. Fangzheng Lin, Heming Sun, Jiro Katto |
ICIP | 3 |
| 2022 | Memory-Efficient Learned Image Compression with Pruned Hyperprior ModuleabstractLearned Image Compression (LIC) gradually became more and more famous in these years. The hyperprior-module-based LIC models have achieved remarkable rate-distortion performance. However, the memory cost of these LIC models is too large to actually apply them to various devices, especially to portable or edge devices. The parameter scale is directly linked with memory cost. In our research, we found the hyperprior module is not only highly over-parameterized, but also its latent representation contains redundant information. Therefore, we propose a novel pruning method named ERHP in this paper to efficiently reduce the memory cost of hyperprior module, while improving the network performance. The experiments show our method is effective, reducing at least 22.6% parameters in the whole model while achieving better rate-distortion performance. Ao Luo, Heming Sun, Jinming Liu 0001, Jiro Katto |
ICIP | 4 |
| 2022 | Improving Multiple Machine Vision Tasks in the Compressed DomainabstractThere is a growing number of images that are analyzed by machines rather than just humans. Recently, most machine vision tasks are based on decoded images which require an image compression (encoding/decoding) framework. However, using the decoded images in the pixel-domain has two drawbacks: 1) the complexity is high for the decoder part, 2) the accuracy (e.g., mIoU, mean absolute error, and average precision) of machine vision tasks will be degraded since decoded images only aim to optimize the human perceived quality (e.g., PSNR) so that information required for machine vision tasks will be lost during the decoding process. In this paper, we improve the machine vision tasks in the compressed domain. 1) A gate module is utilized to effectively select some compressed-domain features. 2) Knowledge distillation is introduced to improve the accuracy. 3) A training strategy is explored to support multiple tasks including the image compression. The experimental results show that we can achieve better rate-accuracy/distortion and lower complexity compared with the state-of-the-art pixel-domain work that can take both machine and human vision tasks. Jinming Liu 0001, Heming Sun, Jiro Katto |
ICPR | 3 |
| 2022 | Fast Intra Mode Decision for VVC Based on Histogram of Oriented GradientabstractThe latest Versatile Video Coding (VVC) standard incorporates a series of effective and complex new intra coding tools, which obtains superior coding efficiency than the High Efficiency Video Coding (HEVC). However, this makes the intra coding more complicated and time-consuming. A fast algorithm for VVC is proposed from two aspects of model selection and early terminating to reduce coding complexity in this paper. The relationship between bins with HOG and intra modes is created for the mode selection, decreasing the planar modes for SATD and RDO. Moreover, we analyze the maximum bins to determine the final modes, and we use the modes of left and upper blocks as a reference for the current CU, which can early terminate RDO. The proposed algorithm is implemented on VVC test model, and the experimental results show that it can achieve 36.61% time savings with only 0.94% BDBR increases averagely, which outperforms other relative existing state-of-the-art methods. Aorui Gou, Heming Sun, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
ISCAS | 3 |
| 2022 | A QP-adaptive Mechanism for CNN-based Filter in Video CodingabstractConvolutional neural network (CNN)-based in-loop filtering have been very successful in video coding. For most existing works, however, a specific model was required for each quantization parameter (QP) band. In this paper, we introduce a generic method for helping CNN-filters deal with variable quantization noises. A feasible solution to this problem can be implemented on CNN by introducing a quantization step (Qstep) into the CNN. As the quantization noise changes, the CNN filter’s ability to suppress noise changes accordingly. The (vanilla) convolution layer can be replaced directly by this method in existing CNN filters. Compared with the VVenC anchor, only one CNN filter is used and achieves about 3.6% BD-rate reduction for the luminance component of random-access configuration. Also, about 0.8% BD-rate reduction has been achieved compared with the previous QP-map method. Chao Liu 0027, Heming Sun, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
ISCAS | 3 |
| 2022 | Semantic Segmentation In Learned Compressed DomainabstractMost machine vision tasks (e.g., semantic segmentation) are based on images encoded and decoded by image compression algorithms (e.g., JPEG). However, these decoded images in the pixel domain introduce distortion, and they are optimized for human perception, making the performance of machine vision tasks suboptimal. In this paper, we propose a method based on the compressed domain to improve segmentation tasks. i) A dynamic and a static channel selection method are proposed to reduce the redundancy of compressed representations that are obtained by encoding. ii) Two different transform modules are explored and analyzed to help the compressed representation be transformed as the features in the segmentation network. The experimental results show that we can save up to 15.8% bitrates compared with a state-of-the-art compressed domain-based work while saving up to about 83.6% bitrates and 44.8% inference time compared with the pixel domain-based method. Jinming Liu 0001, Heming Sun, Jiro Katto |
PCS | 3 |
| 2022 | Improving Latent Quantization of Learned Image Compression with Gradient ScalingabstractLearned image compression (LIC) has shown its superior compression ability. Quantization is an inevitable stage to generate quantized latent for the entropy coding. To solve the non-differentiable problem of quantization in the training phase, many differentiable approximated quantization methods have been proposed. However, the derivative of quantized latent to non-quantized latent are set as one in most of the previous methods. As a result, the quantization error between non-quantized and quantized latent is not taken into consideration in the gradient descent. To address this issue, we exploit the gradient scaling method to scale the gradient of non-quantized latent in the back-propagation. The experimental results show that we can outperform the recent LIC quantization methods. Heming Sun, Lu Yu 0003, Jiro Katto |
VCIP | 3 |
| 2022 | Real-time Learned Image Codec on FPGAabstractThis demo paper gives a real-time learned image codec on FPGA. By using Xilinx VCU128, the proposed system reaches 720P@30fps codec, which is 7.76x faster than prior work. Heming Sun, Qingyang Yi, Fangzheng Lin, Lu Yu 0003, Jiro Katto |
VCIP | 5 |
| 2022 | Pilot Allocation Optimization using Digital Annealer for Multi-cell Massive MIMOabstractFor massive multiple-input multiple-output (MIMO) systems, pilot contamination reduces the data transmission capacity owing to the inter-cell interference of non-orthogonal pilots reusage. To develop efficient mobile communication, it is necessary to mitigate pilot contamination. To address this problem, we propose an annealing-based pilot allocation method using Digital Annealer to provide solution for Ising machine. The proposed method is a max k-cut-based approach, where the graph represents the potential strength of pilot contamination among users in other cells. By using this proposed method, users who have strong relationship with pilot contamination will be assigned different pilots. Experiment results show that the proposed method can realize optimal pilot allocation and mitigate pilot contamination. Compared with conventional methods, the proposal achieves the best performance which can increase the minimum achievable rate and show higher SINR, especially when the numbers of users and cells are large. Daiki Maruyama, Bo Wei 0001, Hang Song 0001, Jiro Katto |
WCNC | 4 |
| 2022 | QA-Filter: A QP-Adaptive Convolutional Neural Network Filter for Video CodingabstractConvolutional neural network (CNN)-based filters have achieved great success in video coding. However, in most previous works, individual models were needed for each quantization parameter (QP) band, which is impractical due to limited storage resources. To explore this, our work consists of two parts. First, we propose a frequency and spatial QP-adaptive mechanism (FSQAM), which can be directly applied to the (vanilla) convolution to help any CNN filter handle different quantization noise. From the frequency domain, a FQAM that introduces the quantization step (Qstep) into the convolution is proposed. When the quantization noise increases, the ability of the CNN filter to suppress noise improves. Moreover, SQAM is further designed to compensate for the FQAM from the spatial domain. Second, based on FSQAM, a QP-adaptive CNN filter called QA-Filter that can be used under a wide range of QP is proposed. By factorizing the mixed features to high-frequency and low-frequency parts with the pair of pooling and upsampling operations, the QA-Filter and FQAM can promote each other to obtain better performance. Compared to the H.266/VVC baseline, average 5.25% and 3.84% BD-rate reductions for luma are achieved by QA-Filter with default all-intra (AI) and random-access (RA) configurations, respectively. Additionally, an up to 9.16% BD-rate reduction is achieved on the luma of sequence BasketballDrill. Besides, FSQAM achieves measurably better BD-rate performance compared with the previous QP map method. Chao Liu 0027, Heming Sun, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Image Process. | 3 |
| 2021 | Performance Evaluations of Channel Estimation Using Deep-learning Based Super-resolutionabstractThanks to breakthrough and evolution of deep learning in computer vision areas, adaptation of deep learning into communication systems are getting lots of attention to researchers. Recently, a channel estimation method by using a deep learning-based image super-resolution (SR) technique, namely ChannelNet, has been proposed. Inspired by this research, in this paper, we propose a deep SR based channel estimation method by applying more accurate deep learning-based SR network architecture, EDSR. In order to enhance intelligibility and reliability of deep SR based channel estimation methods, we evaluate the performance of several deep SR based channel estimation methods (SRCNN, ChannelNet and EDSR) by carrying out practical 5G simulations. From the evaluations, the results conclude that the deep SR based channel estimation methods can potentially improve accuracy of channel estimation and reduce BER characteristics. Daiki Maruyama, Kenji Kanai, Jiro Katto |
CCNC | 3 |
| 2021 | IoT-centric Service Function ChainingOrchestration and its Performance ValidationabstractIn order to simplify deployment and management of IoT services, Network Function Virtualization (NFV) and Service Function Chaining (SFC) are promising solutions, and much researchers have conducted these topics. To enhance the reliability of former research efforts, in this paper, we propose an orchestration framework for IoT-centric SFC by using Docker and Kubernetes. The framework enables an automatic IoT service deployment by satisfying service requirements and computing and network resource constraints. In such deployment, we apply a Virtual Network Function (VNF)/Service Function (SF) placement problem to achieve efficient utilization of the resources. We set an objective function as minimizing both numbers of SF instances and communications and build a mathematical model based on Integer Linear Programming (ILP). To validate it, we implement a model for the framework and evaluate the performances by carrying out a numerical evaluation and a real experiment. From the evaluation results, we confirm that the proposed approach can reduce the number of SF placements and the number of communications among SF instances. Hibiki Sekine, Kenji Kanai, Jiro Katto, Hidehiro Kanemitsu, Hidenori Nakazato |
CCNC | 3 |
| 2021 | FRAB: A Flexible Relaxation Method for Fair, Stable, Efficient Multi-user DASH Video StreamingabstractDynamic adaptive streaming over HTTP (DASH) has been widely adopted in modern video streaming services. In DASH, the core technique is adaptive bitrate (ABR) control which can adjust the requested video bitrate level according to the network conditions to tradeoff between video quality and rebuffering risk. It is a challenge for the ABR methods in the scenarios when multiple DASH streaming users compete over the network bottleneck. This paper proposes a client-side ABR control method, flexible relaxation assisted by buffer (FRAB), to achieve fair, stable and efficient video streaming among different users. The idea of FRAB is to "relax" the change of the video quality based on current buffer level, which can enhance the stability of video streaming. Meanwhile, by flexibly adjusting the relaxation, the efficiency and fairness among all users are improved. FRAB is evaluated in real experiments under three different network conditions and compared with conventional multi-user ABR algorithms. Results indicate FRAB has the best performance in fairness, which reduces the unfairness by a maximum of 69.5% under real-world measured network condition. It also improves the efficiency by 71.3% comparing with PANDA, and enhances the stability by 73.3% comparing with TFDASH. The experiment results demonstrated that the proposed method has superior performances in multi-user DASH video streaming. Bo Wei 0001, Hang Song 0001, Jiro Katto |
ICC | 3 |
| 2021 | Deep Pedestrian Density Estimation For Smart City MonitoringabstractRecently, requirement of city monitoring and maintenance using ICT techniques increases with the help of transportation system. In addition, the spread of COVID-19 has increased the demand for managing pedestrian traffic volume. To contribute to these trends, in this paper, we propose a new pedestrian radar map system in order to estimate pedestrian density on streets and sidewalks. Our system uses e-bikes to collect 360-degree images and visualize pedestrian positions as a radar map. In evaluations, we confirm the accuracies of the radar maps and pedestrian density by using KITTI dataset and by carrying out a field experiment. Kazuki Murayama, Kenji Kanai, Masaru Takeuchi, Heming Sun, Jiro Katto |
ICIP | 5 |
| 2021 | Approximated Reconfigurable Transform Architecture for VVCabstractAs the demand for high-resolution videos grows, the next generation video coding standard Versatile Video Coding introduces many new proposals, including Adaptive Multiple Transforms (AMT), to improve coding efficiency. This paper presents a reconfigurable transform core for the VVC standard where the implementation of 1D DST-VII and DCT-VIII for all transform sizes are enabled. To offer a very low circuit complexity, a simple approximation strategy with a little coding performance loss is proposed. An 8x8 Processing Element (PE) array is employed as the core computational unit, where each PE can be configured dynamically based on the transform type. In addition, the transforms of larger sizes can be realized in the finite PE units with the Partitioned Matrix Multiplication (PMM) scheme. The experimental and synthesis results show that this design can save at least 29.1% area compared with other works in literature with the negligible degradation of video quality and a slight increase in the bit rate. Yixuan Zeng, Heming Sun, Jiro Katto, Yibo Fan |
ISCAS | 3 |
| 2021 | Accelerating Convolutional Neural Network Inference Based on a Reconfigurable Sliced Systolic ArrayabstractConvolutional neural networks (CNNs) have achieved great successes on many computer vision tasks, such as image recognition, video processing, and target detection. In recent years, many hardware designs have been devoted to accelerating CNN inference. In order to further speed up CNN inference and reduce data waste, this work proposed a reconfigurable sliced systolic array: 1) Depending on the number of network nodes in each layer, the slice mode could be dynamically configured to achieve high throughput and resource utilization. 2) To take full advantage of convolution reuse and weight reuse, this work designed a tile-column sliding (TCS) processing dataflow. 3) A four-stage for loop algorithm was employed, which divides the CNN calculation into several parts based on the input nodes and output nodes. The entire CNN inference is carried out using integer-only arithmetic originated from TensorLite. Experimental results prove that these strategies lead to significant improvement in inference performance and energy efficiency. Yixuan Zeng, Heming Sun, Jiro Katto, Yibo Fan |
ISCAS | 3 |
| 2021 | High-QoE DASH Live Streaming Using Reinforcement LearningabstractWith the live video streaming becomes more and more common in daily life such as live meeting and live video call, it is an urgent task to ensure high-quality and low-delay live video streaming service. High user quality of experience (QoE) should be ensured to satisfy the requirement of user, for which latency is one of the important factors. In this paper, a high-QoE live streaming method is proposed with reinforcement learning. Experiments are conducted to evaluate the proposed method. Results demonstrate that the proposal shows the best performance with highest QoE compared with conventional methods in three network conditions. In Ferry case, the QoE is almost twice of the QoE of other methods. Bo Wei 0001, Hang Song 0001, Jiro Katto |
IWQoS | 3 |
| 2021 | Learned Image Compression with Fixed-point ArithmeticabstractLearned image compression (LIC) has achieved superior coding performance than traditional image compression standards such as HEVC intra in terms of both PSNR and MS-SSIM. However, most LIC frameworks are based on floating-point arithmetic which has two potential problems. First is that using traditional 32-bit floating-point will consume huge memory and computational cost. Second is that the decoding might fail because of the floating-point error coming from different encoding/decoding platforms. To solve the above two problems. 1) We linearly quantize the weight in the main path to 8-bit fixed-point arithmetic, and propose a fine tuning scheme to reduce the coding loss caused by the quantization. Analysis transform and synthesis transform are fine tuned layer by layer. 2) We exploit look-up-table (LUT) for the cumulative distribution function (CDF) to avoid the floating-point error. When the latent node follows non-zero mean Gaussian distribution, to share the CDF LUT for different mean values, we restrict the range of latent node to be within a certain range around mean. As a result, 8-bit weight quantization can achieve negligible coding gain loss compared with 32-bit floating-point anchor. In addition, proposed CDF LUT can ensure the correct coding at various CPU and GPU hardware platforms. Heming Sun, Lu Yu 0003, Jiro Katto |
PCS | 3 |
| 2021 | Adaptive Video Transmission Strategy Based on Ising MachineabstractWith the dramatically increasing video streaming in the total network traffic, it is critical to develop effective algorithms to ensure the quality of content delivery service. Adaptive bitrate (ABR) control is the most essential technique which determines the proper bitrate to be chosen based on network conditions, thus realize high-quality video streaming. In this paper, a novel ABR strategy is proposed based on Ising machine by using the quadratic unconstrained binary optimization (QUBO) method and Digital Annealer (DA) for the first time. The proposed method is evaluated by simulation with the real-world measured throughput and compared with other state-of-the-art methods. Experiment results show that the proposed QUBO-based method can outperform the existing methods, which demonstrating the superior of the proposed QUBO-based method. Bo Wei 0001, Hang Song 0001, Jiro Katto |
SenSys | 3 |
| 2021 | Learning in Compressed Domain for Faster Machine Vision TasksabstractLearned image compression (LIC) has illustrated good ability for reconstruction quality driven tasks (e.g. PSNR, MS-SSIM) and machine vision tasks such as image understanding. However, most LIC frameworks are based on pixel domain, which requires the decoding process. In this paper, we develop a learned compressed domain framework for machine vision tasks. 1) By sending the compressed latent representation directly to the task network, the decoding computation can be eliminated to reduce the complexity. 2) By sorting the latent channels by entropy, only selective channels will be transmitted to the task network, which can reduce the bitrate. As a result, compared with the traditional pixel domain methods, we can reduce about 1/3 multiply-add operations (MACs) and 1/5 inference time while keeping the same accuracy. Moreover, proposed channel selection can contribute to at most 6.8% bitrate saving. Jinming Liu 0001, Heming Sun, Jiro Katto |
VCIP | 3 |
| 2021 | Performance Analysis of Adaptive Bitrate Algorithms for Multi-user DASH Video StreamingabstractWith the increasing video demand in daily network traffic, it is an urgent task to develop effective algorithms to facilitate high-quality content delivery service. Recently, numerous adaptive streaming algorithms have been proposed to improve the user perceived experience. However, these algorithms were mainly developed from the perspective of single user. There is not yet systematical evaluation and comparison of the bitrate adaptation methods for multi-user video streaming. Besides, the Quality of Experience (QoE) metrics were not unified.In this work, we propose a new mininet-based testbed framework which is able to conduct real-time video streaming emulation in various multi-user scenarios. Seven state-of-the-art adaptation methods are incorporated into the testbed. Meanwhile, ITU-T P.1203 model, the world's first standard for measuring QoE of HTTP adaptive streaming, is implemented to calculate the mean opinion scores of different methods. Using the developed testbed, the performance of current adaptation methods in multi-user network is analyzed and compared. A variety of experiments are carried out by changing the user number and network conditions, in which the QoE of different users are investigated. It is found that current algorithms perform inconsistently in various network scenarios. In the excessive user and limited bandwidth cases, machine learning and scheduling techniques show superiority in providing high and equal QoE for all users. While in the high-delay case, the buffer-based approaches show robust performance. Overall, the findings of this work give an insight for designing and choosing adaptive streaming strategies in different multi-user network conditions. Bo Wei 0001, Hang Song 0001, Shangguang Wang, Jiro Katto |
WCNC | 4 |
| 2021 | A containerized task clustering for scheduling workflows to utilize processors and containers on clouds
Hidehiro Kanemitsu, Kenji Kanai, Jiro Katto, Hidenori Nakazato |
J. Supercomput. | 3 |
| 2020 | Field Experiments of 28 GHz Band 5G System at Indoor Train Station PlatformabstractRecently, a fifth-generation cellular system (5G) is widely expected to provide plenty of wireless network resources (i.e., broadband capacity). In this paper, to validate 5G system performances, such as physical-layer and TCP-layer throughputs, we carry out a field trial at an actual indoor train station, named Haneda International Airport Terminal Station. In the field trial, we deploy the prototype 5G system (Central Unit, Distribution Unit, Radio Unit and 5G UE (tablet)) on the train station platform and evaluate mobile 5G downlink throughputs. Through the actual measurements, the results confirm that the prototype 5G system can achieve mobile broadband capacity (more than 1 Gbps) even when the UE is located anywhere at the indoor train station platform. Mayuko Okano, Yohei Hasegawa, Kenji Kanai, Bo Wei 0001, Jiro Katto |
CCNC | 5 |
| 2020 | Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention ModulesabstractImage compression is a fundamental research field and many well-known compression standards have been developed for many decades. Recently, learned compression methods exhibit a fast development trend with promising results. However, there is still a performance gap between learned compression algorithms and reigning compression standards, especially in terms of widely used PSNR metric. In this paper, we explore the remaining redundancy of recent learned compression algorithms. We have found accurate entropy models for rate estimation largely affect the optimization of network parameters and thus affect the rate-distortion performance. Therefore, in this paper, we propose to use discretized Gaussian Mixture Likelihoods to parameterize the distributions of latent codes, which can achieve a more accurate and flexible entropy model. Besides, we take advantage of recent attention modules and incorporate them into network architecture to enhance the performance. Experimental results demonstrate our proposed method achieves a state-of-the-art performance compared to existing learned compression methods on both Kodak and high-resolution datasets. To our knowledge our approach is the first work to achieve comparable performance with latest compression standard Versatile Video Coding (VVC) regarding PSNR. More importantly, our approach generates more visually pleasant results when optimized by MS-SSIM. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
CVPR | 4 |
| 2020 | Learned Lossless Image Compression with A Hyperprior and Discretized Gaussian Mixture LikelihoodsabstractLossless image compression is an important task in the field of multimedia communication. Traditional image codecs typically support lossless mode, such as WebP, JPEG2000, FLIF. Recently, deep learning based approaches have started to show the potential at this point. HyperPrior is an effective technique proposed for lossy image compression. This paper generalizes the hyperprior from lossy model to lossless compression, and proposes a L2-norm term into the loss function to speed up training procedure. Besides, this paper also investigated different parameterized models for latent codes, and propose to use Gaussian mixture likelihoods to achieve adaptive and flexible context models. Experimental results validate our method can outperform existing deep learning based lossless compression, and outperform the JPEG2000 and WebP for JPG images. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
ICASSP | 4 |
| 2020 | Scalable Learned Image Compression With A Recurrent Neural Networks-Based HyperpriorabstractRecently learned image compression has achieved many great progresses, such as representative hyperprior and its variants based on convolutional neural networks (CNNs). However, CNNs are not fit for scalable coding and multiple models need to be trained separately to achieve variable rates. In this paper, we incorporate differentiable quantization and accurate entropy models into recurrent neural networks (RNNs) architectures to achieve a scalable learned image compression. First, we present an RNN architecture with quantization and entropy coding. To realize the scalable coding, we allocate the bits to multiple layers, by adjusting the layer-wise lambda values in Lagrangian multiplier-based rate-distortion optimization function. Second, we add an RNN-based hyperprior to improve the accuracy of entropy models for multiple-layer residual representations. Experimental results demonstrate that our performance can be comparable with recent CNN-based hyperprior methods on Kodak dataset. Besides, our method is a scalable and flexible coding approach, to achieve multiple rates using one single model, which is very appealing. Rige Su, Zhengxue Cheng, Heming Sun, Jiro Katto |
ICIP | 4 |
| 2020 | End-To-End Learned Image Compression With Fixed Point Weight QuantizationabstractLearned image compression (LIC) has reached the traditional hand-crafted methods such as JPEG2000 and BPG in terms of the coding gain. However, the large model size of the network prohibits the usage of LIC on resource-limited embedded systems. This paper presents a LIC with 8-bit fixed-point weights. First, we quantize the weights in groups and propose a non-linear memory-free codebook. Second, we explore the optimal grouping and quantization scheme. Finally, we develop a novel weight clipping fine tuning scheme. Experimental results illustrate that the coding loss caused by the quantization is small, while around 75% model size can be reduced compared with the 32-bit floating-point anchor. As far as we know, this is the first work to explore and evaluate the LIC fully with fixed-point weights, and our proposed quantized LIC is able to outperform BPG in terms of MS-SSIM. Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto |
ICIP | 4 |
| 2020 | Fully Neural Network Mode Based Intra Prediction of Variable Block SizeabstractIntra prediction is an essential component in the image coding. This paper gives an intra prediction framework completely based on neural network modes (NM). Each NM can be regarded as a regression from the neighboring reference blocks to the current coding block. (1) For variable block size, we utilize different network structures. For small blocks 4×4 and 8×8, fully connected networks are used, while for large blocks 16×16 and 32×32, convolutional neural networks are exploited. (2) For each prediction mode, we develop a specific pre-trained network to boost the regression accuracy. When integrating into HEVC test model, we can save 3.55%, 3.03% and 3.27% BD-rate for Y, U, V components compared with the anchor. As far as we know, this is the first work to explore a fully NM based framework for intra prediction, and we reach a better coding gain with a lower complexity compared with the previous work. Heming Sun, Lu Yu 0003, Jiro Katto |
VCIP | 3 |
| 2020 | A Pipelined 2D Transform Architecture Supporting Mixed Block Sizes for the VVC StandardabstractFor the next-generation video coding standard Versatile Video Coding (VVC), several new contributions have been proposed to improve the coding efficiency, especially in the transformation operations. This paper proposes a unified $32\times 32$ block-based transform architecture for the VVC standard that enables 2D Discrete Sine Transform-VII (DST-VII) and Discrete Cosine Transform-VIII (DCT-VIII) of all sizes. It mainly gives three contributions: 1) The N-Dimensional Reduced Adder Graph (RAG-n) algorithm is adopted to design the minimal adder-oriented computational units. 2) The storage of the asymmetric transform units can be realized in the dual-port SRAM-based transpose memory. 3) The pipelined 2D transformations of mixed block sizes are achieved with the throughput rate of 32 samples per cycle. The synthesis results indicate that this architecture can reduce area by up to 73.1% compared with other state-of-the-art works. Moreover, power saving ranging from 4.9% to 9.9% can be achieved. Regarding the transpose memory, at least 21.9% of the area can be saved by using SRAM. Yibo Fan, Yixuan Zeng, Heming Sun, Jiro Katto, Xiaoyang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Energy Compaction-Based Image Compression Using Convolutional AutoEncoderabstractImage compression has been an important research topic for many decades. Recently, deep learning has achieved great success in many computer vision tasks, and its use in image compression has gradually been increasing. In this paper, we present an energy compaction-based image compression architecture using a convolutional autoencoder (CAE) to achieve high coding efficiency. Our main contributions include three aspects: 1) we propose a CAE architecture for image compression by decomposing it into several down(up)sampling operations; 2) for our CAE architecture, we offer a mathematical analysis on the energy compaction property and we are the first work to propose a normalized coding gain metric in neural networks, which can act as a measurement of compression capability; 3) based on the coding gain metric, we propose an energy compaction-based bit allocation method, which adds a regularizer to the loss function during the training stage to help the CAE maximize the coding gain and achieve high compression efficiency. The experimental results demonstrate our proposed method outperforms BPG (HEVC-intra), in terms of the MS-SSIM quality metric. Additionally, we achieve better performance in comparison with existing bit allocation methods, and provide higher coding efficiency compared with state-of-the-art learning compression methods at high bit rates. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
IEEE Trans. Multim. | 4 |
| 2020 | Enhanced Intra Prediction for Video Coding by Using Multiple Neural NetworksabstractThis paper enhances the intra prediction by using multiple neural network modes (NM). Each NM serves as an end-to-end mapping from the neighboring reference blocks to the current coding block. For the provided NMs, we present two schemes (appending and substitution) to integrate the NMs with the traditional modes (TM) defined in high efficiency video coding (HEVC). For the appending scheme, each NM is corresponding to a certain range of TMs. The categorization of TMs is based on the expected prediction errors. After determining the relevant TMs for each NM, we present a probability-aware mode signaling scheme. The NMs with higher probabilities to be the best mode are signaled with fewer bits. For the substitution scheme, we propose to replace the highest and lowest probable TMs. New most probable mode (MPM) generation method is also employed when substituting the lowest probable TMs. Experimental results demonstrate that using multiple NMs will improve the coding efficiency apparently compared with the single NM. Specifically, proposed appending scheme with seven NMs can save 2.6%, 3.8%, and 3.1% BD-rate for Y, U, and V components compared with using single NM in the state-of-the-art works. Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto |
IEEE Trans. Multim. | 4 |
| 2019 | A Function Clustering Algorithm for Resource Utilization in Service Function ChainingabstractVirtualized service and network functions are deployed on virtual machines (VMs) to realize essential processing to realize service function chaining (SFC). Issues on SFC is SF allocation to a VM and to minimize the response time and number of function instances. In this paper, we propose an SF clustering-based scheduling algorithm, called "SF-clustering for utilizing virtual CPUs" (SFCUV), to solve the SF allocation and SF selection problems simultaneously. Experimental results show that SF-CUV can utilize vCPUs to minimize the response time. Hidehiro Kanemitsu, Kenji Kanai, Jiro Katto, Hidenori Nakazato |
CLOUD | 3 |
| 2019 | Learning Image and Video Compression Through Spatial-Temporal Energy CompactionabstractCompression has been an important research topic for many decades, to produce a significant impact on data transmission and storage. Recent advances have shown a great potential of learning based image and video compression. Inspired from related works, in this paper, we present an image compression architecture using a convolutional autoencoder, and then generalize image compression to video compression, by adding an interpolation loop into both encoder and decoder sides. Our basic idea is to realize spatial-temporal energy compaction in learning image and video compression. Thereby, we propose to add a spatial energy compaction-based penalty into loss function, to achieve higher image compression performance. Furthermore, based on temporal energy distribution, we propose to select the number of frames in one interpolation loop, adapting to the motion characteristics of video contents. Experimental results demonstrate that our proposed image compression outperforms the latest image compression standard with MS-SSIM quality metric, and provides higher performance compared with state-of-the-art learning compression methods at high bit rates, which benefits from our spatial energy compaction approach. Meanwhile, our proposed video compression approach with temporal energy compaction can significantly outperform MPEG-4, and is competitive with commonly used H.264. Both our image and video compression can produce more visually pleasant results than traditional standards. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
CVPR | 4 |
| 2019 | Perceptual Quality Study on Deep Learning Based Image CompressionabstractRecently deep learning based image compression has made rapid advances with promising results based on objective quality metrics. However, a rigorous subjective quality evaluation on such compression schemes have rarely been reported. This paper aims at perceptual quality studies on learned compression. First, we build a general learned compression approach, and optimize the model. In total six compression algorithms are considered for this study. Then, we perform subjective quality tests in a controlled environment using high-resolution images. Results demonstrate learned compression optimized by MS-SSIM yields competitive results that approach the efficiency of state-of-the-art compression. The results obtained can provide a useful benchmark for future developments in learned image compression. Zhengxue Cheng, Pinar Akyazi, Heming Sun, Jiro Katto, Touradj Ebrahimi |
ICIP | 4 |
| 2019 | A Gamut-Extension Method Considering Color Information Restoration using Convolutional Neural NetworksabstractRecently, Ultra HDTV (UHDTV) services become popular over satellite and on the internet. On the contrary, there are tremendously huge volume of High Definition Television (HDTV) and Standard Definition Television (SDTV) contents stored in broadcasting companies and storage devices. In this paper, we propose a color space conversion (also known as gamut mapping) method from BT. 709 (used for current HDTV broadcast) to BT. 2020 (used for UHDTV broadcast), which estimates and restores lost color information. It learns an end-to-end conversion method from BT. 709 image to BT. 2020 image with restoring lost color information using Convolutional Neural Network (CNN). By experiments, we confirm that our method can achieve 2.31dB gain against the conventional method on average. Masaru Takeuchi, Yusuke Sakamoto, Ryota Yokoyama, Heming Sun, Yasutaka Matsuo, Jiro Katto |
ICIP | 6 |
| 2019 | Performance Evaluations of Viewport Movement Prediction and Rate Adaptation for Tile-Based 360-Degree Video DeliveryabstractRecently, the demand for high quality 360-degree video delivery is increasing, however, 360-degree videos require extremely high bitrate and large network capacity. Therefore, an efficient (i.e., higher quality and lower traffic) 360-degree video delivery is mandatory. To address this fact, this paper introduces and evaluates a tile-based 360-degree video delivery system that equips viewport movement prediction and rate adaptation. Yuya Shinohara, Kenji Kanai, Jiro Katto |
ISM | 3 |
| 2019 | Dual Learning-based Video Coding with Inception Dense BlocksabstractIn this paper, a dual learning-based method in intra coding is introduced for PCS Grand Challenge. This method is mainly composed of two parts: intra prediction and reconstruction filtering. They use different network structures, the neural network-based intra prediction uses the full-connected network to predict the block while the neural network-based reconstruction filtering utilizes the convolutional networks. Different with the previous filtering works, we use a network with more powerful feature extraction capabilities in our reconstruction filtering network. And the filtering unit is the block-level so as to achieve a more accurate filtering compensation. To our best knowledge, among all the learning-based methods, this is the first attempt to combine two different networks in one application, and we achieve the state-of-the-art performance for AI configuration on the HEVC Test sequences. The experimental result shows that our method leads to significant BD-rate saving for provided 8 sequences compared to HM-16.20 baseline (average 10.24% and 3.57% bitrate reductions for all-intra and random-access coding, respectively). For HEVC test sequences, our model also achieved a 9.70% BD-rate saving compared to HM-16.20 baseline for all-intra configuration. Chao Liu 0027, Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
PCS | 6 |
| 2019 | Fast QTMT Partition Decision Algorithm in VVC Intra Coding based on Variance and GradientabstractQuadtree with nested multi-type tree (QTMT) partition structure in Versatile Video Coding (VVC) contributes to superior encoding performance compared to the basic quad-tree (QT) structure in High Efficiency Video Coding (HEVC). However, the improvement of performance leads to an un-avoidable increase of computational complexity. To achieve a balance between coding efficiency and compression quality, we propose a fast intra partition algorithm based on variance and gradient to solve the rectangular partition problem in VVC. First, further splitting of smooth areas is terminated. Then, QT partition is chosen depending on the gradient features extracted by Sobel operator. Finally, one partition from five possible QTMT partitions is directly chosen by computing the variance of variance of sub-CUs. The theoretical basis of our method is that a homogeneous area tends to be predicted with a larger coding unit (CU), and sub-parts of a split CU are prone to have different textures from each other. To our knowledge, this is the first attempt to apply traditional method to accelerating the rectangular partition problem in VVC intra prediction. Experimental results show that the proposed method can save averagely 53.17% encoding time with only 1.62% BDBR increase and 0.09dB BDPSNR loss compared to anchor VTM4.0. Heming Sun, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
VCIP | 3 |
| 2019 | TCP throughput characteristics over 5G millimeterwave network in indoor train stationabstractTo realize highly reliable video surveillance and provide ultrahigh-definition/immersive video streaming, it is planned to adopt the 5G cellular system using millimeter-wave (mmWave) as the wireless-network infrastructure. However, mmWave communication has a challenging issue: mmWave communication is extremely sensitive to obstacles, such as walls, pillars, and even human bodies, and this issue easily increases the packet loss rates and round trip time (RTT) (or disconnection from the base station) due to a no line of sight (NLOS) environment. Therefore, in this work, 5G throughput performances were evaluates in an indoor train station by considering the effect of an NLOS environment caused by blockage by human bodies. In addition, to improve the robustness of TCP transmission in a high-RTT and high-packet-loss environment (e.g., an NLOS environment), a state-of-the-art TCP, TCP-FSO, was used. In the evaluations, the MATLAB 5G library was used to simulate the 5G environment, and a Linux software-based network emulator, Traffic Control, was used to emulate the 5G network. From the evaluations, it the 5G mobile throughput characteristics were confirmed in three different crowded patterns (low, middle, and high density), and the TCP-FSO advantage against CUBIC-TCP was validated. Mayuko Okano, Yohei Hasegawa, Kenji Kanai, Bo Wei 0001, Jiro Katto |
WCNC | 5 |
| 2018 | Performance evaluations of multimedia service function chaining in edge cloudsabstractAs mobile multimedia services have significantly evolved and diversified with the spread of smartphones and Internet of Things (IoT) devices, low-delay multimedia cloud computing is the need of the hour. To address this demand, in this study, we introduce an edge cloud system that equips a multimedia service function chaining capability. A prototype implementation of the proposed edge cloud system has three main features: 1) edge computing deployment by using OpenStack, 2) multimedia service slicing and chaining, and 3) efficient resource management in edge networks. Based on these features, the proposed system achieves lower multimedia processing delay compared to a conventional cloud computing platform. We deploy the proposed system in our laboratory and validate the system performance by using typical multimedia application, such as human detection in video surveillance. Kentaro Imagane, Kenji Kanai, Jiro Katto, Toshitaka Tsuda, Hidenori Nakazato |
CCNC | 3 |
| 2018 | Edge-centric field monitoring system for energy-efficient and network-friendly field sensingabstractTo provide energy-efficient (i.e., longer lifetime of sensors) and network-friendly (i.e., reducing network traffic) field sensing, we propose an edge-centric field monitoring system which applies efficient sensors and camera control. The proposed system detects conditions in a monitoring area and controls sensing frequency (sampling rate) of sensors, and capture rate and encoding rate of surveillance cameras, according to the detected conditions. In addition, the system applies a Multi-access Edge Computing (MEC) platform to provide fast feedback control to the sensors and cameras. In performance evaluations, we assume that the monitoring target is landslide detection and create a miniature “artificial landslide generation” environment in our laboratory. By using the environment, we evaluate the system performance, and evaluation results indicate that the proposed system can reduce network traffic and save energy consumption efficiently. Keigo Ogawa, Kenji Kanai, Masaru Takeuchi, Jiro Katto, Toshitaka Tsuda |
CCNC | 4 |
| 2018 | TRUST: A TCP Throughput Prediction Method in Mobile NetworksabstractThroughput prediction is essential for ensuring high quality of service for video streaming transmissions. However, current methods are incapable of accurately predicting throughput in mobile networks, especially for moving user scenarios. Therefore, we propose a TCP throughput prediction method TRUST using machine learning for mobile networks. TRUST has two stages: user movement pattern identification and throughput prediction. In the prediction stage, the long short-term memory (LSTM) model is employed for TCP throughput prediction. TRUST takes all the communication quality factors, sensor data and scenario information into consideration. Field experiments are conducted to evaluate TRUST in various scenarios. The results indicate that TRUST can predict future throughput with higher accuracy than the conventional methods, which decreases the throughput prediction error by maximum 44% under the moving bus scenario. Bo Wei 0001, Wataru Kawakami, Kenji Kanai, Jiro Katto, Shangguang Wang |
GLOBECOM | 4 |
| 2018 | Machine Learning Based Transportation Modes Recognition Using Mobile Communication QualityabstractIn order to recognize the transportation modes without any additional sensor devices, we propose a recognition method by using communication quality factors. In the proposed method, instead of Global Positioning System (GPS) and accelerometer sensors, we collect mobile TCP throughputs, Received Signal Strength Indicators (RSSIs), and cellular base station IDs (Cell IDs) through in-line network measurement when the user enjoys mobile services, such as video streaming service. In accuracy evaluations, we conduct two different field experiments to collect the data in five typical transportation modes (static, walking, riding a bicycle, a bus and a train,) and then construct the classifiers by applying Support Vector Machine (SVM), k-Nearest Neighbor (k-NN) and Random Forest (RF). Results conclude that these transportation modes can be recognized by using communication quality factors with high accuracy as well as the use of accelerometer sensors. Wataru Kawakami, Kenji Kanai, Bo Wei 0001, Jiro Katto |
ICME | 4 |
| 2018 | Light-Weight Video Coding Based on Perceptual Video Quality for Live StreamingabstractIn video streaming on the internet, effective encoding recipes (i.e. bitrate-resolution pairs) are a main obstacle to deliver high-quality video streams. We developed a method to generate an encoding recipe that considers subjective visual quality with one just-noticeable difference (JND) distance. However, this method requires excessive computation time, which is not directly applicable for live streaming. In this paper, in order to provide a light-weight method for live streaming, we developed three acceleration techniques: resolution extrapolation, VMAF skipping and sampled objective measure calculation. These techniques are heuristic, but greatly contribute to reducing computational cost. Experimental results demonstrate that the proposed method achieves a significant reduction in computation time without significant effects on rate-JND characteristics. Yusuke Sakamoto, Shintaro Saika, Masaru Takeuchi, Tatsuya Nagashima, Zhengxue Cheng, Kenji Kanai, Jiro Katto, Kaijin Wei, Ju Zengwei |
ISM | 7 |
| 2018 | Deep Convolutional AutoEncoder-based Lossy Image CompressionabstractImage compression has been investigated as a fundamental research topic for many decades. Recently, deep learning has achieved great success in many computer vision tasks, and is gradually being used in image compression. In this paper, we present a lossy image compression architecture, which utilizes the advantages of convolutional autoencoder (CAE) to achieve a high coding efficiency. First, we design a novel CAE architecture to replace the conventional transforms and train this CAE using a rate-distortion loss function. Second, to generate a more energy-compact representation, we utilize the principal components analysis (PCA) to rotate the feature maps produced by the CAE, and then apply the quantization and entropy coder to generate the codes. Experimental results demonstrate that our method outperforms traditional image coding algorithms, by achieving a 13.7% BD-rate decrement on the Kodak database images compared to JPEG2000. Besides, our method maintains a moderate complexity similar to JPEG2000. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
PCS | 4 |
| 2018 | Perceptual Quality Driven Adaptive Video Coding Using JND EstimationabstractWe introduce a perceptual video quality driven video encoding solution for optimized adaptive streaming. By using multiple bitrate/resolution encoding like MPEG-DASH, video streaming services can deliver the best video stream to a client, under the conditions of the client's available bandwidth and viewing device capability. However, conventional fixed encoding recipes (i.e., resolution-bitrate pairs) suffer from many problems, such as improper resolution selection and stream redundancy. To avoid these problems, we propose a novel video coding method, which generates multiple representations with constant Just-Noticeable Difference (JND) interval. For this purpose, we developed a JND scale estimator using Support Vector Regression (SVR), and designed a pre-encoder which outputs an encoding recipe with constant JND interval in an adaptive manner to input video. Masaru Takeuchi, Shintaro Saika, Yusuke Sakamoto, Tatsuya Nagashima, Zhengxue Cheng, Kenji Kanai, Jiro Katto, Kaijin Wei, Ju Zengwei |
PCS | 7 |
| 2018 | Performance evaluations of software-defined acoustic MIMO-OFDM transmissionabstractIn recent years, the system using acoustic communication is increasing. However, because acoustic communication uses low frequency, transmission rate is lower than radio wave communication. In wireless communication, MIMO-OFDM is proposed for improvement quality and transmission rate. In this paper, we introduce a software-defined acoustic communication platform by using MATLAB and implement acoustic MIMO-OFDM transmission into the platform. Also, we evaluate BER characteristics in various experimental parameters in MATLAB simulation and real environment. Moreover, we evaluate image quality in actual acoustic image transmission by using the acoustic communication platform and we can successfully transmit the image via acoustic MIMO-OFDM. Airi Sakaushi, Mayuko Okano, Kenji Kanai, Jiro Katto |
WCNC | 4 |
| 2017 | A Pre-Saliency Map Based Blind Image Quality Assessment via Convolutional Neural NetworksabstractIn recent years, various approaches have been investigated towards blind image quality assessment (IQA) with high accuracy and low complexity. In this paper we develop a pre-saliency map based blind IQA method, which takes advantage of saliency information in prior of quality prediction for performance enhancement by two steps. 1) We split the image into patches and design a convolution neural network (CNN) to predict the patch-wise quality score. Then we explore the relation between image saliency information and CNN prediction error to present a statistical analysis. 2) Based on the analysis, we propose a patch quality aggregation algorithm by removing non-salient patches which are likely to bring large prediction error and assigning large weights for salient patches. Experimental results validate that our method can achieve high accuracy (0.978) with subjective quality scores, which outperforms existing IQA methods. Meanwhile, the proposed method can reduce 52.7% computational time than the IQA without pre-saliency map. Zhengxue Cheng, Masaru Takeuchi, Jiro Katto |
ISM | 3 |
| 2017 | QoS and QoE Evaluations of 2K and 4K DASH Contents DistributionsabstractThe increasing demand of mobile applications has brought large amount of mobile traffic. To meet users' requirements for high-quality video delivery, it is an urgent task to provide fair-quality video delivery for various users and situations. In this paper, we evaluate QoS and QoE characteristics and validate QoE unfriendliness in heterogeneous DASH contents distributions to provide QoE-fair video delivery,. In the evaluations, we employ multiple contents with different resolutions, frame rates, and rate-distortion characteristics. By using heterogeneous DASH contents, we evaluate the effect of playout buffer size on QoS and QoE performances. Evaluation results show that smaller playout buffer size contributes to QoE improvement under network congestion. In addition, we confirm that suppression of playback stall is a particularly important factor to achieve QoE fairness compared to other QoS metrics such as initial delay and representation. Tatsuya Nagashima, Kenji Kanai, Jiro Katto |
ISM | 3 |
| 2017 | A History-Based TCP Throughput Prediction Incorporating Communication Quality Features by Support Vector Regression for Mobile NetworkabstractThroughput prediction is one of good solutions to improve quality of mobile applications (e.g., YouTube or Netflix) for video streaming delivery services in mobile networks. This is because such applications require monitoring the network performances to control content quality, thus guarantee quality of service (QoS) and quality of experience (QoE). In this paper, we propose a history-based TCP throughput prediction method incorporating communication quality features using SVR (Support Vector Regression). By taking history of communication quality features such as historical throughput and Received Signal Strength Indication (RSSI) into consideration, the throughput prediction error can be decreased. We conduct experiments with the proposed method and compare the prediction accuracy with a variety of methods in different scenarios of various moving modes of users. Results show that the proposed model could predict throughput effectively in various scenarios and decrease throughput prediction errors by a maximum of 26.47% compared with other methods. Bo Wei 0001, Wataru Kawakami, Kenji Kanai, Jiro Katto |
ISM | 4 |
| 2016 | Enhancement of HCCA utilizing capture effect to support high QoS and DCF friendlinessabstractWith increase of mobile wireless LAN systems, frequency contamination by the overlapping basic service set (OBSS) becomes a critical issue. In this paper, we focus on HCF controlled channel access (HCCA) to alleviate the OBSS problem. HCCA considers single BSS (SBSS) environment only and suppresses traffic of co-existing DCF (distributed coordination function) based WiFi access points (APs). We propose two methods that utilize capture effects to improve the coexistence capability with DCF networks, that we call “DCF friendliness”. The first method adjusts transmission timing in HCCA WLAN by applying inter-AP coordination. The second method enables simultaneous communication with surrounding DCF WLANs by changing the frame interval of HCCA to “DIFS+1 SlotTime.” Simulations show that both the proposed method can achieve higher throughput and better DCF friendliness. Masanori Kanda, Jiro Katto, Tutomu Murase |
CCNC | 2 |
| 2016 | Quality Evaluations of 8K/60P UHDTV Retransmission for a Broadcasting and Communication Integrated PlatformabstractIn this study, we propose a broadcasting and communication integrated platform that can retransmit 8K broadcasting contents from a TV receiving set to mobile devices. We then evaluate the QoS characteristics and the qualities of 8K contents in this platform. Evaluation results conclude that 8K broadcasting contents can be retransmitted to mobile devices smoothly and their qualities are nearly as high as those of assumed 8K broadcasting service in Japan. Rintaro Harada, Shintaro Saika, Masaru Takeuchi, Kenji Kanai, Yasutaka Matsuo, Jiro Katto |
ISM | 6 |
| 2016 | Proactive Content Caching for Mobile Video Utilizing Transportation Systems and Evaluation Through Field ExperimentsabstractIn order to provide high-quality and highly reliable video delivery services for mobile users, especially train passengers, we propose a proactive content caching scheme that uses transportation systems. In our system, we place content servers with cache capability [e.g., content centric networking/named data networking (CCN/NDN)] in every train and station. Video segments encapsulated by MPEG-Dynamic Adaptive Streaming over HTTP (MPEG-DASH) are distributed and pre-cached by the station servers before the trains arrive at the stations. The trains receive content via high-speed wireless transport, such as wireless LANs or millimeter waves, when they stop at the stations. We developed prototype systems based on hypertext transfer protocol and CCN/NDN protocol, evaluate their performance through two field experiments that uses actual trains, and compare with traditional video streaming over cellular networks. Such evaluations indicate that our system can achieve high-quality video delivery without interruption for up to 50 users simultaneously. Kenji Kanai, Takeshi Muto, Jiro Katto, Shinya Yamamura, Tomoyuki Furutono, Takafumi Saito, Hirohide Mikami, Kaoru Kusachi, Toshitaka Tsuda, Wataru Kameyama, Takuro Sato |
IEEE J. Sel. Areas Commun. | 3 |
| 2015 | A highly-reliable buffer strategy based on long-term throughput prediction for mobile video streamingabstractProviding robust video streaming along with efficient wireless resource usage is necessary for mobile users, especially on subway, and mobile carriers. To achieve this, we propose a highly-reliable buffer strategy based on long-term throughput prediction. Our approach has two elements which are called “long-term throughput prediction” and “guaranteed playout buffer filling mechanism.” To avoid any video freeze due to network quality degradation, our approach calculates the optimal amount of playout buffer and schedules video download timing in a theoretical manner. We evaluate its performance via experiments in real environment. Evaluations conclude that our approach can provide highly-reliable video streaming and also achieve to reduce the average playout buffer size on the client. Kenji Kanai, Hidenori Konishi, Yuya Ishizu, Jiro Katto |
CCNC | 4 |
| 2015 | An Adaptive H.265/HEVC Encoding Control for 8K UHDTV Movies Based on Motion Complexity EstimationabstractIn this paper, we propose a method to control H.265/HEVC encoding for 8K UHDTV moving pictures by detecting amount or complexity of object motions. In 8K video, which has very high spatial resolution, motion has a big influence on encoding efficiency and processing time. The proposed method estimates motion features by external process which uses local feature points matching between two frames, selects an optimal prediction mode and determines search ranges of motion vectors. Experiments show we can detect motion complexity of 8K movies by using local feature matching between frames and we can select optimal configurations of encoding. By our method, we achieved highly efficient and low computation encoding. Shota Orihashi, Rintaro Harada, Yasutaka Matsuo, Jiro Katto |
ISM | 4 |
| 2015 | Live Version Identification with Audio Scene Detection
Kazumasa Ishikura, Aiko Uemura, Jiro Katto |
MMM (1) | 3 |
| 2015 | Development of software-defined acoustic communication platform and its evaluationsabstractIn recent years, researches of underwater sensor networks have continued to investigate environment and resources of the sea. Acoustic waves are used instead of radio waves for wireless communication in underwater. However, dedicated hardware is very expensive, experiments on the sea are very time-consuming, and huge water spaces are necessary to study the underwater acoustic communication. In this paper, we present a cheap and tractable software-defined acoustic communication platform running on PCs using MATLAB, and evaluate its characteristics in a variety of communication methods by changing modulation schemes, error correction codes, transmission power and frequency by using commercial speaker and microphone devices. Our current implementation achieves data rate of up to 4.5 Kbps. Ryo Kato, Jiro Katto |
WCNC | 2 |
| 2015 | Implementation evaluation of proactive content caching using DASH-NDN-JSabstractProactive content caching scheme utilizing transportation systems, especially on trains, was proposed in order to provide a robust content delivery with efficient wireless resource usage. This system requires content servers with NDN capability to be placed on every station and trains. The mechanism is to pre-cache the contents that users request to the station server before the train arrives, and the train server caches the content during its stoppage time. With this mechanism, users are able to have a continuous playback of videos while riding on trains. In this paper, we have proposed a browser-based implementation, called DASH-NDN-JS, for this proactive content caching scheme. We evaluate this scheme, and experiment with multiple users to see how it will affect the video quality each user will achieve and the bandwidth consumption between the connections. Our evaluations conclude that the increase of users lowers the video quality, but avoids congestion depending on what video content each user will want to request. Takeshi Muto, Kenji Kanai, Jiro Katto |
WCNC | 3 |
| 2014 | Performance analysis and validation of high QoS route navigation for mobile usersabstractImproving Quality of Service (QoS) in wireless networks is important and necessary for mobile users. We have previously proposed Comfort Route (CR) Navigation, which navigates users to their destinations using high QoS communication areas, such as Wi-Fi APs, rather than the geographical Shortest Route (SR). In this paper, we employ an analytical model to estimate the CR gain in a theoretical manner which assumes that available cellular and Wi-Fi throughputs are uniform within their coverage. The CR gain is computed by using basic parameters, including wireless network bandwidth and transmission time. To validate our model, we compare simulation results and real observation. These results conclude that the CR gain could estimate by using our analytical model. Kenji Kanai, Jiro Katto, Tutomu Murase |
APNOMS | 2 |
| 2014 | Proactive content caching utilizing transportation systems and its evaluation by field experimentabstractProviding robust content delivery along with efficient wireless resource usage is important for next generation wireless networks. To achieve this, we propose a proactive content caching scheme utilizing transportation systems, especially trains. In our system, we place content servers with CCN capability to every train and station. Segments of video contents are pre-cached by the station servers before trains arrive at stations. Trains receive the contents via high-speed wireless transport while they stop at the stations. We develop a prototype system based on IP and CCN Hybrid protocols. We evaluate its performance by field experiment and compare with traditional CDN scenarios using cellular networks. Evaluations conclude that our system can achieve high-speed and high-reliable video delivery without freezing. Kenji Kanai, Takeshi Muto, Hiroto Kisara, Jiro Katto, Toshitaka Tsuda, Wataru Kameyama, Takuro Sato |
GLOBECOM | 4 |
| 2014 | Effects of Audio Compression on Chord Recognition
Aiko Uemura, Kazumasa Ishikura, Jiro Katto |
MMM (2) | 3 |
| 2013 | A study on gait recognition using LPC cepstrum for mobile terminalabstractThe use of mobile terminals has been expanding dramatically in recent years as they evolve from a means of dispatching and gathering information to a highly functional tool that supports personal lifestyles and behavior. A mobile terminal is likely to store various kinds of personal information such as a calendar and contact information as well as key data to carry out online transactions. Losing one's mobile terminal therefore creates the possibility that one's personal information may fall into the wrong hands and be used for malicious purposes. We therefore propose a method of personal authentication using sensor data in a mobile terminal. First, we applied the LPC cepstrum to this authentication and checked for validity. We also evaluated the effectiveness of gait authentication using several frames. Masatsugu Ichino, Hiroki Kasahara, Hideki Yoshii, Kazuhiro Tsurumaru, Naohisa Komatsu, Jiro Katto |
ICIS | 6 |
| 2013 | An adaptive TCP congestion control having RTT-fairness and inter-protocol friendlinessabstractThis paper presents an RTT-fair TCP congestion control using ACK interval measurement and extends the approach to have inter-protocol friendliness, especially with CUBIC TCP. In the previous RTT-fair TCP congestion control, including ours, estimation of RTTs of a competing flow had been a problem. We try to solve this problem by measuring ACK arrival intervals, which are observable parameters by an end host. This approach enables estimation of congestion window behaviors, in addition to RTTs, of a competing flow. We then extend our congestion control to have friendliness to CUBIC-TCP, in addition to classical TCP-Reno, in an adaptive manner. Extensive experiment results are shown for simulations and implementations and effectiveness of our approach is confirmed. Yohei Nemoto, Kazumine Ogura, Jiro Katto |
CCNC | 3 |
| 2013 | BRAEVE: Stable and adaptive BSM rate control over IEEE802.11p vehicular networksabstractIn vehicle-to-vehicle communication, a message named BSM (Basic Safety Message) has a major role to inform a driver about surrounding condition. A vehicle periodically sends the BSM which includes the information of itself. Traffic load issues caused by BSMs would often be raised on heavily congested roads. Therefore, BSM congestion controls to avoid traffic congestion are challenging. This paper proposes a new BSM congestion control named BRAEVE. BRAEVE adapts its BSM generation rate according to the number of neighbor vehicles in communication range and then controls overall network traffic load. Our simulation evaluation shows that BRAEVE provides more uniform recognition of surrounding vehicles than an existing method, which brings much better safety assurance to a driver. Kazumine Ogura, Jiro Katto, Mineo Takai |
CCNC | 2 |
| 2013 | TCP differentiation using version identification and EDCA for low-delay multimedia streamingabstractThis paper presents a TCP differentiation based on TCP version identification and IEEE 802.11e EDCA for low delay multimedia streaming over wireless LAN. It has been known that delay-based TCP can achieve low delay transport as long as it does not compete with loss-based TCP. However, when competition happens, it seriously decreases its rate by itself due to RTT increase. In order to alleviate this problem, we consider combination of TCP version identification and prioritized transport of delay-based TCP by EDCA at an access point. We evaluate this approach by simulations and implementations, and confirm its effectiveness. Kazuhide Sonoda, Kazumine Ogura, Jiro Katto |
CCNC | 3 |
| 2013 | Image Super-resolution Using Registration of Wavelet Multi-scale Components with Affine TransformationabstractWe propose a novel image super-resolution method from digital cinema to 8K ultra high-definition television using registration of wavelet multi-scale components with affine transformation. The proposed method features that an original image is divided into signal and noise components by the wavelet soft-shrinkage with detection of white noise level. The signal component enhances resolution by registration between a signal component and its wavelet multi-scale components with affine transformation and parameters optimization. The affine transformation enhances super-resolution image quality because it increases registration candidates. The noise component enhances resolution with power control considering cinema noise representation. Super-resolution image outputs by synthesis of super-resolved signal and noise components. Experiments show that the proposed method has objectively better PSNR measurement and subjectively better appearance in comparison with conventional super-resolution methods. Yasutaka Matsuo, Ryoki Takada, Shinya Iwasaki, Jiro Katto |
ISM | 4 |
| 2013 | RoCNet: Spatial mobile data offload with user-behavior prediction through delay tolerant networksabstractWe present a robust cellular network (RoCNet) that combines a cellular and an opportunistic networks for spatial uplink mobile data offloading, which focuses on the spatial difference of the traffic load among areas (e.g., business district and residential area in the daytime). RoCNet realizes the spatial data offload by leveraging the store-carry-forward routing mechanism. In the area where traffic load is high, delay-tolerant data originated from a mobile terminal is directly forwarded to a nearby terminal using Bluetooth or wireless LAN instead of being transmitted to a congested cellular base station. When the data is carried by the nearby terminal to other area where the traffic load is low, the data is forwarded to a cellular base station. To enhance the offload effect, it is necessary for data to be forwarded to a terminal that moves to a low traffic load area. In this paper, we use the particle filter to predict user behavior. Before forwarding data between mobile terminals, the terminals exchange prediction results and decide whether or not the data should be forwarded. We conducted a computer simulation whose result shows RoCNet can spatially offload uplink traffic in a traffic concentration area to non-congested areas. As a result, RoCNet can suppress peak traffic by about 20 percent in a traffic-congested base station by distributing traffic to vicinity base stations. Haruki Izumikawa, Jiro Katto |
WCNC | 2 |
| 2012 | A bit-depth scalable video coding approach considering spatial gradation restorationabstractBit-depth scalable coding method is an approach that generates multiplexed bit-streams that can be decoded to two video sequences, for Standard Dynamic Range (SDR) environment and for High Dynamic Range (HDR) environment. This paper presents a bit-depth scalable coding method that is considered as gradation restoration in interlayer prediction process. Our proposed inter-layer prediction uses a histogram interpolation method that enables to generate more mid-gray brightness levels than traditional inverse tone mapping methods with one-to-one correspondence. Masaru Takeuchi, Yasutaka Matsuo, Yuta Yamamura, Jiro Katto, Kazuhisa Iguchi |
ICASSP | 4 |
| 2012 | Chord recognition using Doubly Nested Circle of FifthsabstractThis paper presents a chord recognition method from music signals using chroma vectors and musical knowledge known as “Doubly Nested Circle of Fifths (DNCOF)”. DNCOF represents the relationships of major and minor chords where the neighboring two triads are similar. We obtain a novel feature from chroma vectors by mapping them onto two-dimensional DNCOF coordinate, which we call “DNCOF vectors”. We expect that the DNCOF vectors can contribute to correcting false recognition obtained by the chroma vectors when their mapped positions are apart from one another in the DNCOF coordinate. In this research, we evaluated our proposal using the Beatles' datasets and showed its effectiveness. Aiko Uemura, Jiro Katto |
ICASSP | 2 |
| 2012 | Improving the performance of SIFT using Bilateral Filter and its application to generic object recognitionabstractFeature extraction of images can be applied to image matching, image searching, object recognition, image tracking etc. One of the effective methods to extract features of images is Scale-Invariant Feature Transform (SIFT) [1], In this paper, we indicate problems of SIFT and propose a method to improve its performance by applying Bilateral Filter [2]. In addition, we implement its acceleration by GPGPU (general purpose GPU), apply this method to generic object recognition and perform a comparison experiment. We compare the proposed method with the original method using SIFT and confirm improvement of the identification rate by the proposed method. Tomoaki Yamazaki, Tetsuya Fujikawa, Jiro Katto |
ICASSP | 3 |
| 2012 | Music Part Segmentation in Music TV Programs Based on Chroma Vector AnalysisabstractThis paper presents a music part detection method incorporating chroma vector analysis for use with music TV programs. Results show that envelopes of chroma components of music signals tend to have horizontal (i.e. temporal) correlation in time-frequency representation because music signals have a periodic chord sequences. Based on this fact, we analyze time series of chroma components and attempt to segment music parts in music TV programs from other parts. Experimental results show an F-measure of 0.78, which is better than that obtained using the previous method. Aiko Uemura, Jiro Katto, Kyota Higa, Masumi Ishikawa, Toshiyuki Nomura |
ISM | 2 |
| 2012 | Wavelet domain image super-resolution from digital cinema to ultrahigh definition television by dividing noise componentabstractWe propose a novel wavelet domain image super-resolution method from digital cinema to ultrahigh definition television considering cinema noise component. The proposed method features that spatial resolution of an original image is expanded by synthesis of super-resolved signal and noise components respectively after dividing an original image into signal and noise components. Dividing noise component uses spatio-temporal wavelet decomposition based on frequency spectrum analysis of cinema noise. And super-resolution parameters are optimized by comparing size-reduced super-resolution images with an original image. Experimental results showed that a super-resolution image using the proposed method has a subjectively better appearance and an objectively better peak signal-to-noise ratio measurement than conventional methods. Yasutaka Matsuo, Shinya Iwasaki, Yuta Yamamura, Jiro Katto |
VCIP | 4 |
| 2010 | ToMo: A Two-Layer Mesh/Tree Structure for Live Streaming in P2P Overlay NetworkabstractIn this paper, we introduce a hybrid approach for overlay construction and data delivery in an application-layer multicast. We combine the strong points of a tree-based structure and a mesh-based data delivery to form ToMo, a two-layer hybrid overlay. We try to reduce the number of replicated packets at a source, and reduce an effect when slow connection peers are located near the source. The overlay is constructed in the fashion of a mesh layer over a tree layer. This structure allocates the source to multicast each piece of the packet to a specific group of child peers only. Different from other approaches, we employ only push-based data delivery in order to minimize the latency. The redundancy is avoided by defining a set of well-organized mesh connections. Furthermore, in our approach, the isolated peers affected by parent departure are not facing data loss during the rejoin process since they still receive data from their neighbors via mesh connections. Simulations through ns2 demonstrate the efficiency of this solution. Suphakit Awiphan, Zhou Su 0001, Jiro Katto |
CCNC | 3 |
| 2010 | Hybrid Application Layer Multicast with Hierarchically Distributed NodesabstractThe hybrid application layer multicast (ALM) has been shown its efficiency by leveraging the conventionally main structures of application layer multicast, tree-based and mesh-based. However, how to select the proper node to construct the overlay and how to establish the connection between any two nodes are still unsolved. Therefore, this paper is to design a novel construction algorithm for the hybrid ALM to resolve the above two issues. Firstly, by carrying out the analysis of nodes' characteristics, all nodes are divided into groups and a node priority is proposed to select the super node within each group. Secondly, by using the selected super node, all nodes are hierarchically controlled and different kinds of connections are carried out in the ALM, where the connection between super node and other nodes is set to be a tree to enhance the efficient utilization of network resource while the connection between other normal nodes is decided to be a mesh to reduce the overhead. Simulation results show that the proposal outperforms other conventional methods. Zhou Su 0001, Suphakit Awiphan, Kazumine Ogura, Jiro Katto, Yasuhiko Yasuda |
CCNC | 4 |
| 2010 | A Novel Algorithm to Control Contents Selectively for Vehicular Communication NetworksabstractWith the development of recent vehicular communication technologies, distributing multimedia contents in the vehicular communication networks (VCNs) has become more and more popular, to provide conveniences and entertainment services during the time of driving. However, as multimedia contents are changed and updated dynamically, how to keep the consistency between the original and these replicas in VCNs is very important. Therefore, this paper designs a novel algorithm to control the consistency for the VCNs. In our proposal, after the analyses of the status of road-side units, on-board units and local geographical information, we divide all replicas into two groups, where one is necessary for update and the other are not. Then, we compare the cost to update replicas by using wireless and wired connection, and propose a method to make selection between them. The performance of our proposal is tested by simulation experiments. And the results show that our method can reduce the delay successfully. Zhou Su 0001, Pinyi Ren, Rongtao Xu, Jiro Katto, Yasuhiko Yasuda |
VTC Fall | 4 |
| 2009 | Efficient Construction in ALM with Assignment of Layered Degree and ALM-Bi-CastabstractThis paper designs a tree construction algorithm by distributing the layered steaming over the ALM in order to improve both the throughput and user delay. Firstly, by carrying out theory analysis, we define an out/in degree and the corresponding constraints to manage the layered streaming and nodes to improve the throughput. Besides, a novel method, called ALM-Bi-cast, is also proposed and analyzed to reduce user delay during the data-transmission. Secondly, by using the defined degrees and the ALM-Bi-cast, we present a tree construction algorithm and test it by simulations with NS-2. Simulations show that the proposal can get better results than other conventional methods. Finally, we carry out an implementation of our proposal, by distributing the video encoded by H.263+ over the ALM nodes placed in Tokyo City and Kyusyu Prefecture. Implementation further improves the out-performance of our proposal. Zhou Su 0001, Masato Oguro, Yohei Okada, Jiro Katto, Sakae Okubo |
CCNC | 4 |
| 2009 | Robust Mesh-Based Data Delivery over Multiple Tree-Shaped Routes in P2P Overlay NetworkabstractIn this paper, we introduce a new mesh-based approach for data delivery which is organized over multiple tree-shaped core routes. Given that both tree and mesh approaches have their own strong points, we simply combine them together. We evaluate the proposal through ns-2 simulator. The simulation results demonstrate that our approach can provide higher average received quality and has acceptable data delivery delay when compared with a single tree method. We also show that, over a static overlay, the push-based data delivery on mesh can provide the received quality close to pull-based data delivery method with less latency. As well, it has lower control overhead than the pull-based method when the peer number is large. Suphakit Awiphan, Jiro Katto |
ISORC | 3 |
| 2009 | Feature Analysis and Normalization Approach for Robust Content-Based Music Retrieval to Encoded Audio with Different Bit Rates
Shuhei Hamawaki, Shintaro Funasawa, Jiro Katto, Hiromi Ishizaki, Keiichiro Hoashi, Yasuhiro Takishima |
MMM | 3 |
| 2009 | Key Estimation Using Circle of Fifths
Takahito Inoshita, Jiro Katto |
MMM | 2 |
| 2009 | Priority based selection to improve contents consistency for mobile overlay networkabstractWith the growing use of dynamic content by mobile content distribution systems, how to manage dynamically changing files has become an important issue, since the cached replicas on different mobile sites must be updated if the originals have been changed. Therefore, this paper proposes a priority based selection method to enhance the efficient utilization of network resource and support the client mobility for mobile contents delivery network (M-CDN). On one hand, a consistency priority is calculated by analyzing the characteristics of mobile surrogates. If a given content which has been changed on its original node, only the replicas with the high consistency priority instead of all replicas are updated. On the other hand, an update priority is also proposed. If a replica has been selected for update, the latest version will be sent from the site decided by the update consistency. Simulation results show that the proposed new approach outperforms other conventional methods. Zhou Su 0001, Jiro Katto, Yasuhiko Yasuda |
WCNC | 2 |
| 2008 | Simple Model Analysis and Performance Tuning of Hybrid TCP Congestion ControlabstractThis paper presents simple analytical models of hybrid TCP congestion controls, which switch loss-based mode and delay-based mode adaptively, and tries their performance tuning. We firstly present ideal behavior models of three kinds of TCP congestion controls (loss-based, delay-based and hybrid). We then give abstracted models of the actual hybrid TCP s and consider their performance tuning. Finally, experiments validate analytical expectations and effectiveness of the hybrid TCP. Jiro Katto, Kazumine Ogura, Yuki Akae, Tomoki Fujikawa, Kazumi Kaneko |
GLOBECOM | 1 |
| 2008 | Denoising intra-coded moving pictures using motion estimation and pixel shiftabstractThis paper presents a denoising method of intra-coded pictures using motion estimation and pixel shift. Firstly, we show that pixel-aligned mixture of distorted images which are spatially shifted and differently encoded brings reduction of quantization errors. We show that this effect can be formulated as a special case of Wiener-Hopf equation and independence of quantization errors affects the performance. We then consider its application to denoising of intra-coded pictures by using motion estimation and pixel shift. Experiments using actual image sequences verify that motion estimation is effective in moving regions, pixel shift is effective in static regions and favorable PSNR gains are achieved. Jiro Katto, Junya Suzuki, Shusei Itagaki, Shinichi Sakaida, Kazuhisa Iguchi |
ICASSP | 1 |
| 2007 | Scalable Maintenance for Strong Web Consistency in Dynamic Content Delivery OverlaysabstractContent Delivery Overlays improves end-user performance by replicating Web contents on a group of geographically distributed sites interconnected over the Internet. However, with the development whereby overlay systems can manage dynamically changing files, an important issue to be resolved is consistency management, which means the cached replicas on different sites must be updated if the originals change. In this paper, based on the analytical formulation of object freshness, Web access distribution and network topology, we derive a novel algorithm as follows: (1) For a given content which has been changed on its original server, only a limited number of its replicas instead of all replicas are updated. (2) After a replica has been selected for update, the latest version will be sent from an algorithm-decided site instead of from its original server. Simulation results verify that the proposed algorithm provides much better consistency management than conventional methods with the reduced the old hit ratio and network traffic. Zhou Su 0001, Jiro Katto, Yasuhiko Yasuda |
ICC | 2 |
| 2007 | Support Strong Consistency for Mobile Dynamic Contents Delivery NetworkabstractWith the development whereby mobile content distribution systems can manage dynamically changing files, an important issue to be resolved is consistency management, which means the cached replicas on different mobile sites must be updated if the originals change. This paper is to design an integrated consistency-control algorithm for mobile contents delivery network (M-CDN) to enhance the efficient utilization of network resource and support the client mobility. Firstly, by carrying out an analysis of mobile surrogates' characteristics, for a given content which has been changed on its original node, only a limited number of its replicas instead of all replicas are updated. Secondly, if a replica has been selected for update, the latest version will be sent from an algorithm-decided site instead of from its original server. Simulation results show that the proposal outperforms other conventional methods. Zhou Su 0001, Jiro Katto, Yasuhiko Yasuda |
ISM | 2 |
| 2006 | Feature Space Modification for Content-Based Music Retrieval Based on User PreferencesabstractThis paper proposes a feature space modification method for feature extraction of music, which is effective for the development of a content-based music information retrieval (MIR) system based on user preferences. The proposed method conducts clustering of all songs in the music collection, and utilizes the resulting cluster IDs as training data for feature space modification, and is capable to automatically generate a feature space which is suitable to the content of any music collection. Experiment results prove that the proposed method improves accuracy of user preference based MIR Keiichiro Hoashi, Kazunori Matsumoto, Fumiaki Sugaya, Hiromi Ishizaki, Jiro Katto |
ICASSP (5) | 5 |
| 2005 | Scalable Consistency Management in Dynamic Content Distribution OverlaysabstractContent distribution overlays improves end-user performance by replicating Web contents on a group of geographically distributed sites interconnected over the Internet. However, with the development whereby overlay systems can manage dynamically (Ganguly et al., 2005) changing files, an important issue to be resolved is consistency management, which means the cached replicas on different sites must be updated if the originals change. In this paper, based on the analytical formulation of object freshness time, Web access distribution and network topology, we derive a novel algorithm as follows: (1) for a given content which has been changed at its original server, only a limited number of its replicas instead of all replicas are updated. (2) After a replica has been selected for update, the latest version will be sent from an algorithm-decided site instead of from its original server. Simulation results verify that the proposed algorithm provides much better consistency management than conventional methods with the reduced update overhead and network traffic Zhou Su 0001, Jiro Katto, Yasuhiko Yasuhiko |
CLUSTER | 2 |
| 2005 | An integrated Retrieval and Pre-fetching algorithms for Segmented Streaming in Mobile Peer-to-Peer NetworksabstractIn contrast to conventional P2P systems in wired networks that consist of static peers, mobile P2P are subjected to the limitations of battery power, wireless bandwidth, and the dynamically changed network topology. Challenges arise in how to improve the source discovery and data replication. In this paper, we talk about an integrated searching and prefetching algorithm for the segmented streaming in mobile peer-to-peer (P2P) Networks. Firstly, each stream is divided into several segments and each segment is assigned a priority based on theory analyses. Then, for a given segment, the different number of queries is sent to search it and the length of the query for this segment is also dynamically decided by the segment-priority to avoid the unnecessary overhead. Next, along the path where a stream is sent from the requester node, parts of the nodes on this path are selected to pre-fetch the requested segment to reduce the user delay for the next possible request. Finally, Simulation results show that better performance than the conventional methods can be achieved Zhou Su 0001, Jiro Katto, Yasuhiko Yasuda |
CLUSTER | 2 |
| 2004 | An efficient TCP with explicit handover notification for mobile networksabstractTCP is a popular Internet protocol for reliable end-to-end data delivery, but it cannot be directly applied to wireless networks in which packet loss may be induced by higher BER or handover than congestion. TCP assumes that such packet loss is caused by network congestion and initiates congestion control procedures. In this paper, we present a novel protocol using explicit handover notification to improve TCP performance over wireless links. Additionally, we execute computer simulations using the network simulator and compare it with the other various protocols. Haruki Izumikawa, Ichiro Yamaguchi, Jiro Katto |
WCNC | 3 |
| 1999 | Structure Recovery with Multiple Cameras from Scaled Orthographic and Perspective ViewsabstractThis paper presents a novel framework for Euclidean structure recovery utilizing a scaled orthographic view and perspective views simultaneously. A scaled orthographic view is introduced in order to automatically obtain camera parameters such as camera positions, orientation, and focal length. Scaled orthographic properties enable all camera parameters to be calculated implicitly and perspective properties enable a Euclidean structure to be recovered. The method can recover a Euclidean structure with at least seven point correspondences across a scaled orthographic view and perspective views. Experimental results for both computed and natural images verify that the method recovers structure with sufficient accuracy to demonstrate potential utility. The proposed method can be applied to an interface for 3D modeling, recognition and tracking. Atsushi Marugame, Jiro Katto, Mutsumi Ohta |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | System architecture for synthetic/natural hybrid coding and some experimentsabstractThis paper presents a system architecture for synthetic/natural hybrid coding toward future visual services. Scene-description capability, terminal architecture, and network architecture are discussed by taking account of the standardization activities: MPEG, VRML, ITU-T, and IETF. A consistent strategy to integrate scene-description capability and streaming technologies is demonstrated. Experimental results are also shown, in which synthetic/natural integration is successfully carried out. Jiro Katto, Mutsumi Ohta |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Runlength-based Wavelet Coding with Adaptive Scanning for Low Bit Rate EnvironmentabstractA new wavelet image coding method is presented to exploit both intra/inter-band correlations. While the traditional zerotree approach can exploit correlation between subbands, a wavelet decomposition has properties more suitable for the runlength coding method, which can exploit spatial correlation. The proposed method classifies wavelet coefficients into groups according to their parents' magnitude. Each group is chosen in order of the magnitude and scanned to keep spatial correlation. It outperforms the zerotree-based coding approach especially at low bit rate. Takahiro Kimoto, Jiro Katto, Mutsumi Ohta |
ICIP (2) | 2 |
| 1996 | Novel algorithms for object extraction using multiple camera inputsabstractThis paper presents novel algorithms exploiting multiple camera inputs and segmentation techniques, which can be applied to image fusion, disparity detection and object extraction. Differently focused images, stereo pairs and both of them are used for fusion, disparity detection and object extraction, respectively. Firstly, image fusion is done by segmentation of each image and determination of focused regions per segment. An efficient decision criterion is developed taking the method of auto-focus into consideration. Secondly, disparity detection is executed by recursively applying segmentation and disparity detection per segment. A new clustering criterion is proposed in order to achieve good segmentation and high compression ratio of disparity maps simultaneously. Finally, object extraction is carried out by utilizing both the fusion result and the disparity map. Experiments are carried out, and they show the effectiveness of the proposed algorithms. Jiro Katto, Mutsumi Ohta |
ICIP (2) | 1 |
| 1996 | Structure recovery from scaled orthographic and perspective viewsabstractThis paper presents a novel framework for structure recovery utilizing a scaled orthographic view and perspective views simultaneously. Perspective views lead to precise recovery based on the triangulation principle; however, many parameters such as camera positions, camera poses and focal length, must be measured beforehand. Thus, a scaled orthographic view is introduced as a subsidiary system to achieve them automatically. Camera parameters are calculated implicitly owing to the scaled orthographic properties, and then the structure recovery is done by the orthogonality of camera coordinate systems and the triangulation principle. The experimental results provide sufficient accuracy of structure recovery. Atsushi Marugame, Jiro Katto, Mutsumi Ohta |
ICIP (2) | 2 |
| 1996 | Improved scanning methods for wavelet coefficients of video signals
Fabrice Rossetti, Jiro Katto, Mutsumi Ohta |
Signal Process. Image Commun. | 2 |
| 1995 | An analytical framework for overlapped motion compensationabstractMotion compensated interframe prediction is a key technology used to accomplish efficient video codecs. This paper presents a novel framework which proves the superiority of overlapped motion compensation. The window design problem is revised by introducing a statistical model of the motion estimation process. The result clarifies the relationship between the optimum window and image characteristics in an explicit formula and quantifies the prediction error reduction achieved by overlapped motion compensation. Experimental results using real image sequences support the proposed theory and demonstrate its superiority. Overlapping in warping prediction is also considered and its effectiveness is shown. Jiro Katto, Mutsumi Ohta |
ICASSP | 1 |
| 1995 | Mathematical analysis of MPEG compression capability and its application to rate controlabstractThis paper presents mathematical frameworks of temporal predictive processing in the MPEG video compression standard. The coding gain is derived based on traditional prediction theories. The optimum ordering of three different picture types (I,P,B-pictures) is clarified according to the image source characteristics. A novel framework of the target bit assignment is presented with some experimental backgrounds. The solution consists of simple formulae, but provides drastic SNR gains to the conventional TM5 algorithm. Jiro Katto, Mutsumi Ohta |
ICIP | 1 |
| 1994 | A wavelet codec with overlapped motion compensation for very low bit-rate environmentabstractThe paper describes theories on overlapped motion compensation and their applications to a wavelet codec aimed for the very low bit-rate environment below 64 kb/s. The theories are concerned with the evaluation of prediction efficiency improved by overlapped motion compensation and also its smoothing effect on the discontinuities at block boundaries encountered when motion vectors do not coincide among neighboring blocks. They contribute to determine the optimum window shape for overlapped motion compensation in a developed wavelet codec, which suffers from reduction of coding efficiency when there are such discontinuities in a signal to be transformed. Regarding the wavelet codec, a new scanning method of transform coefficients and alternate use of normal and spatially-reversed basis functions for the cascaded wavelet transform are introduced. Finally, the implementation of the codec on a video image signal processor (VISP) is discussed.> Jiro Katto, Jun-ichi Ohki, Satoshi Nogaki, Mutsumi Ohta |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1991 | Performance evaluation of subband coding and optimization of its filter coefficients
Jiro Katto, Yasuhiko Yasuda |
J. Vis. Commun. Image Represent. | 1 |
| 1991 | Variable bit-rate coding based on human visual system
Jiro Katto, Katsumasa Onda, Yasuhiko Yasuda |
Signal Process. Image Commun. | 1 |