Hasan F. Ates

dblp:70/391 · also Hasan Fehmi Ates · DBLP profile ↗
← Back
31ranked-venue papers
17as first author
7since 2021 · last 2025
0000-0002-6842-1528ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 15 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 DA-Mamba: Domain Adaptive Hybrid Mamba-Transformer Based One-Stage Object Detection
abstract
Recent 2D CNN-based domain adaptation approaches struggle with long-range dependencies due to limited receptive fields, making it difficult to adapt to target domains with significant spatial distribution changes. While transformer-based domain adaptation methods better capture distant relationships through self-attention mechanisms that facilitate more effective cross-domain feature alignment, their quadratic computational complexity makes practical deployment challenging for object detection tasks across diverse domains. Inspired by the global modeling and linear computation complexity of the Mamba architecture, we present the first domain adaptive Mamba-based one-stage object detection model, termed DA-Mamba. Specifically, we combine Mamba’s efficient state-space modeling with attention mechanisms to address domain-specific spatial and channel-wise variations. Our design leverages domain adaptive spatial and channel-wise scanning within the Mamba block to extract highly transferable representations for efficient sequential processing, while cross-attention modules generate long-range, mixed-domain spatial features to enable robust soft alignment across domains. Besides, motivated by the observation that hybrid architectures introduce feature noise in domain adaptation tasks, we propose an entropy-based knowledge distillation framework with margin ReLU, which adaptively refines multi-level representations by suppressing irrelevant activations and aligning uncertainty across source and target domains. Finally, to prevent overfitting caused by the mixed-up features generated through cross-attention mechanisms, we propose entropy-driven gating attention with random perturbations that simultaneously refine target features and enhance model generalization. Extensive experiments demonstrate that DA-Mamba consistently outperforms existing methods across a range of widely recognized domain adaptation benchmarks. Our code is available at https://github.com/enesdoruk/DA-Mamba.
Abdullah Enes Doruk, Hasan F. Ates
ECAI2
2025 Hstr-net: reference based video super-resolution with dual cameras
abstract
Abstract High-spatio-temporal resolution (HSTR) video recording plays a crucial role in enhancing various imagery tasks that require fine-detailed information. State-of-the-art cameras provide this required high frame-rate and high spatial resolution together, albeit at a high cost. To alleviate this issue, this paper proposes a dual camera system for the generation of HSTR video using reference-based super-resolution (RefSR). One camera captures high spatial resolution low frame rate (HSLF) video while the other captures low spatial resolution high frame rate (LSHF) video simultaneously for the same scene. A novel deep learning architecture is proposed to fuse HSLF and LSHF video feeds and synthesize HSTR video frames. The proposed model combines optical flow estimation and (channel-wise and spatial) attention mechanisms to capture the fine motion and complex dependencies between frames of the two video feeds. Simulations show that the proposed model provides significant improvement over existing reference-based SR techniques in terms of PSNR and SSIM metrics. The method also exhibits sufficient frames per second (FPS) for aerial monitoring when deployed on a power-constrained drone equipped with dual cameras. The source code is publicly available at https://github.com/umutsuluhan/HSTRNet .
Hasan Umut Suluhan, Abdullah Enes Doruk, Hasan F. Ates, Bahadir K. Gunturk
Mach. Vis. Appl.3
2023 Deep learning-based blind image super-resolution with iterative kernel reconstruction and noise estimation
Hasan F. Ates, Suleyman Yildirim, Bahadir K. Gunturk
Comput. Vis. Image Underst.1
2022 Dual Camera Based High Spatio-Temporal Resolution Video Generation For Wide Area Surveillance
abstract
Wide area surveillance (WAS) requires high spatiotemporal resolution (HSTR) video for better precision. As an alternative to expensive WAS systems, low-cost hybrid imaging systems can be used. This paper presents the usage of multiple video feeds for the generation of HSTR video as an extension of reference based super resolution (RefSR). One feed captures video at high spatial resolution with low frame rate (HSLF) while the other captures low spatial resolution and high frame rate (LSHF) video simultaneously for the same scene. The main purpose is to create an HSTR video from the fusion of HSLF and LSHF videos. In this paper we propose an end-to-end trainable deep network that performs optical flow (OF) estimation and frame reconstruction by combining inputs from both video feeds. The proposed architecture provides significant improvement over existing video frame interpolation and RefSR techniques in terms of PSNR and SSIM metrics and can be deployed on drones with dual cameras.
Hasan Umut Suluhan, Hasan F. Ates, Bahadir K. Gunturk
AVSS2
2022 Predicting Path Loss Distributions of a Wireless Communication System for Multiple Base Station Altitudes from Satellite Images
abstract
It is expected that unmanned aerial vehicles (UAVs) will play a vital role in future communication systems. Optimum positioning of UAVs, serving as base stations, can be done through extensive field measurements or ray tracing simulations when the 3D model of the region of interest is available. In this paper, we present an alternative approach to optimize UAV base station altitude for a region. The approach is based on deep learning; specifically, a 2D satellite image of the target region is input to a deep neural network to predict path loss distributions for different UAV altitudes. The neural network is designed and trained to produce multiple path loss distributions in a single inference; thus, it is not necessary to train a separate network for each altitude.
Ibrahim Shoer, Bahadir K. Gunturk, Hasan F. Ates, Tuncer Baykas
ICIP3
2022 Iterative Kernel Reconstruction for Deep Learning-Based Blind Image Super-Resolution
abstract
Deep learning based methods have received a great deal of interest in recent years to solve the single image super-resolution (SISR) problem and their performance is proven to be superior when compared to classical SR techniques. Yet, most of these methods fail to generalize well on real life image datasets because they are trained on synthetic datasets with a small range of blur kernels. This makes data-driven approaches inherently weak when it comes to real images. Therefore, applying image super-resolution independently of the blur kernel is still a challenging task. In this paper we propose IKR-Net, Iterative Kernel Reconstruction network, for blind SISR. In the proposed approach, kernel estimation and high resolution image reconstruction are carried out iteratively using deep models. The iterative refinement provides significant improvement in both the reconstructed image and the estimated blur kernel. IKR-Net achieves state-of-the-art results in blind SISR, especially for images with motion blur.
Suleyman Yildirim, Hasan F. Ates, Bahadir K. Gunturk
ICIP2
2021 Deep Learning-Based Blind Image Super-Resolution using Iterative Networks
abstract
Deep learning-based single image super-resolution (SR) consistently shows superior performance compared to the traditional SR methods. However, most of these methods assume that the blur kernel used to generate the low-resolution (LR) image is known and fixed (e.g. bicubic). Since blur kernels involved in real-life scenarios are complex and unknown, per-formance of these SR methods is greatly reduced for real blurry images. Reconstruction of high-resolution (HR) images from randomly blurred and noisy LR images remains a challenging task. Typical blind SR approaches involve two sequential stages: i) kernel estimation; ii) SR image reconstruction based on estimated kernel. However, due to the ill-posed nature of this problem, an iterative refinement could be beneficial for both kernel and SR image estimate. With this observation, in this paper, we propose an image SR method based on deep learning with iterative kernel estimation and image reconstruction. Simulation results show that the proposed method outperforms state-of-the-art in blind image SR and produces visually superior results as well.
Asfand Yaar, Hasan F. Ates, Bahadir K. Gunturk
VCIP2
2020 Signal Relation-Based Physical Layer Authentication
abstract
Most physical-layer authentication techniques use channel information to prevent spoofing attacks. In such techniques, one must estimate the channel information for each authentication procedure. However, when the number of pilots decreases, authentication accuracy also decreases due to low channel estimation quality. This paper proposes a novel signal relation-based authentication method that relies on the detection of received signal symbols and does not require the estimation of channel information in the testing stage. It is noteworthy that the authentication performance of the proposed scheme remains in a good level. We develop two different solutions for the detection of received signal symbols, namely, minimum mean-square error and long short-term memory. Extensive simulation results show the main insights of the proposed signal relation-based authentication method compared to conventional channel-based authentication method.
Mehmet Ali Aygül, Saliha Buyukcorak, Daniel B. da Costa 0001, Hasan F. Ates, Hüseyin Arslan
ICC4
2020 Spectrum Occupancy Prediction Exploiting Time and Frequency Correlations Through 2D-LSTM
abstract
The identification of spectrum opportunities is a pivotal requirement for efficient spectrum utilization in cognitive radio systems. Spectrum prediction offers a convenient means for revealing such opportunities based on the previously obtained occupancies. As spectrum occupancy states are correlated over time, spectrum prediction is often cast as a predictable time-series process using classical or deep learning-based models. However, this variety of methods exploits time-domain correlation and overlooks the existing correlation over frequency. In this paper, differently from previous works, we investigate a more realistic scenario by exploiting correlation over time and frequency through a 2D-long short-term memory (LSTM) model. Extensive experimental results show a performance improvement over conventional spectrum prediction methods in terms of accuracy and computational complexity. These observations are validated over the real-world spectrum measurements, assuming a frequency range between 832-862 MHz where most of the telecom operators in Turkey have private uplink bands.
Mehmet Ali Aygül, Mahmoud Nazzal, Ali Riza Ekti, Ali Gorcin, Daniel B. da Costa 0001, Hasan F. Ates, Hüseyin Arslan
VTC Spring6
2019 Multi-hypothesis contextual modeling for semantic segmentation
Hasan F. Ates, Sercan Sunetci
Pattern Recognit. Lett.1
2017 Improving Semantic Segmentation with Generalized Models of Local Context
Hasan F. Ates, Sercan Sunetci
CAIP (2)1
2016 Battle Damage Assessment based on self-similarity and contextual modeling of buildings in dense urban areas
abstract
Assessment of battle damages is significant both for tactical planning and for after-war relief efforts. In this study damaged buildings are detected using self-similarity descriptor in pre- and post-war satellite images. Detection accuracy is improved by the use of a contextual model that describes the building neighborhoods. Building footprints are utilized for accurate assessment of building-level changes and for the formation of neighborhood context. The Gaza Strip after 2014 Israel-Palestine conflict is analyzed with the suggested method and 84% true positive rate and 19% false positive rate are obtained on the average for detection of damaged buildings with respect to the ground truth data of UNOSAT.
Fatih Kahraman, Mumin Imamoglu, Hasan F. Ates
IGARSS3
2016 Disaster Damage Assessment of Buildings Using Adaptive Self-Similarity Descriptor
abstract
Assessment of damage caused by a disaster is significant for coordinating emergency response teams and planning emergency aid. In this letter, a robust method for rapid building damage assessment is proposed using pre- and postevent EO images and building footprints. The method uses a local self-similarity descriptor (SSD) for change detection in buildings, which is shown to be robust against variations in global illumination and small local deformations. The use of building footprints helps reduce the false alarms due to changes in nonbuilding areas. Footprint is also used to differentiate small and large buildings, extract the boundary region of a building, and adapt the descriptor computation accordingly. It is shown that the adaptive SSD provides a more accurate measure of local damage on the building. The 2010 Haiti Earthquake and Typhoon Haiyan 2013 Philippines are analyzed with the proposed method, and 75/82% true positive rate and 25/15% false positive rate are obtained for detection of collapsed buildings with respect to the ground truth data of UNITAR/UNOSAT and HOT.
Fatih Kahraman, Mumin Imamoglu, Hasan F. Ates
IEEE Geosci. Remote. Sens. Lett.3
2015 Disaster damage assessment for buildings using self-similarity descriptor
abstract
Assessment of damage caused by an earthquake is significant for coordinating emergency response teams and planning emergency aid. In this study, a robust method is proposed for detecting damaged buildings using pre- and post-event satellite images and building footprints. The method uses local self-similarity descriptor for change detection in buildings, which is shown to be robust against variations in illumination and small local deformations. The use of building footprints helps reduce the false alarms due to changes in non-building areas. The 2010 Haiti earthquake is analyzed with the suggested method and 72% true positive rate and 29% false positive rate are obtained for detection of collapsed buildings with respect to the ground truth data of UNITAR/UNOSAT.
Fatih Kahraman, Mumin Imamoglu, Hasan F. Ates
IGARSS3
2013 Decoder-Side Super-Resolution and Frame Interpolation for Improved H.264 Video Coding
abstract
In literature decoder-side motion estimation is shown to improve video coding efficiency of both H.264 and HEVC standards. In this paper we introduce enhanced skip and direct modes for H.264 coding using decoder-side super-resolution (SR) and frame interpolation. P- and B-frames are down sampled and H.264 encoded at lower resolution (LR). Then reconstructed LR frames are super-resolved using decoder-side motion estimation. Alternatively for B-frames, bidirectional true motion estimation is performed to synthesize a B-frame from its reference frames. For P-frames, bicubic interpolation of the LR frame is used as an alternative to SR reconstruction. A rate-distortion optimal mode selection algorithm determines for each MB which of the two reconstructions to use as skip/direct mode prediction. Simulations indicate an average of 1.04 dB PSNR improvement or 23.0% bit rate reduction at low bit rates when compared to H.264 standard. Average PSNR gains reach as high as 3.95 dB depending on the video content and frame rate.
Hasan F. Ates
DCC1
2011 Decoder side true motion estimation for very low bitrate B-frame coding
abstract
In H.264 standard, coding of motion vectors constitutes a significant portion of total bitrate especially at low bitrate regimes. This is because differential coding of motion vectors is inefficient when the bit budget is very low. In this paper, we propose a novel estimation and coding algorithm for motion vectors of B-frames at very low bitrates. In this method, the encoder selects the optimal motion vector from a limited set of candidate vectors that are determined at the decoder side using true motion estimation. Since these candidate vector sets are fixed by the decoder for each macroblock, there is no need for explicit coding of motion information, which reduces the bitrate required for coding. Also, true motion vector estimates are used for improved direct mode coding in B-frames. The algorithm provides an average of 0.68 dB PSNR gain for B-frames when compared to the reference H.264 results at the same bitrates. Simulation results also indicate significant improvement in visual quality of the compressed B-frames.
Hasan F. Ates, Burak Cizmeci
ICIP1
2010 Hierarchical quantization indexing for wavelet and wavelet packet image coding
Hasan F. Ates, Engin Tamer
Signal Process. Image Commun.1
2010 3-D Mesh Geometry Compression With Set Partitioning in the Spectral Domain
abstract
This paper explains the development of a highly efficient progressive 3-D mesh geometry coder based on the region adaptive transform in the spectral mesh compression method. A hierarchical set partitioning technique, originally proposed for the efficient compression of wavelet transform coefficients in high-performance wavelet-based image coding methods, is proposed for the efficient compression of the coefficients of this transform. Experiments confirm that the proposed coder employing such a region adaptive transform has a high compression performance rarely achieved by other state of the art 3-D mesh geometry compression algorithms. A new, high-performance fixed spectral basis method is also proposed for reducing the computational complexity of the transform. Many-to-one mappings are employed to relate the coded irregular mesh region to a regular mesh whose basis is used. To prevent loss of compression performance due to the low-pass nature of such mappings, transitions are made from transform-based coding to spatial coding on a per region basis at high coding rates. Experimental results show the performance advantage of the newly proposed fixed spectral basis method over the original fixed spectral basis method in the literature that employs one-to-one mappings.
Ulug Bayazit, Umut Konur, Hasan F. Ates
IEEE Trans. Circuits Syst. Video Technol.3
2009 Spherical Coding Algorithm for Wavelet Image Compression
abstract
In recent literature, there exist many high-performance wavelet coders that use different spatially adaptive coding techniques in order to exploit the spatial energy compaction property of the wavelet transform. Two crucial issues in adaptive methods are the level of flexibility and the coding efficiency achieved while modeling different image regions and allocating bitrate within the wavelet subbands. In this paper, we introduce the "spherical coder," which provides a new adaptive framework for handling these issues in a simple and effective manner. The coder uses local energy as a direct measure to differentiate between parts of the wavelet subband and to decide how to allocate the available bitrate. As local energy becomes available at finer resolutions, i.e., in smaller size windows, the coder automatically updates its decisions about how to spend the bitrate. We use a hierarchical set of variables to specify and code the local energy up to the highest resolution, i.e., the energy of individual wavelet coefficients. The overall scheme is nonredundant, meaning that the subband information is conveyed using this equivalent set of variables without the need for any side parameters. Despite its simplicity, the algorithm produces PSNR results that are competitive with the state-of-art coders in literature.
Hasan F. Ates, Michael T. Orchard
IEEE Trans. Image Process.1
2008 Fast inter-mode decision and selective quarter-pel refinement in H.264 video coding
abstract
In H.264 video coding standard, there exist several inter - prediction modes that use macroblock partitions with variable block sizes. Choosing a rate-distortion optimal coding mode for each macroblock is essential for the best possible coding performance, but also prohibitive due to the heavy computational complexity associated with the required rate-distortion calculations. Likewise, sub-pel motion refinement improves the coding efficiency, but becomes a major computational bottleneck when integer-pel search is executed fast. In this paper, we present a simple strategy to reduce the complexity of quarter-pel refinement and inter-mode decision with minimum loss of coding efficiency. Based on the results of the half-pel motion estimation step, our method evaluates the likelihood of each inter-coding mode being optimal. Then, quarter-pel refinement and actual rate and distortion are computed for only those coding modes with sufficient chance of being optimal. We claim that this method minimizes optimal mode estimation error at a given level of refinement and mode decision complexity. Simulation results show that the algorithm speeds up quarter-pel search and inter-mode selection modules by a factor of about 6 with less than 0.12 dB PSNR loss.
Hasan F. Ates
ICASSP1
2008 Rate-Distortion and Complexity Optimized Motion Estimation for H.264 Video Coding
abstract
H.264 video coding standard supports several inter- prediction coding modes that use macroblock (MB) partitions with variable block sizes. Rate-distortion (R-D) optimal selection of both the motion vectors (MVs) and the coding mode of each MB is essential for an H.264 encoder to achieve superior coding efficiency. Unfortunately, searching for optimal MVs of each possible subblock incurs a heavy computational cost. In this paper, in order to reduce the computational burden of integer-pel motion estimation (ME) without sacrificing from the coding performance, we propose a R-D and complexity joint optimization framework. Within this framework, we develop a simple method that determines for each MB which partitions are likely to be optimal. MV search is carried out for only the selected partitions, thus reducing the complexity of the ME step. The mode selection criteria is based on a measure of spatiotemporal activity within the MB. The procedure minimizes the coding loss at a given level of computational complexity either for the full video sequence or for each single frame. For the latter case, the algorithm provides a tight upper bound on the worst case complexity/execution time of the ME module. Simulation results show that the algorithm speeds up integer-pel ME by a factor of up to 40 with less than 0.2 dB loss in coding efficiency.
Hasan F. Ates, Yücel Altunbasak
IEEE Trans. Circuits Syst. Video Technol.1
2006 Rate-Distortion and Complexity Joint Optimization for Fast Motion Estimation In H.264 Video Coding
abstract
H.264 video coding standard offers several coding modes including inter-prediction modes that use macroblock partitions with variable block sizes, Choosing a rate-distortion optimal mode among these possibilities contributes significantly to the superior coding efficiency of the H.264 encoder. Unfortunately, searching for optimal motion vectors of each possible subblock incurs a heavy computational cost, In this paper, in order to reduce the complexity of integer-pel motion estimation, we propose a rate-distortion and complexity-joint optimization method that selects for each MB a subset of partitions to evaluate during motion estimation. This selection is based on simple, measures of spatio-temporal activity within the MB. The procedure is optimized to minimize mode estimation error at a certain level of computational complexity. Simulation results show that the algorithm speeds up the motion estimation module by a factor of up to 20 with little loss in coding efficiency.
Hasan F. Ates, Berkay Kanberoglu, Yücel Altunbasak
ICIP1
2006 Low Complexity Inter-Mode Selection for H.264
abstract
The coding efficiency of the H.264/AVC standard enables the transmission of high quality video over bandwidth limited networks. Due to the use of multiple macroblock (MB) partitions, the motion estimation module has extremely high complexity that makes it unpractical for most real-time applications on resource-limited platforms such as hand held devices. In this paper we propose a novel algorithm that significantly reduces the encoding complexity while maintaining high rate distortion performance. The proposed method reduces the motion estimation (ME) computational complexity by accurately predicting the optimal MB partitions and restricting the number of candidate modes based on a-priori probabilities computed from spatio-temporal information. The experimental results show that the speed up of UmHexagonS (one of the most efficient ME algorithms) can be doubled while maintaining the coding efficiency of full search.
Seydou-Nourou Ba, Yücel Altunbasak, Hasan F. Ates
ICIP3
2005 A High Performance Hardware Architecture for an SAD Reuse based Hierarchical Motion Estimation Algorithm for H.264 Video Coding
abstract
In this paper, we present a high performance and low cost hardware architecture for real-time implementation of an SAD reuse based hierarchical motion estimation algorithm for H.264/MPEG4 Part 10 video coding. This hardware is designed to be used as part of a complete H.264 video coding system for portable applications. The proposed architecture is implemented in Verilog HDL. The Verilog RTL code is verified to work at 68 MHz in a Xilinx Virtex II FPGA. The FPGA implementation can process 27 VGA frames (640 /spl times/ 480) or 82 CIF frames (352 /spl times/ 288) per second.
Sinan Yalcin, Hasan F. Ates, Ilker Hamzaoglu
FPL2
2005 SAD reuse in hierarchical motion estimation for the H.264 encoder
abstract
The adoption of multiple macroblock partitions with variable block sizes is one of the main reasons behind the superior coding efficiency of H.264 video coding standard. Unfortunately, in the motion estimation phase, repeating sum of absolute difference (SAD) calculations for every possible block size incurs a heavy computational cost for the encoder. In this paper, in order to reduce the encoder complexity, we propose a hierarchical block matching based motion estimation algorithm that uses a common set of SAD computations for motion estimation of different block sizes. Based on the hierarchical prediction and the median motion vector predictor of H.264, the algorithm defines a limited set of candidate vectors; and the optimal motion vectors for all partitions are chosen from this common set. Simulation results show that hierarchical estimation with SAD reuse reduces the total computations by a factor of 17.6 with slight loss in coding efficiency.
Hasan F. Ates, Yücel Altunbasak
ICASSP (2)1
2005 Wavelet image coding using the spherical representation
abstract
In this paper, we introduce the "spherical representation", which provides a new adaptive framework for modeling and coding the image information in wavelet subbands. Based on this representation, a practical coding algorithm is developed. This coder uses local energy as a direct measure to differentiate between parts of the wavelet subband and to decide how to allocate the available bitrate. As local energy becomes available at finer resolutions, i.e. in smaller size windows, the coder automatically updates its decisions about how to spend the bitrate. We use a hierarchical set of variables to specify and code the local energy up to the highest resolution, i.e. the energy of individual wavelet coefficients. The overall scheme is nonredundant, meaning that the subband information is conveyed using this equivalent set of variables without the need for any side parameters. Despite its simplicity, the algorithm produces PSNR results that are competitive with the state-of-art coders in literature.
Hasan F. Ates, Michael T. Orchard
ICIP (1)1
2005 An adaptive edge model in the wavelet domain for wavelet image coding
Hasan F. Ates, Michael T. Orchard
Signal Process. Image Commun.1
2003 Image interpolation using wavelet-based contour estimation
abstract
Successful image interpolation requires proper enhancement of high frequency content of image pixels around edges. We introduce a simple edge model to estimate high resolution edge profiles from lower resolution values. Pixels around edges are viewed as samples taken from one dimensional (1D) continuous edge profiles according to 1D smooth edge contours defining the sampling instants. The image is highpass filtered by wavelets and subpixel edge locations are estimated by minimizing the modeling error in the wavelet domain. Interpolation is carried out by applying the model, wherever applicable, together with a baseline interpolator (here, bilinear) in order to make edges look sharper without introducing artifacts. The results are compared to bilinear interpolation, and significant improvement in terms of SNR, edge sharpness and contour smoothness is observed.
Hasan F. Ates, Michael T. Orchard
ICASSP (3)1
2003 Nonlinear modeling of wavelet coefficients around edges
abstract
State of the art image coders make use of various methods to exploit intra and inter-band dependencies of wavelet coefficients in order to improve performance. While these efforts achieve considerable bitrate reduction for coding clusters of insignificant coefficients in smooth areas, most of the bitrate is spent on coding wavelet coefficients that are localized around edges in images. Recent research in literature is focused on developing new (linear or nonlinear) representations that deal with the rich and varying structures of pixel values around edges. In this paper, we use a simplified edge model to investigate the nonlinear dependencies that exist among wavelet coefficients, and introduce a nonlinear representation that is geared towards exploiting such dependencies for improved coding performance. Simulations support the relevance of the model, and we discuss our current efforts to incorporate these ideas into an actual image coder.
Hasan F. Ates, Michael T. Orchard
ICIP (1)1
2003 Image interpolation using wavelet-based contour estimation
abstract
Successful image interpolation requires proper enhancement of high frequency content of image pixels around edges. In this paper, we introduce a simple edge model to estimate high resolution edge profiles from lower resolution values. Pixels around edges are viewed as samples taken from one dimensional (1-D) continuous edge profiles according to 1-D smooth edge contours defining the sampling instants. The image is highpass filtered by wavelets and subpixel edge locations are estimated by minimizing the modeling error in the wavelet domain. Interpolation is carried out by applying the model, wherever applicable, together with a baseline interpolator (here, bilinear) in order to make edges look sharper without introducing artifacts. The results are compared to bilinear interpolation, and significant improvement in terms of SNR, edge sharpness and contour smoothness is observed.
Hasan F. Ates, Michael T. Orchard
ICME1
2000 Block motion estimation using wavelet filtering
abstract
Block matching motion compensation achieves savings in residual error energy at the cost of motion vector bit rate. While this tradeoff has proven valuable for many video sequences, there are many obvious examples of blocks for which the cost of motion vectors is not offset by the gains of motion compensation. This paper proposes a multiresolution framework for motion compensated prediction that offers a richer set of options for trading off motion compensation accuracy against the cost of motion vectors. The method improves the prediction at motion boundaries and on covered/uncovered regions, while reducing the bit rate by using less accurate motion vectors in smoother regions. The new algorithm is compared to the full-search block matching in various simulations, and the results show that the algorithm achieves a 10 to 30% reduction in motion vector bit rate. These saving are particularly important in low bit rate applications where motion overhead constitutes a significant percentage of overall bit rate.
Hasan F. Ates, Michael T. Orchard
ICASSP1