Dimitrios Androutsos

dblp:16/6872 · also Dimitri Androutsos · DBLP profile ↗
← Back
57ranked-venue papers
5as first author
5since 2021 · last 2026
0009-0009-4892-6324ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 38 · 3 first-authorComputer networks · 7Artificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 Failing Forward: Understanding Query Failure in Retrieval, Judgment, and Generation
abstract
Modern information retrieval pipelines combine retrieval, LLM-based generation, and LLM-based judgment, and a poor outcome may originate in any of the three stages. Existing work studies these failures in isolation. This paper instead asks whether query difficulty itself transfers across the three stages: are the same queries hard to retrieve, hard to generate for, and hard to judge? Using four years of TREC Deep Learning benchmarks (2019–2022), we define hard-to-retrieve, hard-to-generate, and hard-to-judge query sets under a unified quartile-based operationalization and analyze their overlap, their stability across system configurations, and the linguistic and semantic causes of failure in each task. We find that the three sets overlap only weakly; three-way overlap is at or below the level expected under independence, indicating that difficulty is largely task-conditioned and does not transfer reliably across stages. The overlap structure is nonetheless stable across retrievers, generators, and judging setups, suggesting that task-specific difficulty is driven by query characteristics interacting with each task's inductive biases rather than by model choice. We further induce a data-driven typology of failure causes and show that conditioning generation on task-relevant difficulty cues yields consistent gains in answer quality.
Negar Arabzadeh, Mohammad Hossein Saliminabi, Dimitrios Androutsos, Morteza Zihayat, Ebrahim Bagheri
SIGIR4
2026 LearnDCG: End-to-End Joint Optimization of Ranker and Loss in Neural Ranking
Mohammad Hossein Saliminabi, Dimitrios Androutsos, Ebrahim Bagheri
SIGIR2
2026 Diffusion-based generative modeling for expert team formation
abstract
Forming effective expert teams is central to domains where solving complex problems requires diverse, complementary skills. However, automating this task is highly challenging due to sparse co-occurrence data, long-tailed expert participation, and the combinatorial complexity of unseen skill configurations. Existing graph-based, probabilistic, and neural approaches often struggle with generalization, fairness, and robustness, leading to biased selections that favor historically popular experts over more suitable candidates. To address these challenges, we propose a generative framework for expert team formation based on denoising diffusion probabilistic models. We cast team formation as skill-conditioned imputation (i.e., inpainting), where skills are treated as observed context and the expert component is generated via conditional diffusion sampling. This design enables our method to preserve semantic skill–expert alignment, mitigate data sparsity, and generate diverse yet contextually coherent teams. Extensive experiments on DBLP and DOTA2 datasets show that our model consistently outperforms state-of-the-art baselines, achieving over 3 × higher recall (16.4% vs. 5.0%) and MAP (9.7% vs. 2.2%) on DBLP, while delivering more than 5 × improvement in MRR (13.3% vs. 2.5%) on DOTA2. Fairness analysis further demonstrates that our method reduces average overlap with the top-100 most popular experts to 2.6, compared to 86.7 for the strongest baseline, and achieves near-optimal diversity with NDKL ≈ 0.1 under high non-popular expert ratios. For reproducibility purposes , we made our code and model publicly available at https://github.com/17shiraz/DiffTF .
Mohammad Hossein Saliminabi, Sajad Ebrahimi 0001, Radin Hamidi Rad, Dimitrios Androutsos, Fattane Zarrinkalam, Ebrahim Bagheri
Inf. Process. Manag.4
2025 LLM-as-a-Judge in Entity Retrieval: Assessing Explicit and Implicit Relevance
abstract
Entity retrieval plays a critical role in information access systems, yet the development and evaluation of retrieval models remain constrained by the limited availability of high-quality supervision. While recent work has demonstrated the utility of large language models (LLMs) as relevance assessors in passage and document retrieval, their reliability in the context of entity retrieval-where targets are abstract, underspecified, and often semantically sparse-remains unexplored. In this work, we evaluate LLM-based judgments against two complementary supervision signals: human-annotated relevance labels from the DBpedia-Entity benchmark and implicit feedback from user clicks in the LaQuE dataset. We show that LLMs exhibit strong agreement with expert annotations and replicate user click patterns with over 91% agreement, suggesting alignment with behavioral judgments despite noisy input queries. We further identify and analyze systematic mismatches for user clicks on irrelevant entities. Our findings establish LLMs not only as effective annotators for entity relevance judgment-even when given only the entity title-but also as powerful tools for predicting click-through behavior and simulating explainable user intent. Our code, prompts, and data are publicly available at: https://github.com/17shiraz/ClickLLM
Mohammad Hossein Saliminabi, Negar Arabzadeh, Dimitrios Androutsos, Morteza Zihayat, Ebrahim Bagheri
CIKM4
2025 Unlocking wisdom: enhancing biomedical question answering with domain knowledge
Bita Azad, Mahdiyar Ali Akbar Alavi, Parastoo Jafarzadeh, Faezeh Ensan, Dimitrios Androutsos
Knowl. Inf. Syst.5
2019 Single Image Super-Resolution via Cascaded Parallel Multisize Receptive Field
abstract
Recovering a High-Resolution (HR) image from a Low-Resolution (LR) image is the main concept of image Super-Resolution(SR). These days Convolution Neural Networks(CNN) have become very popular and efficient in generating HR image from a LR image. Although CNNs are widely used with great performance improvements, there is still much room for improvement. There has always been a trade-off between the number of parameters and performance enhancement. Inspired by the Inception architecture of GoogleNet, we have developed a novel approach with an efficiently increased number of receptive field by cascading different kernel sizes to reconstruct the HR image. Experimental results shows that the proposed model outperforms the state-of-art methods.
Debjoy Chowdhury, Dimitrios Androutsos
ICIP2
2018 Edge-Based Loss Function for Single Image Super-Resolution
abstract
In recent years, convolutional neural networks have shown state-of-the-art performance on the task of single-image super-resolution. Although these proposed networks have shown high-quality reconstruction results, the use of the mean-squared error (MSE) loss function for training tends to produce images that are overly smooth and blurry. The MSE does not consider image structures that are often important for achieving high human-perceived image quality. We propose a novel edge-based loss function to improve super-resolution resconstruction of images. Our loss function directly optimizes the edge pixels of the reconstructed image, thus driving the trained network to produce high-quality salient edges and thus sharper images. Extensive quantitative and qualitative results show that our proposed loss function significantly outperforms the MSE.
George Seif, Dimitrios Androutsos
ICASSP2
2017 Human visual system inspired saliency guided edge preserving tone-mapping for high dynamic range imaging
abstract
With the growing popularity of High Dynamic Range Imaging (HDR), the necessity for advanced tone-mapping techniques has greatly increased. We propose a Human Visual System (HVS)-inspired Saliency Guided Edge-preserving Tone Mapping method (HSGETM) that uses saliency region information of an HDR image as input to the guided filter for a base and detail image layer separation. After detail layer enhancement and base layer compression with constant weights, a new edge preserved tone mapped image is composed by adding the layers back together with saturation and exposure adjustments. The filter operation is faster due to the use of the guided filter which has O(N) time operation with N number of pixels. Experimental results demonstrate that the HSGETM has higher edge, and naturalness preserving capability, which is homologous to the HVS, as compared to other state-of-the-art tone mapping approaches.
Nipu R. Barai, Matthew J. Kyan, Dimitrios Androutsos
ICIP3
2016 Depth guided image completion for structure and texture synthesis
abstract
This paper presents a method that separates the image completion process into structure and texture synthesis. A method is first introduced for completing the respective depth map through the use of a morphological diffusion-based operation, ensuring that structure is propagated smoothly across the unwanted region. These estimated depth values are then used to guide completion of the missing color image texture. A fast exemplar based PatchMatch approach is adopted and an extension of the coherence-based objective function introduced by Wexler et al. is used to complete the resulting texture in both the color and depth images. Finally, we provide experimental results to demonstrate the superiority of our method against others in a variety of different scenes.
Michael Ciotta, Dimitrios Androutsos
ICASSP2
2015 Adaptive exposure fusion for high dynamic range imaging
abstract
This paper proposes a novel exposure fusion algorithm that directly fuses exposure bracketed shots into a displayable image. Most techniques targeted for direct fusion do not have an effective exposure control mechanism that can compensate for the limitations in the human visual system (HVS). The proposed algorithm offers a novel approach that adaptively adjusts its parameter for the best viewing experience. Changing the parameter adaptively aims to mitigate the perceptual loss of details caused by Weber's effect, brightening the dark regions while darkening the bright regions. The proposed method yields a perceptually detailed image when compared against other methods of similar nature.
Sidhdharthkumar Patel, Dimitrios Androutsos, Matthew J. Kyan
ICIP2
2015 A framework for estimating relative depth in video
Richard Rzeszutek, Dimitrios Androutsos
Comput. Vis. Image Underst.2
2015 Sequential Subspace Estimator for biometric authentication
Obaidul Malek, Anastasios N. Venetsanopoulos, Dimitrios Androutsos, Lian Zhao
Neurocomputing3
2014 Label propagation through edge-preserving filters
abstract
In this paper we investigate methods for propagating automatically generated or user-defined labels through an image using edge-preserving filters. We focus on the domain transform filter as it has been used for propagation purposes in the past. The method we present addresses some of the numerical issues that arise with using the filter directly and also improve on the results by better respecting the underlying image structure during the label propagation. Finally we also demonstrate how a filter-based approach is preferable to using global optimization for interpolating automatically generated sparse features.
Richard Rzeszutek, Dimitrios Androutsos
ICASSP2
2014 Robust Semi-Automatic Depth Map Generation in Unconstrained Images and Video Sequences for 2D to Stereoscopic 3D Conversion
abstract
We describe a system for robustly estimating synthetic depth maps in unconstrained images and videos, for semi-automatic conversion into stereoscopic 3D. Currently, this process is automatic or done manually by rotoscopers. Automatic is the least labor intensive, but makes user intervention or error correction difficult. Manual is the most accurate, but time consuming and costly. Noting the merits of both, a semi-automatic method blends them together, allowing for faster and accurate conversion. This requires user-defined strokes on the image, or over several keyframes for video, corresponding to a rough estimate of the depths. After, the rest of the depths are determined, creating depth maps to generate stereoscopic 3D content, with Depth Image Based Rendering to generate the artificial views. Depth map estimation can be considered as a multi-label segmentation problem: each class is a depth. For video, we allow the user to label only the first frame, and we propagate the strokes using computer vision techniques. We combine the merits of two well-respected segmentation algorithms: Graph Cuts and Random Walks. The diffusion from Random Walks, with the edge preserving of Graph Cuts should give good results. We generate good quality content, more suitable for perception, compared to a similar framework.
Raymond Phan, Dimitrios Androutsos
IEEE Trans. Multim.2
2014 On the Power Allocation Problem in the Gaussian Interference Channel with Proportional Rate Constraints
abstract
This paper takes an analytical approach to solving the optimization problem of finding the power allocation that maximizes the sum-rate of the Gaussian interference channel with any linear power (interference) constraint and proportional rate constraints. It is proved that the sum-rate of the Gaussian interference channel restricted to proportional rate constraints does not have a critical point and the maximum sum-rate subject to said constraints occurs at the boundary of the domain formed by the plane representing the linear power constraint. This is accomplished by using analytic geometry in higher dimensions to show that the curve of intersection of the sum-rate and the proportional rate constraints is always increasing, and intersects the boundary plane representing the linear power constraint at a unique point. A polynomial time (in the number of users) centralized algorithm that finds this point of optimal power allocation is proposed. This is a significant improvement over existing algorithms for related power allocation problems which have exponential time complexity in the number of users. Two distributed algorithms with linear and constant complexities are also presented. Simulation results supporting the analysis and demonstrating the performances of the algorithms are presented.
Kandasamy Illanko, Alagan Anpalagan, Ekram Hossain 0001, Dimitrios Androutsos
IEEE Trans. Wirel. Commun.4
2012 Low complexity energy efficient power allocation for green cognitive radio with rate constraints
abstract
This paper combines two emerging research areas: green communications and cognitive radio. A green cognitive radio network must be accountable for its energy expenditure. Energy expenditure of a cognitive base station is reduced by maximizing the bits/Joule energy efficiency (EE) of its transmissions. Any high complexity solution to this optimization problem will spend too much energy in computation. This paper presents a low complexity solution to the problem of finding the power allocation that maximizes the EE, while limiting the interference to the primary users and meeting the users' minimum rate requirements. The objective function of the optimization problem is not concave. Charnes-Cooper Transformation is applied to the problem to convert it into a concave program. KKT conditions were analyzed instead of the Lagrangian dual in lieu of low complexity solutions. A power allocation procedure that branches into two main cases depending on the channel gains is proposed. In the first case, an exact solution is obtained by solving a single non-linear equation that produces a common water level. In the second case, a near optimal solution in closed form is given. Simulation results supporting the analytical green solutions are presented.
Kandasamy Illanko, Muhammad Naeem 0001, Alagan Anpalagan, Dimitrios Androutsos
GLOBECOM4
2012 Adaptive 2D to 3D image conversion using a hybrid Graph Cuts and Random Walks approach
abstract
In this paper, we propose an adaptive method for 2D to 3D conversion of images using a user-aided process based on Graph Cuts and Random Walks. Given user-defined labeling that correspond to a rough estimate of depth, the system produces a depth map which, combined with a 2D image can be used to synthesize a stereoscopic image pair. The work presented here is an extension of work done previously combining the popular Graph Cuts and Random Walks image segmentation algorithms. Specifically, we have made the previous approach adaptive, as well as improved the quality of the results. This is achieved by feeding information from the Graph Cuts result into the Random Walks process at two different stages, and using edge and spatial information to adapt various weights. The results show that we can produce good quality stereoscopic 3D image pairs using a simple yet adaptive approach.
Mohammad Fawaz, Raymond Phan, Richard Rzeszutek, Dimitrios Androutsos
ICASSP4
2012 Competitive pricing for spectrum subleasing for future wireless ad hoc networks
abstract
This paper envisions a near future in which the proliferation of wireless ad hoc networks in urban centers causes excessive spectrum pollution on currently allocated unlicensed bands. One solution for this problem is for the operators to lease freshly released spectrum from the regulators and sublease it to agencies in major cities. We consider one such operator who divides an urban area into regions and subleases spectrum with the condition that the interference measured at boundary points should not exceed a threshold. The subleasing pricing structure has a fixed part, as well as a variable part that discounts the price based on the margin between the interference threshold and the actual interference. The slope of the variable part is called the discount rate and is determined by a competition that is modeled as a game within a game. For a fixed discount rate, the competition between the customers forms a strategic game. The end result of this game becomes the input to the Stackelberg game between the customers as a whole on the one side and the operator on the other side. We derive the mild condition under which the strategic game of the customers has a unique Nash equilibrium, and obtain an explicit closed form solution for the equilibrium point. This result is then used to derive the best response of the operator and the optimum (Stackelberg equilibrium) discount rate the operator would want to offer. Numerical results obtained through simulations that support the analysis are also provided.
Kandasamy Illanko, Alagan Anpalagan, Dimitrios Androutsos
ICC3
2012 Depth estimation for semi-automatic 2D to 3D conversion
abstract
The conversion of monoscopic footage into stereoscopic or multiview content is a difficult and time consuming task. A number of semi-automatic methods have been developed to speed up the process and provide some control to the user. However these methods require that the user provide detailed labels indicating the relative depth of objects in the scene. In this paper we present a method to automatically estimate depth in such a way that it is amenable to semi-automatic conversion. The method is designed to simplify the depth labelling task so that the user does not have to provide as many depth labels.
Richard Rzeszutek, Raymond Phan, Dimitrios Androutsos
ACM Multimedia3
2011 Semi-automatic 2D to 3D image conversion using a hybrid Random Walks and graph cuts based approach
abstract
In this paper, we present a semi-automated method for converting conventional 2D images to stereoscopic 3D. User-defined strokes that correspond to a rough estimate of the depth values in the scene are defined for the image of interest. With these strokes, our system thus determines what the depth values are for the rest of the image, producing a depth map that is ultimately used to create a stereoscopic image pair. Our work is based on a similar scheme which employs Random Walks. However, the related work is quite complex, with many processing steps required to produce the final stereoscopic im age pair. Combined with the evident shortcomings of the related work, but noting the merits of Random Walks, we propose a system that is a hybrid between Random Walks, and the popular Graph Cuts segmentation paradigm. Both segmentation algorithms are used to generate a final cohesive depth map, thus combining the merits of both frameworks together. The generated results show that we can produce good quality stereoscopic image pairs, while using a much more simplified method in comparison to the related Random Walks scheme.
Raymond Phan, Richard Rzeszutek, Dimitrios Androutsos
ICASSP3
2011 Convex Structure of the Sum Rate on the Boundary of the Feasible Set for Coexisting Radios
abstract
The power allocation that maximizes the sum rate of transceivers operating in the same frequency band is a difficult non-convex problem. In our earlier work, we proved that for transceivers operating under a total power constraint, the maximum sum rate occurs at the boundary of the feasible set formed by the hyper plane representing the power constraint. This finding is nontrivial considering that we are dealing with an interference limited system. In this paper, we study the convex structure of the sum rate on the boundary of the power constraint hyper plane. For two transceivers, we prove that the sum rate is always convex on the line created by the power constraint equality. In the case of three transceivers, we identify a region in the middle of the plane created by the power constraint equality, where the sum rate is concave. This is significant because it is in the middle of the boundary plane that the power allocation can be expected to be fair to all three users. We also provide a power allocation protocol and an algorithm that distribute the power among the transceivers with fairness. Simulation results are provided to support the theorems proven in the paper as well as to demonstrate the convergence of the algorithm to the global maximum sum rate. Results of the algorithm are compared with solutions based on Game theory.
Kandasamy Illanko, Alagan Anpalagan, Dimitrios Androutsos
ICC3
2011 Stackelberg Game on the Boundary of Coexistence
abstract
This paper combines Convex analysis and Game theory to investigate the problem of maximizing the sum rate of transceivers operating in the same frequency band. In our earlier work, we proved that for transceivers operating under a total power constraint, the power distribution that maximizes the sum rate lies on the boundary of the feasible set formed by the power constraint. In this paper, we first prove that for two users, the sum rate is convex on the boundary formed by the line segment representing the power constraint, and the maximum sum rate is achieved when all the power is allocated to one of the users. Obviously, such a power allocation is unfair to the other user. We consider a scenario in which the first user is willing to sell some of the power allocated to him to the second user. We use Stackelberg Game theory to analyze this scenario and prove the existence of a unique competitive equilibrium. We derive the best response functions, and determine the optimum price the first user must charge and the optimum amount of power the second user should buy at this price. We also use simulation to obtain the best response functions and the equilibrium point, and demonstrate their agreement with our analytical results.
Kandasamy Illanko, Alagan Anpalagan, Dimitrios Androutsos
ICC3
2011 Semi-automatic 2D to 3D image conversion using scale-space Random Walks and a graph cuts based depth prior
abstract
In this paper, we present a semi-automated method for converting conventional 2D images into stereoscopic 3D. User-defined strokes corresponding to a rough estimate of the depth values in the scene are defined for the image of interest. With these, our system determines the depth values for the rest of the image, producing a depth map that can be used to create stereoscopic 3D image pairs. Our work is based on a similar scheme, using the Random Walks segmentation paradigm. However, the related work is quite complex, with many processing steps required to produce the final stereoscopic image pair. Combined with its evident shortcomings, but noting the merits, we propose a system employing Random Walks, while incorporating information from the popular Graph Cuts segmentation paradigm. Thus, a final cohesive depth map is produced, combining the merits of both. The results show that we can produce good quality stereoscopic image pairs, while using a much more simplified method in comparison to the related work.
Raymond Phan, Richard Rzeszutek, Dimitrios Androutsos
ICIP3
2011 Semi-automatic synthetic depth map generation for video using random walks
abstract
We present a method for easily generating depth maps from monoscopic (i.e. “2D”) video footage in order to convert them into stereoscopic, or “3D”, footage. Our method uses user-defined strokes for a number of keyframes in the original footage and interpolates between the keyframes to provide a sparse labelling for each frame. We then apply the Random Walks algorithm to the footage to provide depth estimates based on the input provided by the user. These depth maps can then be used to generate novel views through depth-based image rendering.
Richard Rzeszutek, Raymond Phan, Dimitrios Androutsos
ICME3
2010 Dual Methods for Power Allocation for Radios Coexisting in Unlicensed Spectra
abstract
The power allocation that maximizes the sum rate of transceivers operating in the same frequency band is a difficult non-convex problem. Lack of a convex structure excludes the direct application of Lagrangian dual techniques as the duality gap might not be zero. This paper advances current knowledge by introducing three significant steps in finding a solution. First, we show that for transceivers operating under a total power constraint, the maximum sum rate occurs at the boundary of the feasible set formed by the hyper plane representing the power constraint. This conclusion is nontrivial considering that we are dealing with an interference limited system. Second, we prove that the duality gap is zero for this problem, despite the lack of concavity of the objective. We do this by showing that the maximum sum rate is concave in the power constraint. Third, we propose an iterative algorithm that finds the optimal power allocation by solving the dual problem. Simulation results are provided to support the theorems proven in the paper as well as to demonstrate the convergence of the algorithm to the global maximum sum rate. Results of the algorithm are also compared with solutions based on Game theory.
Kandasamy Illanko, Alagan Anpalagan, Dimitrios Androutsos
GLOBECOM3
2010 An Optimal and Fair Distributed Algorithm for Power Allocation for Radios Coexisting in Unlicensed Spectra
abstract
This paper presents a simple synchronous distributed power allocation algorithm that maximizes the total transmission rate of a number of radios operating in an unlicensed band, with either a total power constraint or individual power constraints. The redistribution of power by the algorithm also results in a fairer rate distribution. The algorithm is not based on Game theory or Lagrangian dual. Rather it uses the sensitivity of each user's rate to changes in the power levels of all users in the system, to steer the power distribution towards the global maximum sum rate. The algorithm's complexity scales with the number of users in the system. Simulation results demonstrate that the algorithm does converge to the global maximum sum rate and, at the same time, redistributes the power among the users to achieve a more equitable rate distribution. Results of the algorithm are also compared with solutions based on Game theory.
Kandasamy Illanko, Alagan Anpalagan, Dimitrios Androutsos
ICC3
2010 Content-based retrieval of logo and trademarks in unconstrained color image databases using Color Edge Gradient Co-occurrence Histograms
Raymond Phan, Dimitrios Androutsos
Comput. Vis. Image Underst.2
2010 Self-Organizing Maps for Topic Trend Discovery
abstract
The large volume of data on the Internet makes it extremely difficult to extract high-level information, such as recurring or time-varying trends in document content. Dimensionality reduction techniques can be applied to simplify the analysis process but the amount of data is still quite large. If the analysis is restricted to just text documents then Latent Dirichlet Allocation (LDA) can be used to quantify semantic, or topical, groupings in the data set. This paper proposes a method that combines LDA with the visualization capabilities of Self-Organizing Maps to track topic trends over time. By examining the response of a map over time, it is possible to build a detailed picture of how the contents of a dataset change.
Richard Rzeszutek, Dimitrios Androutsos, Matthew J. Kyan
IEEE Signal Process. Lett.2
2009 Robust affine invariant shape image retrieval using the ICA Zernike Moment Shape Descriptor
abstract
In this paper, we proposed a new affine invariant region-based shape descriptor, the ICA Zernike Moment Shape Descriptor (ICAZMSD). IndependentComponent Analysis (ICA) is first used to turn the original shape into a canonical form, in which the effects of scaling and skewing are eliminated. Next, the properties of the Zernike transform is used to further eliminate the effects of any possible rotation and reflection of the canonical shapes, in extracting the Zernike moments as the affine invariant region-based descriptors. Using the proposed ICAZMSD as shape feature, shape-based image retrieval experiments on a 4000 complex shape image database and a 5600 simple shape image database, show promising retrieval rates of 99.80% and 92.25%, respectively.
Ye Mei, Dimitrios Androutsos
ICIP2
2009 Interactive rotoscoping through scale-space random walks
abstract
We present a novel rotoscoping method that provides accurate object boundaries while still remaining simple to use. We have designed this method so that a user familiar with existing, spline-based rotoscoping techniques will not need to significantly modify their workflow. Our method utilizes a modified version of the random walks segmentation algorithm, known as scale-space random walks (SSRW), to provide a level of accuracy simply not possible using splines alone. Furthermore, due to the nature of the SSRW algorithm, we are able to provide a realtime, or near realtime, level of interaction with the user.
Richard Rzeszutek, Thomas F. El-Maraghi, Dimitrios Androutsos
ICME3
2009 Robust Affine Invariant Region-Based Shape Descriptors: The ICA Zernike Moment Shape Descriptor and the Whitening Zernike Moment Shape Descriptor
abstract
In this letter, we proposed two new affine invariant region-based shape descriptors, the ICA Zernike moment shape descriptor (ICAZMSD) and the whitening Zernike moment shape descriptor (WZMSD). Either independent component analysis (ICA) or whitening, is first used to turn the original shape into a canonical form, in which the effects of scaling and skewing are eliminated. Next, the properties of the Zernike transform are used to further eliminate the effects of any possible rotation and reflection of the canonical shapes, in extracting the Zernike moments as the affine invariant region-based descriptors. Using the proposed ICAZMSD as shape feature, shape-based image retrieval experiments on a 4000 complex shape image database and on a 5600 simple shape image database, show retrieval rates of 99.80% and 92.25%, respectively. Using the proposed WZMSD as shape feature, the corresponding retrieval rates are 99.79% and 92.22%, respectively. The proposed WZMSD has almost equal performance to the proposed ICAZMSD, while having lower computational requirements.
Ye Mei, Dimitrios Androutsos
IEEE Signal Process. Lett.2
2008 Logo and trademark detection in images using Color Wavelet Co-occurrence Histograms
abstract
The use of histograms to characterize edge information in an image is a common technique for image indexing and retrieval. Two techniques that have recently shown promising success are the edge gradient histogram (EGH) and the co-occurrence edge color histogram (CECH). In this paper, we present a system for logo and trademark retrieval from a database of color logo images. Through the use of a 5-dimensional co-occurrence histogram, the system proposed captures the co-occurrence of colors and wavelet decomposition coefficients of pairs of pixels. The result is a more precise characterization of the spatial distribution of edge information in an image than the ones produced by the EGH and the CECH. We call this 5-dimensional co-occurrence histogram the color wavelet co-occurrence histogram (CWCH). Results demonstrate that our retrieval system performs better than both the EGH and the CECH.
Ali Hesson, Dimitrios Androutsos
ICASSP2
2008 Unconstrained logo and trademark retrieval in general color image databases using Color Edge Gradient Co-occurrence Histograms
abstract
In this paper, we present a logo and trademark retrieval system for general, unconstrained, color image databases, extending the color edge co-occurrence histogram (CECH) object detection scheme. We introduce more accurate information to the CECH, by virtue of incorporating color edge detection using vector order statistics. This produces a more accurate representation of edges in color images, in comparison to the simple color pixel difference classification of edges as seen in the CECH. Our proposed method is thus reliant on edge gradient information, thus we call it the color edge gradient co-occurrence histogram (CEGCH). We also introduce a novel color quantization scheme based in the hue-saturation-value (HSV) color space, illustrating that it is more suitable for image retrieval in comparison to the color quantization scheme introduced with the CECH. Results illustrate that our retrieval system retrieves logos and trademarks with good accuracy, outperforming the use of the CECH in image retrieval with higher precision and recall.
Raymond Phan, John Chia, Dimitrios Androutsos
ICASSP3
2008 Indexing and retrieval of compound color objects using co-occurrence histograms of color and wavelet features
abstract
In this paper, we present a system for the retrieval and indexing of images of compound color objects. Compound color objects are objects that consist of a specific set of colors that are spatially arranged in a unique way. Examples of compound color objects include flags, trademarks, logos, and cartoons. In this paper, we apply our proposed technique to logos and trademarks. We introduces a 5-dimensional cooccurrence histogram that captures color and texture information simultaneously. We use the multi-resolution analysis feature with the Coiflet wavelet to capture the texture information in the image. We call this 5-dimensional histogram the color wavelet co-occurrence histogram (CWCH). We show that the CWCH performs better than the edge gradient histogram (EGH) and the color edge co-occurrence histogram (CECH) and the MPEG-7 edge histogram descriptor (EHD).
Ali Hesson, Dimitrios Androutsos
ICIP2
2008 Wavelet-based color texture retrieval using the independent component color space
abstract
In this paper, we propose a wavelet based color texture retrieval method using the independent component color space. In color texture retrieval, the product of low dimensional marginal distributions of wavelet coefficients from different color layers are preferred to substitute or approximate their high dimensional joint distributions in order to avoid the curse of dimensionality. However, the RGB color spaces is a highly correlated color space and the extractedwavelet coefficients from different layers are also correlated, which means such a substitution or approximation will not be adequate. To solve the problem, we use independent component analysis to decorrelate the R, G and B layers into three new independent layers before applying wavelet decomposition on the color texture images. In the feature extraction(FE) step of the proposed method, generalized Gaussian density (GGD) are used to model the marginal distribution of wavelet coefficients, and the extracted model parameters are used as features. In the similarity measurement (SM) step of the proposed method, the Kullback-Leibler distance(KLD) is calculated as feature distance, using the extracted model parameters of the query texture images and those of the images in the database. Experimental results on a database of 1120 color texture images indicate that the proposed method greatly overperforms its RGB based counterpart that ignores the inter-layer correlation, and its counterpart which uses the I1I2I3 colorspace.
Ye Mei, Dimitrios Androutsos
ICIP2
2008 Color texture retrieval using wavelet decomposition on the Hue/Saturation plane
abstract
This paper introduces a new color texture retrieval method, based on wavelet decomposition of color texture images on the Hue/Saturation plane. The complex Hue/Saturation representation, used in the proposed method, in the Cartesian coordinates solves the problem of the traditional Hue representation in the angular coordinate and implicitly contains the correlation between the two chromatic parts of color pixel. Histograms of the extracted complex wavelet coefficients are used as features to represent the color texture images. The Jeffery divergence is used as a distance measure between the features of the query texture images and those of the images in the database. Experimental results show the new method overperforms its countparts that do not consider the correlations between color layers.
Ye Mei, Dimitrios Androutsos
ICME2
2008 Affine invariant shape descriptors: The ICA-Fourier descriptor and the PCA-Fourier descriptor
abstract
In this paper, we propose two new affine invariant shape descriptors, the ICA-Fourier descriptor and the PCA-Fourier descriptor. We tested the descriptors by using them as features for shape based silhouette image retrieval. Experiments on a 1000 silhouette image database show promising retrieval rates of 95.41% and 93.63%, using the ICA-Fourier descriptor and the PCA-Fourier descriptor, respectively. The relationship between those two descriptors are also explained. The proposed PCA-Fourier descriptor is computationally more efficient than its ICA counterpart, while having comparable performance.
Ye Mei, Dimitrios Androutsos
ICPR2
2008 Color Image Watermarking Using Multidimensional Fourier Transforms
abstract
This paper presents two vector watermarking schemes that are based on the use of complex and quaternion Fourier transforms and demonstrates, for the first time, how to embed watermarks into the frequency domain that is consistent with our human visual system. Watermark casting is performed by estimating the just-noticeable distortion of the images, to ensure watermark invisibility. The first method encodes the chromatic content of a color image into the CIE chromaticity coordinates while the achromatic content is encoded as CIE tristimulus value. Color watermarks (yellow and blue) are embedded in the frequency domain of the chromatic channels by using the spatiochromatic discrete Fourier transform. It first encodes and as complex values, followed by a single discrete Fourier transform. The most interesting characteristic of the scheme is the possibility of performing watermarking in the frequency domain of chromatic components. The second method encodes the components of color images and watermarks are embedded as vectors in the frequency domain of the channels by using the quaternion Fourier transform. Robustness is achieved by embedding a watermark in the coefficient with positive frequency, which spreads it to all color components in the spatial domain and invisibility is satisfied by modifying the coefficient with negative frequency, such that the combined effects of the two are insensitive to human eyes. Experimental results demonstrate that the two proposed algorithms perform better than two existing algorithms - ac- and discrete cosine transform-based schemes.
Tsz Kin Tsui, Xiao-Ping Zhang 0002, Dimitrios Androutsos
IEEE Trans. Inf. Forensics Secur.3
2006 Color Image Watermarking Using the Spatio-Chromatic Fourier Transform
abstract
In this paper, a color watermarking algorithm is proposed. The method encodes the chromatic content of a color image as CIE a*b*chromaticity coordinates whereas the achromatic content is encoded as CIE L tristimulus value. Color watermarks (yellow and blue) are embedded in the frequency domain of the chromatic channels by using the Spatio Chromatic Discrete Fourier Transform (SCDFT). It first encodes a*and b*as complex values, followed by a single discrete Fourier Transform. Watermark casting is performed by estimating the Just-Noticeable distortion (JND) of the images, to ensure watermark invisibility. The most interesting characteristics of the new scheme is the possibility of performing watermarking in the frequency domain of chromatic components.
Tsz Kin Tsui, Xiao-Ping Zhang 0002, Dimitrios Androutsos
ICASSP (2)3
2006 Quaternion image watermarking using the spatio-chromatic fourier coefficients analysis
abstract
In this paper, a new color watermarking algorithm that uses the quaternion fourier transform (QFT) to mark the La∗b∗ components of color images is presented. First, we propose an interpretation of the QFT coefficients using the Spatio-Chromatic Fourier analysis, so the effects of any changes to the coefficients can be predicted. Next, watermark casting is performed by modifying the positive and negative coeffi-cients together. The idea is twofold: Robustness is achieved by embedding a color watermark in the coefficient with pos-itive frequency, which spreads it to all components in the spatial domain. On the other hand, invisibility is satisfied by modifying the coefficient with negative frequency, such that the combined effects of the two are insensitive to human eyes.
Tsz Kin Tsui, Xiao-Ping Zhang 0002, Dimitrios Androutsos
ACM Multimedia3
2006 A distributed fault-tolerant MPEG-7 retrieval scheme based on small world theory
abstract
A small world search agent employing peer-to-peer (P2P) concepts borrowed from sociology is employed for performing image retrievals in a small world distributed media index. The Small World Indexing Method (SWIM) allows for a highly networked architecture where index information does not exist as a separate entity on a specific server, but rather is stored within the actual media objects themselves. Since each media object is only responsible for a small portion of the overall index, the loss of portions of the overall network (data objects) accounts for only a small degradation in the overall retrieval performance. Building upon previous work, the graceful degradation which is provided by the SWIM system is addressed here for retrievals which are performed using small world user agents on a large set of MPEG-7 described images.
Panagiotis Androutsos, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
IEEE Trans. Multim.2
2005 Graceful image retrieval performance degradation using small world distributed indexing
abstract
A small world search agent employing peer-to-peer concepts borrowed from sociology is employed for performing image retrievals in a small world distributed media index. The small world indexing method (SWIM) allows for a highly networked architecture where index information does not exist as a separate entity on a specific server, but rather is stored within the actual media objects themselves. Since each media object is only responsible for a small portion of the overall index, the loss of portions of the overall network (data objects) accounts for only a small degradation in the overall retrieval performance. Building upon previous work, the graceful degradation which is provided by the SWIM system is addressed here for retrievals which are performed using small world user agents on a large set of MPEG-7 described images.
Panagiotis Androutsos, Anastasios N. Venetsanopoulos, Dimitrios Androutsos
ICIP (3)3
2004 Distributed MPEG-7 image indexing using small world user agents
abstract
An open peer-to-peer architecture for performing distributed image indexing and retrieval is proposed. The system employs a sociological model of human acquaintance networks (small world theory) and concepts derived from the nature of the World Wide Web. Retrieval is performed using agents, and a node hopping algorithm is employed that exploits node referrals established from descriptor data stored locally by each node. A general framework for this small world image miner (SWIM) is presented along with a realization using MPEG-7 color structure descriptor data for 2400 images. Results related to search agent path length and network node degree are presented.
Panagiotis Androutsos, Azadeh Kushki, Konstantinos N. Plataniotis, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
ICIP4
1999 A Novel Vector-Based Approach to Color Image Retrieval Using a Vector Angular-Based Distance Measure
Dimitrios Androutsos, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos
Comput. Vis. Image Underst.1
1999 Adaptive fuzzy systems for multichannel signal processing
abstract
Processing multichannel signals using digital signal processing techniques has received increased attention lately due to its importance in applications such as multimedia technologies and telecommunications. The objective of this paper is twofold: 1) to introduce adaptive filtering techniques to the reader who is just beginning in this area and 2) to provide a review for the reader who may be well versed in signal processing. The perspective of the topic offered here is one that comes primarily from work done in the field of multichannel (color) image processing. Hence, many of the techniques and works cited here relate to image processing with the emphasis placed primarily on filtering algorithms based on fuzzy concepts, multidimensional scaling, and order statistics-based designs. It should be noted, however, that multichannel signal processing is a very broad field and thus contains many other approaches that have been developed from different perspectives, such as transform domain filtering, classical least-square approaches, neural networks, and stochastic methods, just to name a few. We present a general formulation based on fuzzy concepts, which allows the use of adaptive weights in the filtering structure, and we discuss different filter designs. The strong potential of fuzzy adaptive filters for multichannel signal applications, such as color image processing, is illustrated with several examples.
Konstantinos N. Plataniotis, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
Proc. IEEE2
1998 Extraction of detailed image regions for content-based image retrieval
abstract
We present a technique for coarsely extracting the regions of natural color images which contain directional detail, e.g., edges, texture, etc., which we then use for image database indexing. As a measure of color activity, we use a perceptually modified distance measure based on the sum-of-angles criterion. We then apply histogram thresholding techniques to separate the image into smooth color regions and busy regions where edge, texture and colour activity exists. Database indices are then created from the busy regions using the directional detail histogram technique and retrieval is performed using these.
Dimitrios Androutsos, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos
ICASSP1
1998 Distance Measures for Color Image Retrieval
abstract
We address the issue of image database retrieval based on color using various vector distance metrics. Our system is based on color segmentation where only a few representative color vectors are extracted from each image and used as image indices. These vectors are then used with vector distance measures to determine similarity between a query color and a database image. We test numerous popular vector distance measures in our system and find that directional measures provide the most accurate and perceptually relevant retrievals.
Dimitrios Androutsos, Konstantinos N. Plataniotis, Anastasios N. Venetsanopoulos
ICIP (2)1
1998 Efficient color image indexing and retrieval using a vector-based scheme
abstract
Color is the characteristic which is most used for image indexing and retrieval. Due to its simplicity, the color histogram remains the most commonly used method for color indexing and retrieval. However, the lack of good perceptual histogram similarity measures, the global color content of histograms and the erroneous retrieval results due to gamma nonlinearity, calls for improved methods. We implement a vector angular-based distance measure for image retrieval based on color. We build distance vectors in a multidimensional query space in which the retrieval ranking of each image is determined. Our system exhibits high flexibility by allowing all types of queries, including query by color, query by multiple colors and query by example. In addition, colors can be excluded in a query, without requiring an additional level of analysis.
Dimitrios Androutsos, Anastasios N. Venetsanopoulos, Konstantinos N. Plataniotis
MMSP1
1998 Adaptive multichannel filters for colour image processing
Konstantinos N. Plataniotis, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
Signal Process. Image Commun.2
1997 A new time series classification approach
abstract
A new approach to the problem of time series classification is discussed. A new adaptive classification scheme is introduced and compared with existing approaches, such as the Bayesian approach and the incremental credit assignment approach. Simulation results are included to demonstrate the effectiveness of the new methodology.
Konstantinos N. Plataniotis, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
ICASSP2
1997 Multichannel filters for image processing
Konstantinos N. Plataniotis, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
Signal Process. Image Commun.2
1997 Color image processing using adaptive multichannel filters
abstract
New adaptive filters for color image processing are introduced and analyzed. The proposed adaptive methodology constitutes a unifying and powerful framework for multichannel signal processing. Using the proposed methodology, color image filtering problems are treated from a global viewpoint that readily yields and unifies previous, seemingly unrelated, results. The new filters utilize Bayesian techniques and nonparametric methodologies to adapt to local data in the color image. The principles behind the new filters are explained in detail. Simulation studies indicate that the new filters are computationally attractive and have excellent performance.
Konstantinos N. Plataniotis, Dimitrios Androutsos, Sri Vinayagamoorthy, Anastasios N. Venetsanopoulos
IEEE Trans. Image Process.2
1997 Application of Active Contours for Photochromic Tracer Flow Extraction
abstract
This paper addresses the implementation of image processing and computer vision techniques to automate tracer flow extraction in images obtained by the photochromic dye technique. This task is important in modeled arterial blood flow studies. Currently, it is performed via manual application of B-spline curve fitting. However, this is a tedious and error-prone procedure and its results are nonreproducible. In the proposed approach, active contours, snakes, are employed in a new curve-fitting method for tracer flow extraction in photochromic images. An algorithm implementing snakes is introduced to automate extraction. Utilizing correlation matching, the algorithm quickly locates and localizes all flow traces in the images. The feasibility of the method for tracer flow extraction is demonstrated. Moreover, results regarding the automation algorithm are presented showing its accuracy and effectiveness. The proposed approach for tracer flow extraction has potential for real-system application.
Dimitrios Androutsos, Panos E. Trahanias, Anastasios N. Venetsanopoulos
IEEE Trans. Medical Imaging1
1996 Multichannel filtering for color image processing
abstract
This paper addresses the problem of noise attenuation for multichannel data, such as color images. The proposed filter utilizes adaptive data dependent non-parametric techniques. Simulation results indicate that the new filter suppresses impulsive as well as Gaussian noise and preserves edges and details.
Konstantinos N. Plataniotis, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
ICIP (1)2
1996 Fuzzy adaptive filters for multichannel image processing
Konstantinos N. Plataniotis, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
Signal Process.2
1996 A new time series classification approach
Konstantinos N. Plataniotis, Dimitrios Androutsos, Anastasios N. Venetsanopoulos, Demetrios G. Lainiotis
Signal Process.2
1996 An adaptive nearest neighbor multichannel filter
abstract
This paper addresses the problem of noise attenuation for multichannel data. The proposed filter utilizes adaptively determined data-dependent coefficients based on a novel distance measure which combines vector directional with vector magnitude filtering. The special case of color image processing is studied as an important example of multichannel signal processing.
Konstantinos N. Plataniotis, Sri Vinayagamoorthy, Dimitrios Androutsos, Anastasios N. Venetsanopoulos
IEEE Trans. Circuits Syst. Video Technol.3