VLDB 2026 Research / reviewers in the wild / expert
Zhiqian Wang
dblp:90/4036
· DBLP profile ↗
24ranked-venue papers
11as first author
2since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-authorArtificial intelligence and machine learning · 13 · 5 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Speech recognition and synthesis · 30% Deep learning architectures and training · 30% Image recognition and object detection · 14% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 50% Cloud and datacenter computing · 50% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 50% Audio and music processing · 50% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Computer vision › Image recognition and object detection
object recognition |
0.0 | 2 | 1998 | Pictorial Recognition of Objects Employing Affine Invariance in the Frequency Domain · IEEE Trans. Pattern Anal. Mach. Intell. 1998 Pictorial Recognition Using Affine-Invariant Spectral Signatures · CVPR 1997 |
Computer vision › Video understanding and tracking › activity recognition
human activity recognition |
0.0 | 1 | 2002 | Human Activity Recognition Using Multidimensional Indexing · IEEE Trans. Pattern Anal. Mach. Intell. 2002 |
Computer vision › 3D vision › motion estimation
3d motion estimation |
0.0 | 1 | 2000 | Estimation of 3-D motion using eigen-normalization and expansion matching · IEEE Trans. Image Process. 2000 |
Computer vision › 3D vision
motion estimation |
0.0 | 1 | 2000 | Estimation of 3-D motion using eigen-normalization and expansion matching · IEEE Trans. Image Process. 2000 |
Computer vision › Image recognition and object detection
object detection |
0.0 | 1 | 1999 | Generic Object Detection using Model Based Segmentation · CVPR 1999 |
Computer vision › 3D vision
affine invariance |
0.0 | 1 | 1998 | Pictorial Recognition of Objects Employing Affine Invariance in the Frequency Domain · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Computer vision › Face, body and person analysis › face recognition › robust face recognition
pose-invariant recognition |
0.0 | 1 | 1998 | Pictorial Recognition of Objects Employing Affine Invariance in the Frequency Domain · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Computer vision › Image recognition and object detection › object recognition › invariant object recognition
affine-invariant recognition |
0.0 | 1 | 1997 | Pictorial Recognition Using Affine-Invariant Spectral Signatures · CVPR 1997 |
Image and video processing
edge detection |
0.0 | 1 | 1996 | Optimal Ramp Edge Detection Using Expansion Matching · IEEE Trans. Pattern Anal. Mach. Intell. 1996 |
Image and video processing › feature extraction
image feature extraction |
0.0 | 1 | 1996 | Optimal Ramp Edge Detection Using Expansion Matching · IEEE Trans. Pattern Anal. Mach. Intell. 1996 |
Audio and music processing
sound source localization |
0.0 | 1 | 1996 | Conveying visual information with spatial auditory patterns · IEEE Trans. Speech Audio Process. 1996 |
Audio and music processing
spatial audio |
0.0 | 1 | 1996 | Conveying visual information with spatial auditory patterns · IEEE Trans. Speech Audio Process. 1996 |
Interaction techniques and input › non-visual interaction › auditory interaction
auditory display |
0.0 | 1 | 1996 | Conveying visual information with spatial auditory patterns · IEEE Trans. Speech Audio Process. 1996 |
Computer vision › Face, body and person analysis
human pose estimation |
0.0 | 1 | 2002 | Human Activity Recognition Using Multidimensional Indexing · IEEE Trans. Pattern Anal. Mach. Intell. 2002 |
Computer vision › Segmentation and scene understanding › image segmentation › boundary-aware segmentation
edge-based segmentation |
0.0 | 1 | 2001 | Detection and segmentation of generic shapes based on affine modeling of energy in eigenspace · IEEE Trans. Image Process. 2001 |
Computer vision › 3D vision › camera calibration
camera model |
0.0 | 1 | 2000 | Estimation of 3-D motion using eigen-normalization and expansion matching · IEEE Trans. Image Process. 2000 |
Computer vision › Segmentation and scene understanding › 3d segmentation
shape segmentation |
0.0 | 1 | 1999 | Generic Object Detection using Model Based Segmentation · CVPR 1999 |
Computer vision › Segmentation and scene understanding › image segmentation
model-based segmentation |
0.0 | 1 | 1998 | Pictorial Recognition of Objects Employing Affine Invariance in the Frequency Domain · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Computer vision › 3D vision
object representation |
0.0 | 1 | 1997 | Pictorial Recognition Using Affine-Invariant Spectral Signatures · CVPR 1997 |
Methods — techniques the papers use, named apart from their topics
batch dispatch · 0.5GPU-based inference · 0.5expansion matching · 0.0singular value decomposition · 0.0sequence-based voting · 0.0hash table · 0.0raster scanning · 0.0auditory image · 0.0principal component analysis · 0.0eigenspace energy modeling · 0.0affine parameter search · 0.0eigen-normalization · 0.0discriminative signal-to-noise ratio · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EDA-Net: Efficient deformable attention network for aero-engine damage detection and recognition
Jujian Lv, Zhiqian Wang, Ruiyin Guo, Kaihan Lin, Rongjun Chen 0001, Jia Wen Li 0001 |
Expert Syst. Appl. | 2 |
| 2026 | GMA-Net: A Lightweight Real-Time Network for Aero-Engine Damage Detection on Embedded IoT DevicesabstractBorescope inspection for aero-engine damage is a non-destructive endoscopic technique based on IoT devices and optical means, and serves as an important measure to ensure flight safety. However, borescope images of aero-engines present multiple challenges, including various damage types, large scale variations, numerous small targets. These issues pose significant obstacles to improving the accuracy of intelligent damage detection. Especially when the algorithm needs to be deployed on embedded devices for real-time processing, it is necessary to achieve coordinated optimization of detection accuracy and processing speed. To address the above problems, this paper proposes an efficient lightweight aero-engine damage detection model named GMA-Net for real-time detection on embedded devices. First, the C3Ghost module is introduced to construct the backbone network, which significantly reduces redundant computation while enhancing the model’s feature expression capability and receptive field. Second, the lightweight SegNext Attention module is embedded, which strengthens the model’s perception of tiny damage features through a combination of strip convolutions. Finally, a novel detection head structure is designed. By integrating the Detect MBConv module with lightweight depthwise separable convolution and the attention mechanism, it improves the efficiency of feature capture for multi-scale damage targets. Experiments on the self-built aero-engine damage dataset demonstrate that GMA-Net achieves promising results. Compared with the state-of-the-art lightweight detector YOLOv13n, GMA-Net reduces the number of parameters by 46.5% and lowers the computational cost (GFLOPs) by 37.1%. It also achieves a 0.86% improvement in mAP@50–95 and a 7.17% increase in recall. In addition, the model reaches 22.19 FPS on the Jetson Nano embedded device, meeting the efficiency requirements of practical intelligent aero-engine damage detection. The source code related to this paper is available at: https://github.com/Evking-silver/GMA-Net. Kaihan Lin, Zhiqian Wang, Ruiyin Guo, Jujian Lv, Xianxian Zeng, Rongjun Chen 0001, Jia Wen Li 0001 |
IEEE Internet Things J. | 2 |
| 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and MandarinabstractWe show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale. Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu |
ICML | 34 |
| 2010 | Effect of squint imaging on beam position design of space borne SARabstractRange migration of space borne SAR at large squint angle is much greater than the side-looking SAR, and longer echo receiving window is needed. Thus, the traditional beam position design method is invalid. In this paper, the method of drawing zebra map is improved by taking the range migration into consideration. The maximum and the minimum slant ranges during the synthetic time are derived. This paper also analyses the relation between the effective swath width and the range beam width at large squint angle. Simulation for X-SAR system proves that a given azimuth resolution limits the squint angle. STK and echo simulation are used to verify the validity of the improved beam position design method. Zhiqian Wang, Ze Yu 0002, Wei Yang 0004 |
IGARSS | 1 |
| 2002 | Shape Description and Invariant Recognition Employing Connectionist ApproachabstractThis paper presents a new approach for shape description and invariant recognition by geometric-normalization implemented by neural networks. The neural system consists of a shape description network, a normalization network and a recognition stage based on fuzzy pyramidal neural networks. The description network uses a novel approach for hierarchical shape segmentation and representation which expands the image shapes into localized feature tokens. These feature tokens form a compact description of the shape and its components that include information on their location, size and orientation. The description network, which is composed of a novel pyramidal architecture called the Vectorial Gradual Lattice Pyramid, processes in parallel a new vectorial scale space representation of the shape. A novel measure called Cancellation Energy is used to determine the feature tokens. The normalization network utilizes the location, size and orientation information in the feature tokens to geometric-normalize the shape or its components with respect to these parameters. The recognition network which has a pyramidal structure, uses a fuzzy representation of these normalized feature tokens to achieve robust invariant recognition. Experimental results demonstrate robust recognition in large variations of scale, rotation, translation and also in moderate affine transformations and partial occlusion. Jezekiel Ben-Arie, Zhiqian Wang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2002 | Human Activity Recognition Using Multidimensional IndexingabstractIn this paper, we develop a novel method for view-based recognition of human action/activity from videos. By observing just a few frames, we can identify the activity that takes place in a video sequence. The basic idea of our method is that activities can be positively identified from a sparsely sampled sequence of a few body poses acquired from videos. In our approach, an activity is represented by a set of pose and velocity vectors for the major body parts (hands, legs, and torso) and stored in a set of multidimensional hash tables. We develop a theoretical foundation that shows that robust recognition of a sequence of body pose vectors can be achieved by a method of indexing and sequencing and it requires only a few pose vectors (i.e., sampled body poses in video frames). We find that the probability of false alarm drops exponentially with the increased number of sampled body poses. So, matching only a few body poses guarantees high probability for correct recognition. Our approach is parallel, i.e., all possible model activities are examined at one indexing operation. In addition, our method is robust to partial occlusion since each body part is indexed separately. We use a sequence-based voting approach to recognize the activity invariant to the activity speed. Jezekiel Ben-Arie, Zhiqian Wang, Purvin Pandit, Shyamsundar Rajaram |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Hierarchical Shape Description and Similarity-Invariant Recognition Using Gradient PropagationabstractThis paper presents a novel hierarchical shape description scheme based on propagating the image gradient radially. This radial propagation is equivalent to a vectorial convolution with sector elements. The propagated gradient field collides at centers of convex/concave shape components, which can be detected as points of high directional disparity. A novel vectorial disparity measure called Cancellation Energy is used to measure this collision of the gradient field, and local maxima of this measure yield feature tokens. These feature tokens form a compact description of shapes and their components and indicate their central locations and sizes. In addition, a Gradient Signature is formed by the gradient field that collides at each center, which is itself a robust and size-independent description of the corresponding shape component. Experimental results demonstrate that the shape description is robust to distortion, noise and clutter. An important advantage of this scheme is that the feature tokens are obtained pre-attentively, without prior understanding of the image. The hierarchical description is also successfully used for similarity-invariant recognition of 2D shapes with a multi-dimensional indexing scheme based on the Gradient Signature. Jezekiel Ben-Arie, Zhiqian Wang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2001 | Detection and segmentation of generic shapes based on affine modeling of energy in eigenspaceabstractThis paper presents a novel approach for detection and segmentation of man made generic shapes in cluttered images. The set of shapes to be detected are members of affine transformed versions of basic geometric shapes such as rectangles, circles etc. The shape set is represented by its vectorial edge map transformed over a wide range of affine parameters. We use vectorial boundary instead of regular boundary to improve the robustness to noise, background clutter and partial occlusion. Our approach consists of a detection stage and a verification stage. In the detection stage, we first derive the energy from the principal eigenvectors of the set. Next, an a posteriori probability map of energy distribution is computed from the projection of the edge map representation in a vectorial eigen-space. Local peaks of the posterior probability map are located and indicate candidate detections. We use energy/probability based detection since we find that the underlying distribution is not Gaussian and resembles a hypertoroid. In the verification stage, each candidate is verified using a fast search algorithm based on a novel representation in angle space and the corresponding pose information of the detected shape is obtained. The angular representation used in the verification stage yields better results than a Euclidean distance representation. Experiments are performed in various interfering distortions, and robust detection and segmentation are achieved. Zhiqian Wang, Jezekiel Ben-Arie |
IEEE Trans. Image Process. | 1 |
| 2000 | Detection and Segmentation of Generic Shapes Based on Vectorial Affine Modeling of Energy in EigenspaceabstractThis paper presents a novel approach for detection and segmentation of generic shapes in cluttered images. We use vectorial eigenvectors to compactly represent a large set of possible appearances of primitive shapes. A posterior energy probability map of the image is calculated in the vectorial eigenspace to yield a relative similarity measure. The detection of genetic shapes is realized by detecting local peaks of the probability map. We find that eigenspace energy is more suitable for representation of sparse sets such as our affine set. At each local probability maxima, a fast search approach based on a novel representation by an angle space is employed to determine the best matching between models and the underlying sub-image. We find that angular representation in multidimensional search corresponds better to Euclidean distance than conventional projection and yields improved classification of noisy shapes. Experiments are performed in various interfering distortions, and robust detection and segmentation are achieved. Zhiqian Wang, Jezekiel Ben-Arie |
ICPR | 1 |
| 2000 | Estimation of 3-D motion using eigen-normalization and expansion matchingabstractThis correspondence describes a novel approach to three-dimensional (3-D) motion estimation of planar objects based on eigen-normalization, expansion matching (EXM), and a scaled orthographic projection model. Our approach leads to a comprehensive temporal description of all the degrees of freedom in 3-D (three rotations and three translations). Experiments with video streams show robust estimation of the real 3-D rotations and translations of the objects in motion. Jezekiel Ben-Arie, Zhiqian Wang |
IEEE Trans. Image Process. | 2 |
| 1999 | Generic Object Detection using Model Based SegmentationabstractThis paper presents a novel approach for detection and segmentation of generic shapes in cluttered images. The underlying assumption is that generic objects that are man made, frequently have surfaces which closely resemble standard model shapes such as rectangles, semi-circles etc. Due to the perspective transformations of optical imaging systems, a model shape may appear differently in the image with various orientations and aspect ratios. The set of possible appearances can be represented compactly by a few vectorial eigenbases that are derived from a small set of model shapes which are affine transformed in a wide parameter range. Instead of regular boundary of standard models, we apply a vectorial boundary which improves robustness to noise, background clutter and partial occlusion. The detection of generic shapes is realized by detecting local peaks of a similarity measure between the image edge map and an eigenspace combined set of the appearances. At each local maxima, a fast search approach based on a novel representation by an angle space is employed to determine the best matching between models and the underlying subimage. We find that angular representation in multidimensional search corresponds better to Euclidean distance than conventional projection and yields improved classification of noisy shapes. Experiments are performed in various interfering distortions, and robust detection and segmentation are achieved. Zhiqian Wang, Jezekiel Ben-Arie |
CVPR | 1 |
| 1998 | 3D Motion Estimation using Expansion Matching and KL based Canonical ImagesabstractThis paper describes a novel approach to 3D motion estimation of planar objects based on eigen-normalization, expansion matching (EXM) and a scaled orthographic projection model. Our approach leads to a comprehensive temporal description of all degrees of freedom in 3D (3 rotations and 3 translations). The 3D motion parameters of the objects are approximated by the corresponding affine parameters. The objects in each frame of a video sequence are normalized to a set of canonical images using principal component normalization procedure. The normalization approach here is based on principal components of the intensity weighted spatial values and not on the intensity values as in works such as eigenfaces. The canonical images generated differ only in orientation. Expansion matching (EXM) is then used to find the differences in orientation. Affine transformations between the shapes also are derived. The pose of the shape in 3D space can therefore be estimated. Experiments on video sequences of planar and quasi-planar objects show robust estimation of the real 3D rotations and translations of the objects in motion. Zhiqian Wang, Jezekiel Ben-Arie |
ICIP (1) | 1 |
| 1998 | Model based Segmentation and Detection of Affine Transformed Shapes in Cluttered Images
Zhiqian Wang, Jezekiel Ben-Arie |
ICIP (3) | 1 |
| 1998 | Pictorial Recognition of Objects Employing Affine Invariance in the Frequency DomainabstractDescribes an efficient approach to pose invariant pictorial object recognition employing spectral signatures of image patches that correspond to object surfaces which are roughly planar. Based on singular value decomposition (SVD), the affine transform is decomposed into slant, tilt, swing, scale, and 2D translation. Unlike previous log-polar representations which were not invariant to slant, our log-log sampling configuration in the frequency domain yields complete affine invariance. The images are preprocessed by a novel model-based segmentation scheme that detects and segments objects that are affine-similar to members of a model set of basic geometric shapes. The segmented objects are then recognized by their signatures using multidimensional indexing in a pictorial dataset represented in the frequency domain. Experimental results with a dataset of 26 models show 100 percent recognition rates in a wide range of 3D pose parameters and imaging degradations: 0-360/spl deg/ swing and tilt, 0-82/spl deg/ of slant, more than three octaves in scale change, window-limited translation, high noise levels (0 dB), and significantly reduced resolution (1:5). Jezekiel Ben-Arie, Zhiqian Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Pictorial Recognition Using Affine-Invariant Spectral SignaturesabstractThis paper describes an efficient approach to pose invariant object recognition employing pictorial recognition of image patches. A complete affine invariance is achieved by a representation which is based on a new sampling configuration in the frequency domain. Employing Singular Value Decomposition (SVD), the affine transform is decomposed into slant, tilt, swing, scale and 2D translation. From this decomposition, we derive an affine invariant representation that allows to recognize image patches that correspond to object surfaces which are roughly planar-invariant to their pose in space. The representation is in the form of Spectral Signatures that are derived from a set of Cartesian logarithmic-logarithmic (log-log) sampling configuration in the frequency domain. Unlike previous log-polar representations which are not invariant to slant (i.e. foreshortening only in one direction), our new configuration yields complete affine invariance. The proposed log-log configuration can be employed both globally or locally by a Gabor or Fourier transforms. Local representation enables to recognize separately several objects in the same image. The actual signature recognition is performed by multidimensional indexing in a pictorial dataset represented in the frequency domain. The recognition also provides 3D pose information. Jezekiel Ben-Arie, Zhiqian Wang |
CVPR | 2 |
| 1997 | SVD and log-log frequency sampling with Gabor kernels for invariant pictorial recognitionabstractThis paper presents an efficient scheme for affine-invariant object recognition. Affine invariance is obtained by a representation which is based on a new sampling configuration in the frequency domain. We discuss the decomposition of affine transform into slant, tilt, swing, scale and 2D translation by applying singular value decomposition (SVD). The affine invariant spectral signatures (AISS) are derived from a set of Cartesian logarithmic-logarithmic (log-log) sampling configuration in the frequency domain. The AISS enables the recognition of image patches that correspond to roughly planar object surfaces-regardless of their poses in space. Unlike previous log-polar representations which are not invariant to slant (i.e. foreshortening only in one direction), the AISS yields a complete affine invariance. The proposed log-log configuration can be employed either by a global Fourier transform or by a local Gabor transform. Local representation enables one to recognize separately several objects in the same image. The actual signature recognition is performed by multi-dimensional indexing in a pictorial dataset. 3D pose information is also derived as a by-product. Zhiqian Wang, Jezekiel Ben-Arie |
ICIP (3) | 1 |
| 1996 | Affine invariant shape representation and recognition using Gaussian kernels and multi-dimensional indexingabstractThis paper presents a new approach for object recognition using affine-invariant recognition of image patches that correspond to object surfaces that are roughly planar. A novel set of affine-invariant spectral signatures (AISSs) are used to recognize each surface separately invariant to its 3D pose. These local spectral signatures are extracted by convolving the image with a novel configuration of Gaussian kernels. The spectral signature of each image patch is then matched against a set of iconic models using multi-dimensional indexing (MDI) in the frequency domain. Affine-invariance of the signatures is achieved by a new configuration of Gaussian kernels with modulation in two orthogonal axes. The proposed configuration of kernels is Cartesian with varying aspect ratios in two orthogonal directions. The kernels are organized in subsets where each subset has a distinct orientation. Each subset spans the entire frequency domain and provides invariance to slant, scale and limited translation. The complete set of orientations is utilized to achieve invariance to rotation and tilt. Hence, the proposed set of kernels achieve complete affine-invariance. Jezekiel Ben-Arie, Zhiqian Wang, K. Raghunath Rao |
ICASSP | 2 |
| 1996 | On the use of the Karhunen-Loeve transform and expansion matching for generalized feature detectionabstractA novel generalized feature extraction method based on the expansion matching (EXM) method and the Karhunen-Loeve (KL) transform is presented. This yields an efficient method to locate a large variety of features with a single pass of parallel filtering operations. The EXM method is used to design optimal detectors for different features. The KL representation is used to define an optimal basis for representing these EXM feature detectors with minimum truncation error. Input images are then analyzed with the resulting KL bases. The KL coefficients obtained from the analysis are used to efficiently reconstruct the response due to any combination of feature detectors. The method is successfully applied to real images and extracts a variety of arc and edge features as well as more complex junction features formed by combining two or more arcs or line features. Dibyendu Nandy, Jezekiel Ben-Arie, Nebojsa Jojic, Zhiqian Wang, K. Raghunath Rao |
ICASSP | 4 |
| 1996 | Iconic recognition with affine-invariant spectral signaturesabstractThis paper presents a new approach for object recognition using affine-invariant recognition of image patches that correspond to object surfaces that are roughly planar. A novel set of affine-invariant spectral signatures (AISSs) are used to recognize each surface separately invariant to its 3D pose. These local spectral signatures are extracted by correlating the image with a novel configuration of Gaussian kernels. The spectral signature of each image patch is then matched against a set of iconic models using multidimensional indexing (MDI) in the frequency domain. Affine-invariance of the signatures is achieved by a new configuration of Gaussian kernels with modulation in two orthogonal axes. The proposed configuration of kernels is Cartesian with varying aspect ratios in two orthogonal directions. The kernels are organized in subsets where each subset has a distinct orientation. Each subset spans the entire frequency domain and provides invariance to slant, scale and limited translation. The complete set of orientations is utilized to achieve invariance to rotation and tilt. Hence, the proposed set of kernels achieve complete affine-invariance. Jezekiel Ben-Arie, Zhiqian Wang, K. Raghunath Rao |
ICPR | 2 |
| 1996 | A generalized expansion matching based feature extractorabstractA novel and efficient generalized feature extraction method is presented based on the expansion matching (EXM) method and the Karhunen-Loueve (KL) transform. The EXM method is used to design optimal detectors for different features. The KL representation is used to define an optimal basis for representing these EXM feature detectors with minimum truncation error. Input images are then analyzed with the resulting KL basis set. The KL coefficients obtained from the analysis are used to efficiently reconstruct the response due to any combination of feature detectors. The method is applied to real images and successfully extracts a variety of arc and edge features as well as complex junction features formed by combining two or more arc or line features. Zhiqian Wang, K. Raghunath Rao, Dibyendu Nandy, Jezekiel Ben-Arie, Nebojsa Jojic |
ICPR | 1 |
| 1996 | Optimal Ramp Edge Detection Using Expansion MatchingabstractIn practical images, ideal step edges are actually transformed into ramp edges, due to the general low pass filtering nature of imaging systems. This paper discusses the application of the expansion matching (EXM) method for optimal ramp edge detection. EXM optimizes a novel matching criterion called discriminative signal-to-noise ratio (DSNR) and has been shown to robustly recognize templates under conditions of noise, severe occlusion, and superposition. We show that our ramp edge detector performs better than the ramp detector obtained from Canny's criteria in terms of DSNR and is relatively easier to derive for various noise levels and slopes. Zhiqian Wang, K. Raghunath Rao, Jezekiel Ben-Arie |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | Conveying visual information with spatial auditory patternsabstractHumans can perceive their spatial surroundings both by the visual and the auditory senses. Even though the auditory system has much lower spatial resolution, it can still serve as an alternative or supplementary information channel for the visual modality. We investigate the feasibility of conveying elaborate visual information by auditory patterns. A system is developed that transforms 2-D binary images into "auditory images." Such images are based on slow raster scanning of a perceptual auditory surface around the subject while emitting sounds that correspond to the brightness level of the images at each location. The parameters involved in such a system are quantitatively investigated with respect to their effect on performance. These parameters include resolution, level difference, sound color contrast, speed, surface curvature, etc. The graphic percepts of synthetic auditory images that correspond to simple shapes are also analyzed. The experimental results show that sound localization can be used to convey visual information quite successfully for simple shapes and low-resolution patterns. Zhiqian Wang, Jezekiel Ben-Arie |
IEEE Trans. Speech Audio Process. | 1 |
| 1995 | Robust shape description and recognition by gradient propagationabstractThis paper presents a novel hierarchical shape description scheme based on propagating the gradient of the image. The propagated gradient field collides at centers of convex/concave shape components, which can be detected as points of high directional disparity. A novel vectorial disparity measure called cancelation energy is used to measure this collision of the gradient field, and local maxima of this measure yield feature tokens. These feature tokens form a compact description of shapes and their components and indicate their central location and size. In addition, a gradient signature is formed by the gradient field that collides at each center, which is itself a robust and size-independent description of the corresponding shape component. Experimental results demonstrate that the shape description is robust to distortion, noise and clutter. An important advantage of this scheme is that the feature tokens are obtained pre-attentively, without prior understanding of the image. The hierarchical description is also successfully used for similarity-invariant recognition of 2D shapes with a multi-dimensional indexing scheme based on the gradient signature. Jezekiel Ben-Arie, K. Raghunath Rao, Zhiqian Wang |
ICIP (3) | 3 |
| 1995 | Optimal DSNR detector for ramp edgesabstractIn practical images, ideal step edges are actually transformed into exponential ramp edges, due to the general low pass filtering nature of imaging systems. This paper discusses the application of a newly developed expansion matching method for optimal ramp edge detection. Expansion matching optimizes a novel matching criterion called discriminative signal to noise ratio (DSNR). The DSNR criterion represents the desirable qualities of a sharp matching response with good localization and minimal off-center response. These requirements are consistent with the three criteria of signal-to-noise ratio, localization, and multiple response suppression used by Canny (1986) and others for optimal edge detection. We compare the optimal ramp edge detector based on DSNR with the ramp edge detector derived from Canny's criteria. We show that our ramp edge detector performs better than the ramp detector obtained from Canny's criteria in terms of DSNR and is relatively easier to derive for various amounts of noise and slopes. Zhiqian Wang, K. Raghunath Rao, Jezekiel Ben-Arie |
ICIP | 1 |