EDBT 2026 Demo / reviewers in the wild / expert
Murtaza Taj
dblp:02/1709
· DBLP profile ↗
33ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-2353-4462ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal AI in healthcare: Review of vision-language foundation models for real-world medical applications
Taha Razzaq, Murtaza Taj, Asim Iqbal |
J. Biomed. Informatics | 2 |
| 2025 | Localization Lens for Improving Medical Vision-Language Models
Hasan Farooq, Murtaza Taj, Mehwish Nasim, Arif Mahmood |
MICCAI (9) | 2 |
| 2025 | CATVis: Context-Aware Thought Visualization
Tariq Mehmood, Hamza Ahmad, Muhammad Haroon Shakeel, Murtaza Taj |
MICCAI (1) | 4 |
| 2024 | Camera Calibration Through Geometric Constraints from Rotation and Projection MatricesabstractThe process of camera calibration involves estimating the intrinsic and extrinsic parameters, which are essential for accurately performing tasks such as 3D reconstruction, object tracking and augmented reality. In this work, we propose a novel constraints-based loss for measuring the intrinsic (focal length: $\left(f_{x}, f_{y}\right)$ and principal point: $\left.\left(p_{x}, p_{y}\right)\right)$ and extrinsic (baseline: (b), disparity: (d), translation: $\left(t_{x}, t_{y}, t_{z}\right)$, and rotation specifically pitch: $\left(\theta_{p}\right)$) camera parameters. Our novel constraints are based on geometric properties inherent in the camera model, including the anatomy of the projection matrix (vanishing points, image of world origin, axis planes) and the orthonormality of the rotation matrix. Thus we proposed a novel Unsupervised Geometric Constraint Loss (UGCL) via a multitask learning framework. Our methodology is a hybrid approach that employs the learning power of a neural network to estimate the desired parameters along with the underlying mathematical properties inherent in the camera projection matrix. This distinctive approach not only enhances the interpretability of the model but also facilitates a more informed learning process. Additionally, we introduce a new CVGL Camera Calibration dataset, featuring over 900 configurations of camera parameters, incorporating 63, 600 image pairs that closely mirror real-world conditions. By training and testing on both synthetic and real-world datasets, our proposed approach demonstrates improvements across all parameters when compared to the state-of-the-art (SOTA) benchmarks. The code and the updated dataset can be found here: https://github.com/CVLABLUMS/CVGL-Camera-Calibration. Muhammad Waleed, Murtaza Taj |
ICIP | 3 |
| 2024 | Mending of Spatio-Temporal Dependencies in Block Adjacency Matrix
Osama Ahmad, Omer Abdul Jalil, Usman Nazir, Murtaza Taj |
ICONIP (7) | 4 |
| 2024 | Stereollax Net: Stereo Parallax-Based Deep Learning Network for Building Height EstimationabstractAccurate estimation of building heights is crucial for effective urban planning and resource management as it provides essential geometric information about the urban landscape. Many end-to-end deep learning-based networks have been proposed for image-to-height mapping using high-resolution non-optical and optical remote sensing imagery. In this study, we develop a novel deep-learning architecture that incorporates a stereo parallax-based mathematical formulation for building height estimation. We estimate stereo formulation parameters include differential parallax (ΔP) image, average photo-base (b), and satellite height (hs). The final height map is computed by utilizing these parameters in the stereo parallax equation, thus combining closed-form solutions within the learning paradigm. Moreover, to improve the estimation of ΔP, we also introduce a multi-scale differential shortcut connections module (MSDSC). The MSDSC module integrates high-frequency components into lower-resolution baseline decoder features while converting them into high-resolution decoder features. To establish the efficacy of our proposed stereo parallax-based deep learning network (Stereollax Net), we train and evaluate our method on densely populated cities of China (42-Cities dataset) and on the IEEE Data Fusion Contest 2018 dataset (DFC2018). Our proposed Stereollax Net is trained only with RGB imagery and compared with state-of-the-art methods that utilize both panchromatic and multi-spectral (RGB and Near-infrared) satellite imagery. The qualitative and quantitative results demonstrate that our Stereollax Net surpasses existing state-of-the-art (SOTA) algorithms, achieving superior performance with fewer data and training parameters by a considerable margin. The code will be made publicly available via the GitHub repository. Sana Jabbar, Murtaza Taj |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Stereoential Net: Deep Network for Learning Building Height Using Stereo Imagery
Sana Jabbar, Murtaza Taj |
ICONIP (13) | 2 |
| 2022 | Camera Calibration Through Camera Projection LossabstractCamera calibration is a necessity in various tasks including 3D reconstruction, hand-eye coordination for a robotic interaction, autonomous driving, etc. In this work we propose a novel method to predict extrinsic (baseline, pitch, and translation), intrinsic (focal length and principal point offset) parameters using an image pair. Unlike existing methods, instead of designing an end-to-end solution, we proposed a new representation that incorporates camera model equations as a neural network in a multi-task learning framework. We estimate the desired parameters via novel camera projection loss (CPL) that uses the camera model neural network to reconstruct the 3D points and uses the reconstruction loss to estimate the camera parameters. To the best of our knowledge, ours is the first method to jointly estimate both the intrinsic and extrinsic parameters via a multi-task learning methodology that combines analytical equations in learning framework for the estimation of camera parameters. We also proposed a novel CVGL Camera Calibration dataset using CARLA Simulator [1]. Empirically, we demonstrate that our proposed approach achieves better performance with respect to both deep learning-based and traditional methods on 8 out of 10 parameters evaluated using both synthetic and real data. Our code and generated dataset are available at https://github.com/thanif/Camera-Calibration-through-Camera-Projection-Loss. Talha Hanif Butt, Murtaza Taj |
ICASSP | 2 |
| 2022 | Neural Network Pruning Through Constrained Reinforcement LearningabstractNetwork pruning reduces the size of neural networks by removing neurons such that the performance drop is minimal. Traditional pruning approaches focus on designing metrics to quantify the usefulness of a neuron which is often tedious and sub-optimal. More recent methods have instead focused on training auxiliary networks to automatically learn how useful each neuron is however, they often do not take computational limitations into account. We propose a general methodology for pruning neural networks. Our approach can prune neural networks to respect pre-defined computational budgets on arbitrary, possibly non-differentiable, functions. Furthermore, we only assume the ability to be able to evaluate these functions for different inputs, and hence they do not need to be fully specified beforehand. We achieve this by proposing a novel pruning strategy via constrained reinforcement learning algorithms. We prove the effectiveness of our approach via comparison with state-of-the-art methods on standard image classification datasets. Specifically, we reduce 83−92.90% of total parameters on various variants of VGG while achieving comparable or better performance. We also achieved 75.09% reduction in parameters on ResNet18 without incurring any loss in accuracy. The code is also available at our github repository1. Shehryar Malik, Muhammad Umair Haider, Omer Iqbal, Murtaza Taj |
ICPR | 4 |
| 2021 | Teacher-Class Network: A Neural Network Compression Mechanism
Shaiq Munir Malik, Fnu Mohbat, Muhammad Umair Haider, Muhammad Musab Rasheed, Murtaza Taj |
BMVC | 5 |
| 2021 | Comprehensive Online Network Pruning Via Learnable Scaling FactorsabstractOne of the major challenges in deploying deep neural network architectures is their size which has an adverse effect on their inference time and memory requirements. Deep CNNs can either be pruned width-wise by removing filters or depth-wise by removing layers and blocks. Width wise pruning (filter pruning) is commonly performed via learnable gates or switches and sparsity regularizers whereas pruning of layers has so far been performed arbitrarily by manually designing a smaller network usually referred to as a student network. We propose a comprehensive pruning strategy that can perform both width-wise as well as depth-wise pruning. This is achieved by introducing gates at different granularities (neuron, filter, layer, block) which are then controlled via an objective function that simultaneously performs pruning at different granularity during each forward pass. Our approach is applicable to wide-variety of architectures without any constraints on spatial dimensions or connection type (sequential, residual, parallel or inception). Our method has resulted in a compression ratio of 70% to 90% without noticeable loss in accuracy when evaluated on benchmark datasets. Muhammad Umair Haider, Murtaza Taj |
ICIP | 2 |
| 2021 | Spatio-Temporal Crop Classification On Volumetric DataabstractLarge-area crop classification using multi-spectral imagery is a widely studied problem for several decades and is generally addressed using classical Random Forest classifier. Recently, deep convolutional neural networks (DCNN) have been proposed. However, these methods only achieved results comparable with Random Forest. In this work, we present a novel CNN based architecture for large-area crop classification. Our methodology combines both spatio-temporal analysis via 3D CNN as well as temporal analysis via 1D CNN. We evaluated the efficacy of our approach on Yolo and Imperial county benchmark datasets. Our combined strategy outperforms both classical as well as recent DCNN based methods in terms of classification accuracy by 2% while maintaining a minimum number of parameters and the lowest inference time. Muhammad Usman Qadeer, Salar Saeed, Murtaza Taj, Abubakr Muhammad |
ICIP | 3 |
| 2021 | Trash Detection on Water Channels
Mohbat Tharani, Abdul Wahab Amin, Fezan Rasool, Mohammad Maaz, Murtaza Taj, Abubakar Muhammad |
ICONIP (1) | 5 |
| 2021 | Statistically correlated multi-task learning for autonomous driving
Waseem Abbas 0002, Murtaza Taj, Arif Mahmood |
Neural Comput. Appl. | 3 |
| 2020 | A Residual-Dyad Encoder Discriminator Network for Remote Sensing Image MatchingabstractWe propose a new method for remote sensing image matching. The proposed method uses an encoder subnetwork of an autoencoder pretrained on the GTCrossView data to construct image features. A discriminator network trained on the University of California Merced land-use/land-cover data set (LandUse) and the high-resolution satellite scene data set (SatScene) computes a match score between a pair of computed image features. We also propose a new network unit, called residual-dyad, and empirically demonstrate that networks that use residual-dyad units outperform those that do not. We compare our approach with both traditional and more recent learning-based schemes on the LandUse and SatScene data sets, and the proposed method achieves the state-of-the-art result in terms of mean average precision and average normalized modified retrieval rank (ANMRR) metrics. Specifically, our method achieves an overall improvement in performance of 11.26% and 22.41%, respectively, for LandUse and SatScene benchmark data sets. Numan Khurshid, Mohbat Tharani, Murtaza Taj, Faisal Z. Qureshi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Adaptively Weighted Multi-task Learning Using Inverse Validation LossabstractMulti-task learning aims to enhance the performance of a model by inductive transfer of information among tasks. However, joint optimization of multiple tasks is challenging due to unbalanced data ranges and variations in task difficulties which can cause the model to converge only for a single task which has large values. To address these problems, we propose a novel weighting scheme based on validation loss. The proposed weighted scheme is evaluated on three datasets, including publicly available Comma.ai and Udacity benchmark dataset and GTA-V dataset. Our experiments demonstrate the superior performance of the proposed approach compared to the current state-of-the-art methods. Waseem Abbas 0002, Murtaza Taj |
ICASSP | 2 |
| 2019 | Using 3D Residual Network for Spatio-temporal Analysis of Remote Sensing DataabstractIn this paper, we propose an approach to recognize spatio-temporal changes from remote sensing data. Instead of performing independent analysis on each instance of satellite imagery, we proposed a 3D Convolutional Neural Network (CNN) based on the ResNet architecture. Our approach takes as input a 3D spatio-temporal block comprising of spatial as well as temporal data from multiple years. We predict four key transition classes namely construction, destruction, cultivation and decultivation. In our proposed architecture, we introduced Leaky ReLU instead of ReLU which improves the overall performance as it solves the dying ReLU problem. We also provided dataset and annotations1for these four classes and have evaluated the efficacy of our approach on data from three different cities. Muhammad Ahmed Bhimra, Usman Nazir, Murtaza Taj |
ICASSP | 3 |
| 2019 | Point Cloud Segmentation Using Hierarchical Tree for Architectural ModelsabstractOver the past few years, gathering massive volume of 3D data has become straightforward due to the proliferation of laser scanners and acquisition devices. Segmentation of such large data into meaningful segments, however, remains a challenge. Raw scans usually have missing data and varying density. In this work, we present a simple yet effective method to semantically decompose and reconstruct 3D models from point clouds. Using a hierarchical tree approach, we segment and reconstruct planar as well as non-planar scenes in an outdoor environment. This tree uses an exclusive energy function and a 3D convolutional neural network, HollowNets, to classify the segments. We test the efficacy of our proposed approach on a variety of complex real and synthetic data samples, obtaining an improvement of 7.9% in mean IOU over the state of the art approaches. Omair Hassaan, Abeera Shamail, Zain Butt, Murtaza Taj |
ICASSP | 4 |
| 2019 | Patch-Based Generative Adversarial Network Towards Retinal Vessel Segmentation
Waseem Abbas 0002, Muhammad Haroon Shakeel, Numan Khurshid, Murtaza Taj |
ICONIP (4) | 4 |
| 2019 | Cross-View Image Retrieval - Ground to Aerial Image Retrieval Through Deep Learning
Numan Khurshid, Talha Hanif Butt, Mohbat Tharani, Murtaza Taj |
ICONIP (2) | 4 |
| 2019 | Accurate Localization Algorithm in Wireless Sensor Networks in the Presence of Cross Technology Interference
Usman Nazir, Ijaz Haider Naqvi, Murtaza Taj |
ICONIP (5) | 3 |
| 2017 | More for less: Insights into convolutional nets for 3D point cloud recognitionabstractWith the recent breakthrough in commodity 3D imaging solutions such as depth sensing, photogrammetry, stereoscopic vision and structured light, 3D shape recognition is becoming an increasingly important problem. A longstanding question is what should be the format of the 3D shape (such as voxel, mesh, point-cloud etc.) and what could be a good generic feature representation for shape recognition. This question is particularly important in the context of convolutional neural network (CNN) whose efficacy and complexity depends upon the choice of input shape format and the design of network. It has been seen that both 3D voxel representation as well as collection of rendered views on 2D images have produced competing results. Similarly, it have been seen that networks with few million parameters and networks with several hundred million parameters have similar performance. In this work we compare these solutions and provide an analysis on the factors resulting in increase in the parameters without significantly improving accuracy. On the basis of the above analysis we propose a representation method (point cloud to 2D grid) and architecture that results in much less parameters for the CNN but has competing accuracy. Usama Shafiq, Murtaza Taj, Mohsen Ali |
ICIP | 2 |
| 2014 | 2D articulated human pose tracking: A hybrid approachabstractIn tracking, there are two fundamental ways to solve the correspondence problem, either as a low-level feature matching or through high-level object matching. Most of the 2D pose tracking methods are based on high-level object matching. This makes them highly dependent on the object detectors, which are typically trained in specific views, limiting pose trackers to those view-points only. We propose a systematic approach for 2D pose tracking that combines low-level feature matching and high-level object matching approaches in a unified framework. We utilized brightness constancy assumption to find the corresponding pixels in two consecutive frames. We combine this tracking with frontal and profile pose detectors through a decoding and fusion strategy, to enable continuous pose estimation and tracking over wide range of view-points. The added advantage of our approach is, we not only track each limb, we can also track an articulated joint between them without requiring any 3D estimate of the skeleton. In addition to being computationally efficient, this hybrid tracking framework generalizes to unseen pose variations and compares favorably with existing work. Murtaza Taj |
ICIP | 2 |
| 2014 | Efficient 2D human pose estimation using mean-shiftabstractIn 2D pose estimation, each limb is parametrized by it position(2D), scale(1D) and orientation(1D). One of the key bottlenecks is the exhaustive search in this 4D limb space where only a few maxima in the space are desired. To reduce the search space, we reformulate this problem in terms of finding the modes of a likelihood distribution and solve it using the Mean-Shift algorithm. Ours is the first paper in the pose estimation community to use such an approach. In addition, we describe a complete top-down approach that estimates limbs in a sequential pair-wise manner. This allows us to use Kinematic Constraints before processing, requiring us to perform search in only a small sub-region of the image for each limb. We finally devise a PCA based pose validation criteria that enables us to prune invalid hypotheses. Combining these search-space reduction techniques allows our method to generate results at par with the state-of-the-art, while saving more than 80% computations when compared to full image search. Abdul Rafay Khalid, Murtaza Taj |
ICIP | 3 |
| 2012 | Interaction recognition in wide areas using audiovisual sensorsabstractWe present an event recognition framework to detect interactions among objects, for example people, using a network of cameras and associated microphone pairs. The complementarity of the video and audio modalities is exploited to cover wide areas. In particular, object movements in portions of the scene that are not covered by the cameras' fields of view are estimated using the input from microphones. After estimating trajectories using audio-visual features, we recognize interactions based on a Coupled Hidden Markov Model Maximum a Posteriori (CHMM-MAP) approach. The states of the CHMM are initialized via Gaussian Mixture Model (GMM) clustering on a multi-dimensional feature space. Evaluation and comparison with three alternative methods demonstrate the effectiveness of the proposed CHMM-MAP trained on multiple features on both synthetic and real data. Murtaza Taj, Andrea Cavallaro |
ICIP | 1 |
| 2010 | Content and task-based view selection from multiple video streams
Fahad Daniyal, Murtaza Taj, Andrea Cavallaro |
Multim. Tools Appl. | 2 |
| 2009 | Audio-assisted trajectory estimation in non-overlapping multi-camera networksabstractWe present an algorithm to improve trajectory estimation in networks of non-overlapping cameras using audio measurements. The algorithm fuses audiovisual cues in each camera's field of view and recovers trajectories in unobserved regions using microphones only. Audio source localization is performed using stereo audio and cycloptic vision (STAC) sensor by estimating the time difference of arrival (TDOA) between microphone pair and then by computing the cross correlation. Audio estimates are then smoothed using Kalman filtering. The audio-visual fusion is performed using a dynamic weighting strategy. We show that using a multi-modal sensor with combined visual (narrow) and audio (wider) field of view can enable extended target tracking in non-overlapping camera settings. In particular, the weighting scheme improves performance in the overlapping regions. The algorithm is evaluated in several multi-sensor configurations using synthetic data and compared with state of the art algorithm. Murtaza Taj, Andrea Cavallaro |
ICASSP | 1 |
| 2008 | Object and Scene-Centric Activity Detection Using State Occupancy Duration ModelingabstractWe propose a video event analysis framework based on object segmentation and tracking, combined with a Hidden Semi-Markov Model (HSMM) that uses state occupancy duration modeling. The observations generated by a multi-object detector and tracker are used as emitting symbols and the corresponding probabilities are computed using multivariate Gaussians. Next, we recognize events by estimating the most likely object state sequence using a HSMM decoding strategy, based on the Viterbi algorithm. Moreover,the duration distribution enforces the state transition after certain time and hence better models the events constrained on time intervals. We demonstrate and evaluate the proposed framework on a dataset of approximately 20 K frames, and show that the duration modeling improves the event detection results by 7% to 11%, compared to state-of-the-art HMMs. Murtaza Taj, Andrea Cavallaro |
AVSS | 1 |
| 2008 | Efficient Multitarget Visual Tracking Using Random Finite SetsabstractWe propose a filtering framework for multitarget tracking that is based on the probability hypothesis density (PHD) filter and data association using graph matching. This framework can be combined with any object detectors that generate positional and dimensional information of objects of interest. The PHD filter compensates for missing detections and removes noise and clutter. Moreover, this filter reduces the growth in complexity with the number of targets from exponential to linear by propagating the first-order moment of the multitarget posterior, instead of the full posterior. In order to account for the nature of the PHD propagation, we propose a novel particle resampling strategy and we adapt dynamic and observation models to cope with varying object scales. The proposed resampling strategy allows us to use the PHD filter when a priori knowledge of the scene is not available. Moreover, the dynamic and observation models are not limited to the PHD filter and can be applied to any Bayesian tracker that can handle state-dependent variances. Extensive experimental results on a large video surveillance dataset using a standard evaluation protocol show that the proposed filtering framework improves the accuracy of the tracker, especially in cluttered scenes. Emilio Maggio, Murtaza Taj, Andrea Cavallaro |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Relative Position Estimation of Non-Overlapping CamerasabstractWe present an algorithm for the estimation of the relative camera position in a network of cameras with non-overlapping fields of view. The algorithm estimates the missing trajectory information in the unobserved areas of the multi-sensor configuration using both parametric and non-parametric algorithms. First, Kalman filtering is used to estimate the trajectories in the unobserved regions. Next, linear regression estimates the position of the target based upon the motion model generated from the measured positions in the field of view of each sensor. Finally, the relative orientation of the sensors is calculated using the observed and estimated target position from adjacent cameras. We demonstrate the algorithm on both synthetic and real data. Nadeem Anjum, Murtaza Taj, Andrea Cavallaro |
ICASSP (2) | 2 |
| 2007 | Hands-On Experience in Image Processing: The Automated Lecture CameramanabstractWe present a system for live recorded lecture-based distance learning delivery that was designed and implemented based on the framework defined in A. Cavallaro et al, 2005. The system is built by students based on previous students' projects and is deployed in a real distance learning scenario. The video capturing process of a lecture is automated using a robotic camera that tracks the movements of a lecturer during the delivery of a traditional class. The robotic camera is guided by the results of an image processing module based on face detection. The video of the lecturer is synchronized with the presentation slides and with the audio of the lecture. The system was evaluated based on students' feedback. Andrea Cavallaro, Ruchira Chandrasekera, Murtaza Taj |
ICASSP (3) | 3 |
| 2007 | Multi-Modal Particle Filtering Tracking using Appearance, Motion and Audio LikelihoodsabstractWe propose a multi-modal object tracking algorithm that combines appearance, motion and audio information in a particle filter. The proposed tracker fuses at the likelihood level the audio-visual observations captured with a video camera coupled with two microphones. Two video likelihoods are computed that are based on a 3D color histogram appearance model and on a color change detection, whereas an audio likelihood provides information about the direction of arrival of a target. The direction of arrival is computed based on a multi-band generalized cross-correlation function enhanced with a noise suppression and reverberation filtering that uses the precedence effect. We evaluate the tracker on single and multi-modality tracking and quantify the performance improvement introduced by integrating audio and visual information in the tracking process. Matteo Bregonzio, Murtaza Taj, Andrea Cavallaro |
ICIP (5) | 2 |
| 2007 | Multi-Camera Scene Analysis using an Object-Centric Continuous Distribution Hidden Markov ModelabstractWe propose a multi-camera event detection framework that can operate on a common ground plane as well as on the image plane. The proposed event detector is based on an object-centric state modeling that uses a continuous distribution hidden Markov model (CDHMM). Video objects are first detected using statistical change detection and then tracked using graph matching. Next, the algorithm recognizes events by estimating the most likely object state sequence using a HMM decoding strategy, based on the Viterbi algorithm. We demonstrate and evaluate the proposed framework on standard event detection datasets with single and multiple cameras, with both overlapping and non-overlapping fields of view. Murtaza Taj, Andrea Cavallaro |
ICIP (4) | 1 |