Muhammad Wajahat Hussain

dblp:149/1429 · also Wajahat Hussain · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-2899-7493ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Robot navigation and mapping · 70% 3D vision · 17% Motion planning and robot control · 14%
Computer graphics and multimedia
1 paper
Audio and music processing · 77% Computational photography and imaging · 23%

Topics — the 6 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
localization
0.912025
Help Me Through: Imitation Learning Based Active View Planning to Avoid SLAM Tracking Failures · IEEE Trans. Robotics 2025
Robotics › Robot navigation and mapping › view planning
next-best-view planning
0.912025
Help Me Through: Imitation Learning Based Active View Planning to Avoid SLAM Tracking Failures · IEEE Trans. Robotics 2025
Robotics › Robot navigation and mapping › SLAM
visual SLAM
0.912025
Help Me Through: Imitation Learning Based Active View Planning to Avoid SLAM Tracking Failures · IEEE Trans. Robotics 2025
Robotics › Motion planning and robot control
robot control
0.312025
Help Me Through: Imitation Learning Based Active View Planning to Avoid SLAM Tracking Failures · IEEE Trans. Robotics 2025
Robotics › Motion planning and robot control
robot learning
0.312025
Help Me Through: Imitation Learning Based Active View Planning to Avoid SLAM Tracking Failures · IEEE Trans. Robotics 2025
Computer vision › 3D vision › 3d scene understanding › monocular 3d perception
single image 3d understanding
0.212015
Single Image 3D without a Single 3D Image · ICCV 2015

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.9online learning · 0.9imitation learning · 0.9audio-visual descriptor · 0.4acoustic echo modeling · 0.4unsupervised learning · 0.2
YearPublicationVenuePosition
2025 Low Cost 3D Motion Capture of Rapid Maneuvers Using a Single High Speed Camera
Muhammad Wajahat Hussain, Mahum Naveed, Taimoor Hasan Khan, Muhammad Latif Anjum, Shahzad Rasool, Adnan Maqsood
Comput. Animat. Virtual Worlds1
2025 Help Me Through: Imitation Learning Based Active View Planning to Avoid SLAM Tracking Failures
abstract
Large-scale evaluation of state-of-the-art visual SLAM has shown that its tracking performance degrades considerably if the camera view is not adjusted to avoid the low-texture areas. Deep reinforcement learning based approaches have been proposed to improve the robustness of visual tracking in such unsupervised settings. Our extensive analysis reveals the fundamental limitations of reinforcement learning based active view planning, especially in transition scenarios (entering/exiting the room, texture-less walls, and lobbies). In challenging transition scenarios, the agent generally remains unable to cross the transition during training, limiting its ability to learn the maneuver. We propose human-supervised RL training (imitation learning) and achieve significantly improved performance after$\sim$50 hours of supervised training. To reduce longer human supervision requirements, we also explore fine-tuning our network with an online learning policy. Here, we use limited human-supervised training ($\sim$20 hours), and fine-tune the network with unsupervised training ($\sim$45 hours), obtaining encouraging results. We also release our multi-model, human supervised training dataset. The dataset contains challenging and diverse transition scenarios and can aid the development of imitation learning policies for consistent visual tracking. We also release our implementation11https://tinyurl.com/478maz6u.
Kanwal Naveed, Muhammad Wajahat Hussain, Irfan Hussain, Muhammad Latif Anjum
IEEE Trans. Robotics2
2024 Deeper Introspective SLAM: How to Avoid Tracking Failures Over Longer Routes?
abstract
Large scale active exploration has recently revealed limitations of visual SLAM’s tracking ability. Active view planning methods based on reinforcement learning have been proposed to improve visual tracking robustness.In this work, we expose the limitations of deep reinforcement learning-based visual SLAM over longer routes. We demonstrate that additional modalities (depth, scene layout) offer little improvement. Furthermore, reward shaping is not the main reason behind the shortsightedness of the state-of-the-art visual SLAM tracker. We propose a novel video vision transformer-based architecture that improves the farsightedness of the visual tracker, which results in the completion of longer routes with efficient paths.Out of 60 challenging routes, our approach manages to complete 56 routes, which is a three-fold improvement over the state-of-the-art active view mapping (DI-SLAM) baseline. Interestingly, ORB-SLAM3 was unable to complete a single route without tracking failure. Our code is available at https://tinyurl.com/w935spuz.
Kanwal Naveed, Muhammad Latif Anjum, Muhammad Wajahat Hussain
IROS3
2024 Targeted adversarial attack on classic vision pipelines
Kainat Riaz, Muhammad Latif Anjum, Muhammad Wajahat Hussain, Rohan Manzoor
Comput. Vis. Image Underst.3
2022 On Smart Gaze Based Annotation of Histopathology Images for Training of Deep Convolutional Neural Networks
abstract
Unavailability of large training datasets is a bottleneck that needs to be overcome to realize the true potential of deep learning in histopathology applications. Although slide digitization via whole slide imaging scanners has increased the speed of data acquisition, labeling of virtual slides requires a substantial time investment from pathologists. Eye gaze annotations have the potential to speed up the slide labeling process. This work explores the viability and timing comparisons of eye gaze labeling compared to conventional manual labeling for training object detectors. Challenges associated with gaze based labeling and methods to refine the coarse data annotations for subsequent object detection are also discussed. Results demonstrate that gaze tracking based labeling can save valuable pathologist time and delivers good performance when employed for training a deep object detector. Using the task of localization of Keratin Pearls in cases of oral squamous cell carcinoma as a test case, we compare the performance gap between deep object detectors trained using hand-labelled and gaze-labelled data. On average, compared to 'Bounding-box' based hand-labeling, gaze-labeling required 57.6% less time per label and compared to 'Freehand' labeling, gaze-labeling required on average 85% less time per label.
Komal Mariam, Osama Mohammed Afzal, Muhammad Wajahat Hussain, Muhammad Umar Javed, Amber Kiyani, Nasir M. Rajpoot, Syed Ali Khurram, Hassan Aqeel Khan
IEEE J. Biomed. Health Informatics3
2021 Grey is the new RGB: How good is GAN-based image colorization for image compression?
Aroosh Fatima, Muhammad Wajahat Hussain, Shahzad Rasool
Multim. Tools Appl.2
2021 A First Look at Private Communications in Video Games using Visual Features
abstract
Abstract Internet privacy is threatened by expanding use of automated mass surveillance and censorship techniques. In this paper, we investigate the feasibility of using video games and virtual environments to evade automated detection, namely by manipulating elements in the game environment to compose and share text with other users. This technique exploits the fact that text spotting in the wild is a challenging problem in computer vision. To test our hypothesis, we compile a novel dataset of text generated in popular video games and analyze it using state-of-the-art text spotting tools. Detection rates are negligible in most cases. Retraining these classifiers specifically for game environments leads to dramatic improvements in some cases (ranging from 6% to 65% in most instances) but overall effectiveness is limited: the costs and benefits of retraining vary significantly for different games, this strategy does not generalize, and, interestingly, users can still evade detection using novel configurations and arbitrary-shaped text. Communicating in this way yields very low bitrates (0.3-1.1 bits/s) which is suited for very short messages, and applications such as microblogging and bootstrapping off-game communications (dialing). This technique does not require technical sophistication and runs easily on existing games infrastructure without modification. We also discuss potential strategies to address efficiency, bandwidth, and security constraints of video game environments. To the best of our knowledge, this is the first such exploration of video games and virtual environments from a computer vision perspective.
Abdul Wajid, Nasir Kamal, Muhammad Sharjeel, Raaez Muhammad Sheikh, Huzaifah Bin Wasim, Muhammad Hashir Ali, Muhammad Wajahat Hussain, Syed Taha Ali, Latif Anjum
Proc. Priv. Enhancing Technol.7
2019 Adversarial Examples for Handcrafted Features
Zohaib Ali, Muhammad Latif Anjum, Muhammad Wajahat Hussain
BMVC3
2016 Dealing with small data and training blind spots in the Manhattan world
abstract
Leveraging Manhattan assumption we generate metrically rectified novel views from a single image, even for non-box scenarios. Our novel views enable the already trained classifiers to handle training data missing views (blind spots) without additional training. We demonstrate this on end-to-end scene text spotting under perspective. Additionally, utilizing our fronto-parallel views, we discover unsuspended invariant mid-level patches given a few widely separated training examples (small data domain). These invariant patches outperform various baselines on small data image retrieval challenge.
Muhammad Wajahat Hussain, Javier Civera 0001, Luis Montano, Martial Hebert
WACV1
2015 Single Image 3D without a Single 3D Image
abstract
Do we really need 3D labels in order to learn how to predict 3D? In this paper, we show that one can learn a mapping from appearance to 3D properties without ever seeing a single explicit 3D label. Rather than use explicit supervision, we use the regularity of indoor scenes to learn the mapping in a completely unsupervised manner. We demonstrate this on both a standard 3D scene understanding dataset as well as Internet images for which 3D is unavailable, precluding supervised learning. Despite never seeing a 3D label, our method produces competitive results.
David F. Fouhey, Muhammad Wajahat Hussain, Abhinav Gupta 0001, Martial Hebert
ICCV2
2015 Layout aware visual tracking and mapping
abstract
Nowadays real time visual Simultaneous Localization And Mapping (SLAM) algorithms exist and rely on consistent measurements across multiple views. In indoor environments, where majority of robot's activity takes place, severe occlusions can occur, e.g., when turning around a corner or moving from one room to another. In these situations, SLAM algorithms can not establish correspondences across views, which leads to failures in camera localization or map construction. This work takes advantage of the recent scene box layout descriptor to make the above mentioned SLAM systems occlusion aware. This room box reasoning helps the sequential tracker to reason about possible occlusions and therefore look for matches in only potentially visible features instead of the entire map. This increases the life of the tracker, as it does not consider itself lost under the occlusion state. Additionally, focusing on the potentially visible portion of the map, i.e., the current room features, it improves the computational efficiency without compromising the accuracy. Finally, this room level reasoning helps in better image selection for bundle adjustment. The image bundle coming from the same room has little occlusion, which leads to better dense reconstruction. We demonstrate the superior performance of layout aware SLAM on several long monocular sequences acquired in difficult indoor situations, specifically in a room-room transition and turning around a corner.
Marta Salas, Muhammad Wajahat Hussain, Alejo Concha, Luis Montano, Javier Civera 0001, J. M. M. Montiel
IROS2
2014 Grounding Acoustic Echoes in Single View Geometry Estimation
abstract
Extracting the 3D geometry plays an important part in scene understanding. Recently, robust visual descriptors are proposed for extracting the indoor scene layout from a passive agent’s perspective, specifically from a single image. Their robustness is mainly due to modelling the physical interaction of the underlying room geometry with the objects and the humans present in the room. In this work we add the physical constraints coming from acoustic echoes, generated by an audio source, to this visual model. Our audio-visual 3D geometry descriptor improves over the state of the art in passive perception models as we show in our experiments.
Muhammad Wajahat Hussain, Javier Civera 0001, Luis Montano
AAAI1