EDBT 2026 Demo / reviewers in the wild / expert
Manas Kamal Bhuyan
dblp:04/4787
· DBLP profile ↗
72ranked-venue papers
2as first author
32since 2021 · last 2026
0000-0003-2152-5466ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EffiSign network: a comprehensive approach for sign language recognition
Bhumika Karsh, Rabul Hussain Laskar, Ram Kumar Karsh, Manas Kamal Bhuyan |
Multim. Tools Appl. | 4 |
| 2026 | CTPNet: Achieving Real-Time Semantic Segmentation on Resource-Constrained Edge Devices for Autonomous Driving
Nadeem Atif, Saquib Mazhar, Mohammed Ameen, Shaik Rafi Ahamed, Manas Kamal Bhuyan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Validating polyp and instrument segmentation methods in colonoscopy through Medico 2020 and MedAI 2021 ChallengesabstractAutomatic analysis of colonoscopy images has been an active field of research motivated by the importance of early detection of precancerous polyps. However, detecting polyps during the live examination can be challenging due to various factors such as variation of skills and experience among the endoscopists, lack of attentiveness, and fatigue leading to a high polyp miss-rate. Therefore, there is a need for an automated system that can flag missed polyps during the examination and improve patient care. Deep learning has emerged as a promising solution to this challenge as it can assist endoscopists in detecting and classifying overlooked polyps and abnormalities in real time, improving the accuracy of diagnosis and enhancing treatment. In addition to the algorithm’s accuracy, transparency and interpretability are crucial to explaining the whys and hows of the algorithm’s prediction. Further, conclusions based on incorrect decisions may be fatal, especially in medicine. Despite these pitfalls, most algorithms are developed in private data, closed source, or proprietary software, and methods lack reproducibility. Therefore, to promote the development of efficient and transparent methods, we have organized the “Medico automatic polyp segmentation (Medico 2020)” and “MedAI: Transparency in Medical Image Segmentation (MedAI 2021)” competitions. The Medico 2020 challenge received submissions from 17 teams, while the MedAI 2021 challenge also gathered submissions from another 17 distinct teams in the following year. We present a comprehensive summary and analyze each contribution, highlight the strength of the best-performing methods, and discuss the possibility of clinical translations of such methods into the clinic. Our analysis revealed that the participants improved dice coefficient metrics from 0.8607 in 2020 to 0.8993 in 2021 despite adding diverse and challenging frames (containing irregular, smaller, sessile, or flat polyps), which are frequently missed during a routine clinical examination. For the instrument segmentation task, the best team obtained a mean Intersection over union metric of 0.9364. For the transparency task, a multi-disciplinary team, including expert gastroenterologists, accessed each submission and evaluated the team based on open-source practices, failure case analysis, ablation studies, usability and understandability of evaluations to gain a deeper understanding of the models’ credibility for clinical deployment. The best team obtained a final transparency score of 21 out of 25. Through the comprehensive analysis of the challenge, we not only highlight the advancements in polyp and surgical instrument segmentation but also encourage subjective evaluation for building more transparent and understandable AI-based colonoscopy systems. Moreover, we discuss the need for multi-center and out-of-distribution testing to address the current limitations of the methods to reduce the cancer burden and improve patient care. • We present a detailed analysis of the Medico 2020 and MedAI 2021 challenges that are aimed at advancing automated polyp and instrument segmentation in colonoscopy for early colorectal cancer diagnosis by using novel deep learning methods. • To the best of our knowledge, MedAI 2021 is the first challenge to evaluate the transparency in both GI endoscopy and colonoscopy. Through the challenge, we invited the participants to list package dependencies and architecture code (with instructions for building, compiling, and training) and share trained model weights in a standardized format. Additionally, we invited participants to include the code for model evaluation and provide repository licensing information to enable others to use the code and the trained model responsibly. Moreover, we asked the participants to explain model predictions using intermediate heatmaps, perform ablation studies, conduct a thorough failure analysis, and share their code for reproducing the results. Finally, we performed a subjective evaluation by including an expert gastroenterologist in the group and gave the final transparency score based on the usefulness and understandability of the results. Our initiative aims to promote transparency in AI research and foster the development of reliable, interpretable, and trustworthy algorithms for use in medical image segmentation. • We provide a comparative analysis of the 34 proposed methods in both challenges (3 subtasks), covering small details of each team in the form of Tables, qualitative and quantitative results (failure analysis), and an in-depth analysis of the findings. • We explore trust, safety, interpretability, transparency, and generalizability issues and provide future strategies to overcome the current limitations of developed algorithms. Debesh Jha, Vanshali Sharma, Debapriya Banik, Debayan Bhattacharya, Kaushiki Roy, Steven Alexander Hicks, Nikhil Kumar Tomar, Vajira Thambawita, Adrian Krenzer, Ge-Peng Ji, Sahadev Poudel, George Batchkala, Saruar Alam, Awadelrahman M. A. Ahmed, Quoc-Huy Trinh, Zeshan Khan, Tien-Phat Nguyen, Shruti Shrestha, Sabari Nathan, Jeonghwan Gwak, Ritika Kumari Jha, Zheyuan Zhang 0001, Alexander Schlaefer, Debotosh Bhattacharjee, Manas Kamal Bhuyan, Pradip K. Das, Deng-Ping Fan, Sravanthi Parasa, Sharib Ali, Michael Riegler 0001, Pål Halvorsen, Thomas de Lange, Ulas Bagci |
Medical Image Anal. | 25 |
| 2025 | Despeckling of Synthetic Aperture Radar Images Using Linear-Angular Attention TransformerabstractSynthetic aperture radar (SAR) images are often contaminated by speckle noise, a type of multiplicative noise resulting from the imaging process. SAR image despeckling is a crucial preprocessing step for satellite imaging, enhancing image visualization and facilitating downstream analysis. In this study, we propose a linear–angular attention transformer network for SAR despeckling. The approach efficiently captures both local and global context in linear time within a multiscale transformer-convolutional neural network architecture. Our transformer integrates nonlocal denoising and multiscale feature extraction in a single model. Using a smoothing loss and fast nonlocal post-processing, the model achieved a 17% improvement in structural similarity index, a 61% enhancement in contrast-to-noise ratio, and an increase above 100% in the equivalent number of looks metric across three datasets compared to multiple state-of-the-art baselines, even when trained on a small dataset and for only 13 epochs. Comparison with different multi-head self-attention mechanisms revealed the effectiveness of linear–angular attention as a step towards green AI, showing both quantitative and qualitative performance improvements. Unlike models that rely on optical images for training and lack domain-specific features for real SAR despeckling, the proposed network is trained directly on SAR images in a self-supervised manner. Souraja Kundu, Manas Kamal Bhuyan, Karl F. MacDorman, Neeraj Kumar Sharma 0007, Manish Bhatt |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Cross-Modality Medical Image Registration with Local-Global Spatial Correlation
Souraja Kundu, Yuji Iwahori, Manas Kamal Bhuyan, Manish Bhatt, Boonserm Kijsirikul, Aili Wang 0001, Akira Ouchi, Yasuhiro Shimizu |
ICPR (11) | 3 |
| 2024 | Sign Language Gesture Recognition Using YOLOv9 for Medical Attention of Hard of Hearing PopulationabstractFor the hard-of-hearing population, communicating medical issues to doctors who do not understand sign language can be challenging. To address this problem, researchers have focused extensively on sign language gesture recognition. However, existing methods often struggle with conversion time and the accuracy of recognizing similar gestures. In this paper, we propose a sign language gesture recognition system utilizing the Yolov9 model. The system's performance was evaluated using a publicly available Kaggle dataset, achieving an impressive Mean Average Precision ([email protected]) of 99.5%. Experimental results show that the proposed method surpasses state-of-the-art techniques in both efficiency and accuracy. Pranjal Gogoi, Bhumika Karsh, Ram Kumar Karsh, Rabul Hussain Laskar, Manas Kamal Bhuyan |
TENCON | 5 |
| 2024 | AttV19Net: Attention Based VGG 19 Network for Hand Gesture Recognition with a Prototype
Bhumika Karsh, Rabul Hussain Laskar, Ram Kumar Karsh, Manas Kamal Bhuyan |
TENCON | 4 |
| 2024 | Glaucoma detection with explainable AI using convolutional neural networks based feature extraction and machine learning classifiersabstractAbstract Glaucoma is an eye disease that damages the optic nerve as a result of vision loss, it is the leading cause of blindness worldwide. Due to the time‐consuming, inaccurate, and manual nature of traditional methods, automation in glaucoma detection is important. This paper proposes an explainable artificial intelligence (XAI) based model for automatic glaucoma detection using pre‐trained convolutional neural networks (PCNNs) and machine learning classifiers (MLCs). PCNNs are used as feature extractors to obtain deep features that can capture the important visual patterns and characteristics from fundus images. Using extracted features MLCs then classify glaucoma and healthy images. An empirical selection of the CNN and MLC parameters has been made in the performance evaluation. In this work, a total of 1,865 healthy and 1,590 glaucoma images from different fundus datasets were used. The results on the ACRIMA dataset show an accuracy, precision, and recall of 98.03%, 97.61%, and 99%, respectively. Explainable artificial intelligence aims to create a model to increase the user's trust in the model's decision‐making process in a transparent and interpretable manner. An assessment of image misclassification has been carried out to facilitate future investigations. Vijaya Kumar Velpula, Diksha Sharma, Lakhan Dev Sharma, Amarjit Roy, Manas Kamal Bhuyan, Sultan Alfarhood, Mejdl S. Safran |
IET Image Process. | 5 |
| 2024 | Semi-supervised generative adversarial networks for improved colorectal polyp classification using histopathological images
Pradipta Sasmal, Vanshali Sharma, Allam Jaya Prakash, Manas Kamal Bhuyan, Kiran Kumar Patro, Nagwan Abdelsamee, Hayam Alamro, Yuji Iwahori, Ryszard Tadeusiewicz, U. Rajendra Acharya, Pawel Plawiak |
Inf. Sci. | 4 |
| 2024 | A Multi-Scale Attention Framework for Automated Polyp Localization and Keyframe Extraction From Colonoscopy VideosabstractColonoscopy video acquisition has been tremendously increased for retrospective analysis, comprehensive inspection, and detection of polyps to diagnose colorectal cancer (CRC). However, extracting meaningful clinical information from colonoscopy videos requires an enormous amount of reviewing time, which burdens the surgeons considerably. To reduce the manual efforts, we propose a first end-to-end automated multi-stage deep learning framework to extract an adequate number of clinically significant frames, i.e., keyframes from colonoscopy videos. The proposed framework comprises multiple stages that employ different deep learning models to select keyframes, which are high-quality, non-redundant polyp frames capturing multi-views of polyps. In one of the stages of our framework, we also propose a novel multi-scale attention-based model, YcOLOn, for polyp localization, which generates ROI and prediction scores crucial for obtaining keyframes. We further designed a GUI application to navigate through different stages. Extensive evaluation in real-world scenarios involving patient-wise and cross-dataset validations shows the efficacy of the proposed approach. The framework removes 96.3% and 94.02% frames, reduces detection processing time by 38.28% and 59.99%, and increases mAP by 2% and 5% on the SUN database and the CVC-VideoClinicDB, respectively. The source code is available at https://github.com/Vanshali/KeyframeExtractionNote to Practitioners—The widespread acceptance of colonoscopy procedures as a gold standard for CRC screening is constrained by the massive amount of data recorded during the process that needs to be manually reviewed. Such manual procedures are burdensome and induce human diagnostic errors. This article suggests an automated framework to extract keyframes (important frames) from colonoscopy videos that can efficiently represent the clinically relevant information captured in the video streams. This is achieved by the automated removal of uninformative and highly correlated frames, which do not add to clinical findings. The approach ensures diversity among keyframes and provides clinicians with a multi-view of polyps for easy resection. In addition, the proposed multi-scale attention-based model improves the polyp localization performance, which further helps in refining the keyframe selection process. The comprehensive experimental results corroborate that discarding insignificant frames can enhance polyp detection and localization performance and reduce computational requirements. The study estimates 30% to 60% time saving for clinicians during video screening. In clinical practices, the proposed automated framework and our designed GUI would enable surgeons to visualize the essential data better with minimal manual interventions and assist in precise polyp resection. Vanshali Sharma, Pradipta Sasmal, Manas Kamal Bhuyan, Pradip K. Das, Yuji Iwahori, Kunio Kasugai |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2023 | Can Adversarial Networks Make Uninformative Colonoscopy Video Frames Clinically Informative? (Student Abstract)abstractVarious artifacts, such as ghost colors, interlacing, and motion blur, hinder diagnosing colorectal cancer (CRC) from videos acquired during colonoscopy. The frames containing these artifacts are called uninformative frames and are present in large proportions in colonoscopy videos. To alleviate the impact of artifacts, we propose an adversarial network based framework to convert uninformative frames to clinically relevant frames. We examine the effectiveness of the proposed approach by evaluating the translated frames for polyp detection using YOLOv5. Preliminary results present improved detection performance along with elegant qualitative outcomes. We also examine the failure cases to determine the directions for future work. Vanshali Sharma, Manas Kamal Bhuyan, Pradip K. Das |
AAAI | 2 |
| 2023 | Polyp Size and Shape Estimation by Using an Endoscopic Hood InformationabstractMedical doctors identify benign or malignant colon polyps by their size and shape. Since this identification is related to whether or not resection surgery is necessary, a technology for estimating absolute size and shape from endoscopic images is required. In previous research, the method for recovering the size and shape of polyps using a blood vessel region as a reference object has been proposed. This paper proposes a method to recover the size and shape of polyps by using a cylindrical endoscopic hood as a reference object. The experimental results confirm that the proposed method can estimate reflectance factor without using blood vessel information. Ryo Kikuchi, Yuji Iwahori, Kenji Funahashi, Manas Kamal Bhuyan, Aili Wang 0001, Naotaka Ogasawara, Kunio Kasugai |
KES | 4 |
| 2023 | Block attention network: A lightweight deep network for real-time semantic segmentation of road scenes in resource-constrained devices
Saquib Mazhar, Nadeem Atif, Manas Kamal Bhuyan, Shaik Rafi Ahamed |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Efficient hand segmentation for rehabilitation tasks using a convolution neural network with attention
H. Pallab Jyoti Dutta, Manas Kamal Bhuyan, Debanga Raj Neog, Karl F. MacDorman, Rabul Hussain Laskar |
Expert Syst. Appl. | 2 |
| 2023 | Unsupervised story segmentation and indexing of broadcast news video
Pranabjyoti Haloi, Manas Kamal Bhuyan, Dibyajyoti Chatterjee, Pooja Rani Borah |
Multim. Tools Appl. | 2 |
| 2023 | A DWT-based encoder-decoder network for Specularity segmentation in colonoscopy images
Vanshali Sharma, Manas Kamal Bhuyan, Pradip K. Das, Kangkana Bora |
Multim. Tools Appl. | 2 |
| 2023 | Detection, tracking, and recognition of isolated multi-stroke gesticulated characters
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Manas Kamal Bhuyan |
Pattern Anal. Appl. | 4 |
| 2022 | Keyframe Selection from Colonoscopy Videos to Enhance Visualization for Polyp DetectionabstractColonoscopy video acquisition and recording have been increasingly performed for comprehensive diagnosis and retrospective analysis of colorectal cancer (CRC). Reviewing video streams helps detect and inspect polyps, the precursor to CRC. However, visualizing these streams in their raw form puts a considerable burden on clinicians as most of the frames are clinically insignificant and are not useful for pathological interpretation. For improved visualization of diagnostically significant information, we have proposed an automated framework that discards the uninformative frames from raw videos. Our approach initially extracts high-quality colonoscopy frames using a deep learning model to assist clinicians in visualizing data in a refined form. Subsequently, our work validates the effectiveness of keyframe selection by employing polyp detection models. All the evaluations are performed either patient-wise or cross-dataset to suffice the real-time requirements. Experimental results show that the keyframe extraction saves reviewing time and enhances the detection performances. The proposed approach achieves a polyp detection F1-score of 79.78% (patient-wise) and 89.22% (cross-dataset) on the SUN and CVC-VideoClinicDB databases, respectively. Vanshali Sharma, Pradipta Sasmal, Manas Kamal Bhuyan, Pradip K. Das |
IV | 3 |
| 2022 | Improvement of Polyp Detection using MUNIT for Image GenerationabstractThis paper treats an image translation using Multimodal Unsupervised Image-to-Image Translation (MUNIT) from the white light source image to the Narrow Band Imaging (NBI) of the endoscope images and proposes a method to improve the detection performance by the Single Shot Multibox Detector (SSD) which trains dataset of white light source image by adding the generated NBI-like images. The proposed approach makes it possible to generate an NBI-like image by keeping the existing polyp, inner wall and specific features of coloring and brightness of the original endoscope image. It is shown that increasing the number of data of the generated images achieves the better performance. The performance of the proposed approach was evaluated using the actual endoscope images which include the polyps of the various shapes. Recall of 80.62% and precision of 93.47% were obtained as a result through the evaluation of computer experiments. Yuji Iwahori, Tsubasa Ooto, Hiroyasu Usami, Shinji Fukui 0001, Manas Kamal Bhuyan, Aili Wang 0001, Naotaka Ogasawara, Kunio Kasugai |
KES | 5 |
| 2022 | Lymph-node Detection and Metastasis Classification from CT Images using a Single U-Net ModelabstractThe presence or absence of cancer metastasis in the lymph-nodes using an AI-based approach has been important these days in the medical field of gastroenterological surgery. The medical field needs the introduction of machine learning to help the knowledge and skill of surgery of medical doctors who are taking operations by checking CT scans or MRI images for cancer disease. Recent machine learning based researches on lymph-node are mainly either detection of the location of lymph-nodes or classification of cancer metastasis. This paper proposes a method to perform both tasks at the same time with a single U-Net model. Obtained results in the experiments show that the classification accuracy is further improved compared with the previous approach while keeping the detection ratio as almost the same level as that of the related paper. Kosuke Suzuki, Yuji Iwahori, Kenji Funahashi, Manas Kamal Bhuyan, Akira Ouchi, Yasuhiro Shimizu |
KES | 4 |
| 2022 | Design and development of a vision-based system for detection, tracking and recognition of isolated dynamic bare hand gesticulated charactersabstractAbstract Detection and tracking are the vital stages to form the gesture trajectory in gesture recognition. It becomes more challenging when the variations in illumination, pose, position, occlusion, scale, speed, blurring effect and complex environment are introduced. Additionally, the background feature domination effect affects the existing deep learning models. A semantic segmentation model is implemented in this work to detect the bare hand to overcome these challenges. A pre‐trained network VGG‐16 is utilized by training with the proposed NITS S‐Net database. Evaluation of the SegNet model is done on EgoHands, Oxford and OUHands databases. To track the bare hand, a SegNet‐based detection and tracking approach is proposed using Kalman filter and point‐tracker. This model achieves 97.01% accuracy (a relative improvement of ~8% from the baseline models) at 0.068 s per frame computational time on NITS hand gesture database VIIIB. The gesticulated characters, that is, alphabets, numbers, operators, special characters, are gesticulated without any constraints on the pattern/strokes. To recognize these 95 multi‐stroke gestures, a deep convolutional neural network (DCNN) is presented using AlexNet. The DCNN model achieves 97.60% (a relative improvement of ~14% from the baseline models) accuracy on the NITS hand gesture database VIIIB merged. Evaluation of the handwritten EMNIST merged (balanced) database resulted in average recognition accuracy of 91.60%. Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Manas Kamal Bhuyan |
Expert Syst. J. Knowl. Eng. | 4 |
| 2022 | Development of an intelligent recognition system for dynamic mid-air gesticulation of isolated alphanumeric keys
Anish Monsley K., Kuldeep Singh Yadav, Rabul Hussain Laskar, Manas Kamal Bhuyan |
Expert Syst. Appl. | 4 |
| 2022 | A selective region-based detection and tracking approach towards the recognition of dynamic bare hand gesture using deep neural network
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Songhita Misra, Manas Kamal Bhuyan |
Multim. Syst. | 5 |
| 2022 | A curvelet-based multi-sensor image denoising for KLT-based image fusion
Amit Vishwakarma, Manas Kamal Bhuyan |
Multim. Tools Appl. | 2 |
| 2022 | An unsupervised approach of colonic polyp segmentation using adaptive markov random fields
Pradipta Sasmal, Manas Kamal Bhuyan, Soumayan Dutta, Yuji Iwahori |
Pattern Recognit. Lett. | 2 |
| 2021 | Contextual Emotion Learning ChallengeabstractEmotion recognition via vision has been deeply associated with facial expressions, and the inference of emotions has, more often than not, been based on the same. However, context, both environmental and social, plays an imperative role in emotion recognition but has not been incorporated widely so far. The meaning of emotion might entirely switch when shifted from one setting to another if only facial expressions are taken into account. Moreover, there exists no study in the Indian context about the same. To cater to this issue, we generate and introduce the Indian Contextual Emotion Recognition (ICER) dataset based on the multi-ethnic Indian context. This paper summarises the Contextual Emotion Learning Challenge (CELC 2021) organized in conjunction with the 16th IEEE Conference on Automatic Face and Gesture Recognition (FG) 2021. We outline the tasks posed in the challenge, the novel dataset, along with its challenges and the evaluation method. Lastly, we conclude by discussing the possible future directions. Jainendra Shukla, Puneet Gupta 0002, Aniket Bera, Arka Sarkar, Prakhar Goel, Shubhangi Butta, Anup Kumar Gupta 0001, Snehil Sanyal, Debanga Raj Neog, Manas Kamal Bhuyan, Kalyani Marathe, Linda G. Shapiro, Alex Colbrn, Varchita Lalwani |
FG | 10 |
| 2021 | Automatic Generation of Polyp Image using Depth Map for Endoscope DatasetabstractIn recent years, opportunities for diagnosis using endoscopy aiming a less invasive treatment are increasing following the disease rate of colorectal cancer. Computer-aided diagnosis has been developed based on deep learning methodology, it aiming to improve the accuracy of diagnosis and support immature medical doctors. To satisfy the learning dataset, this paper proposes a data augmentation methodology where automatic image generation of polyp images using Pix2Pix and depth map obtained from the original image. The problem of lack of the learning dataset of polyp images can be solved by the proposed approach and the effectiveness of the generated data was confirmed by the quantitative evaluation with the improved performance of SSD (Single Shot Multibox Detector) in the experiments. Haruki Yamane, Shinji Fukui 0001, Yuji Iwahori, Hiroyasu Usami, Manas Kamal Bhuyan, Naotaka Ogasawara, Kunio Kasugai |
KES | 5 |
| 2021 | Recognition of isolated characters across different input interfaces using 2D DCNNabstractRecognition of the characters has gained much attention due to its potential applications like document analysis, license plate detection, house number detection, virtual text entry system, etc., in pattern recognition. However, it is very challenging to recognize the characters under the variations in pattern, style, translation, scale, rotation. This work develops a computationally efficient deep learning model to recognize handwritten, printable, and gesticulated characters. For gesture, the NITS gesticulated database having 60 characters (10 digits, 26 English uppercase alphabets, 4 operations, 18 special symbols) is proposed with the variation in pattern, style, scale in this work. To evaluate the ability and robustness of the proposed model, the handwritten characters (MNIST, EMNIST), printable characters (SVHN, Chars74) databases are considered. This network achieves 94.55%, 89.54%, 87.33, and 93.90% recognition accuracy on NITS gesticulated, EMNIST merge (balanced), SVHN, and Chars74 databases. Kuldeep Singh Yadav, Anish Monsley K., Saharul Alom Barlaskar, Naseem Ahmad, Rabul Hussain Laskar, Manas Kamal Bhuyan |
TENCON | 6 |
| 2021 | Segregation of meaningful strokes, a pre-requisite for self co-articulation removal in isolated dynamic gesturesabstractAbstract Gesture formation, a pre‐processing step, has its importance when variations in patterns, scale, and speed come into play. Self co‐articulations are intentional movements performed by an individual to complete a gesture, whose presence in the trajectory alters its original meaning. For recognition, most researchers have directly used the trajectory formed along with these self co‐articulated strokes, with a few removing it using visible trait‐like velocity. Usage of velocity has shortcomings as gesturing in air differs from gesturing over a solid surface; hence, we propose a gesture formation model, which incorporates global and local measures to remove these self co‐articulations. The global measure uses Euclidean distance, instantaneous velocity, and polarity calculated from the complete gesture, while the local measure segments the gesture into stroke‐level segments by using the minimum–maximum‐polarity algorithm and applies the selective bypass rules. The proposed model, when experimented on gestures patterns with premeditated speed variation, has a mean error rate of 0.0069 and 7.40% self co‐articulations;individuals’ natural gesticulation has a mean error rate of 0.0371 and 12.07% self co‐articulations. Experimentation on each gesture of NITS hand gesture databases showed a relative improvement of 40% (accuracy 97%) over the existing baseline models. Anish Monsley K., Kuldeep Singh Yadav, Songhita Misra, Manas Kamal Bhuyan, Rabul Hussain Laskar |
IET Image Process. | 5 |
| 2021 | Skin detection in video under uncontrolled illumination
Biplab Ketan Chakraborty, Manas Kamal Bhuyan, Karl F. MacDorman |
Multim. Tools Appl. | 2 |
| 2021 | A framework for continuous fingerspelling spotting for H.264/AVC compressed videos using spatio-temporal Markov random field
Anjan Kumar Talukdar, Manas Kamal Bhuyan |
Multim. Tools Appl. | 2 |
| 2021 | Multi-level uncorrelated discriminative shared Gaussian process for multi-view facial expression recognition
Manas Kamal Bhuyan, Yuji Iwahori |
Vis. Comput. | 2 |
| 2020 | Interdependent Multi-task Learning for Simultaneous Segmentation and Detection
Mahesh Reginthala, Yuji Iwahori, Manas Kamal Bhuyan, Yoshitsugu Hayashi, Witsarut Achariyaviriya, Boonserm Kijsirikul |
ICPRAM | 3 |
| 2020 | Mediastinal Lymph Node Detection using Deep Learning
Jayant P. Singh, Yuji Iwahori, Manas Kamal Bhuyan, Hiroyasu Usami, Taihei Oshiro, Yasuhiro Shimizu |
ICPRAM | 3 |
| 2020 | Colorectal Polyp Classification Based On Latent Sharing Features Domain from Multiple Endoscopy ImagesabstractAs a method to judge the benign or malignant polyp from endoscope images, some methods have been proposed using an ultra-high magnification endoscope. The ultra-high magnification endoscope enables the diagnosis at the cell level. However, it tends to spend many times for diagnosis and requires specific expensive devices. There are three types of images that are taken for diagnosis by the regular endoscope: white light, dye, and narrowband image (NBI) in general. This paper proposes a benign/malignant polyp classification method using these images taken by the regular endoscope. Each image features derived from endoscope images are extracted by adapting a pre-trained CNN to each domain. Finally, polyps are classified using extracted features. Experiments confirmed that the proposed method enabled the classification of benign or malignant colorectal polyps with over 90% accuracy. Hiroyasu Usami, Yuji Iwahori, Yoshinori Adachi, Manas Kamal Bhuyan, Aili Wang 0001, Satoshi Inoue, Masahide Ebi, Naotaka Ogasawara, Kunio Kasugai |
KES | 4 |
| 2020 | Image specific discriminative feature extraction for skin segmentation
Biplab Ketan Chakraborty, Manas Kamal Bhuyan |
Multim. Tools Appl. | 2 |
| 2020 | Image mosaicking using improved auto-sorting algorithm and local difference-based harris features
Amit Vishwakarma, Manas Kamal Bhuyan |
Multim. Tools Appl. | 2 |
| 2019 | Polyp Classification and Clustering from Endoscopic Images using Competitive and Convolutional Neural Networks
Avish Kabra, Yuji Iwahori, Hiroyasu Usami, Manas Kamal Bhuyan, Naotaka Ogasawara, Kunio Kasugai |
ICPRAM | 4 |
| 2019 | Contour-Aware Residual W-Net for Nuclei SegmentationabstractNuclei segmentation is an important pre-processing step for any vision based cytopathological diagnostic system which extracts information from nuclei to perform tasks such as cancer detection. A cell nuclei segmentation pipeline should be robust, accurate and fast. We propose a deep learning based model, Contour-Aware Residual W-Net (WRC-Net), which consists of double U-Net, [5] or W-Net. The first U-Net learns to predict nuclei boundaries and the second generates the segmentation map. Our model can accurately segment a 128x128 dimensional image in less than 0.05s. Our model can learn from a very limited training data with as low as a single training image. We tested our model on real HE (Hematoxylin and Eosin) stained cell images and it showed better overall performance against previous state-of-the-art nuclei segmentation methods. Sushmita Das, Ankur Deka, Yuji Iwahori, Manas Kamal Bhuyan, Takashi Iwamoto, Jun Ueda |
KES | 4 |
| 2019 | Detecting and Removing Specular Reflectance Components Based on Image LinearizationabstractShape from Shading and Photometric Stereo are famous approaches to recover the 3D shape from image(s). These approaches can obtain 3D shape from observed gray scale image(s) but it is necessary to estimate reflectance parameters for objects when some specular reflectance components are observed. Specular reflectance is difficult to be handled as diffuse reflectance in general. It is expected to remove specular reflectance component without assuming any kind of reflectance function (reflectance model). This paper proposes a method to remove specular reflection components using 4 observed images taken under 4 different light source directions using relighting approach of diffused reflectance based on image linearization. Results are demonstrated via computer simulation and experiments. Ryosuke Nakao, Yuji Iwahori, Yoshinori Adachi, Aili Wang 0001, Manas Kamal Bhuyan, Boonserm Kijsirikul |
KES | 5 |
| 2019 | Automatic Detection of Aortic Valve Opening Using Seismocardiography in Healthy IndividualsabstractAccurate detection of fiducial points in a seismocardiogram (SCG) is a challenging research problem for its clinical application. In this paper, an automated method for detecting aortic valve opening (AO) instants using the dorso-ventral component of the SCG signal is proposed. This method does not require electrocardiogram (ECG) as a reference signal. After preprocessing the SCG, multiscale wavelet decomposition is carried out to get signal components in different wavelet subbands. The subbands having possible AO peaks are selected by a newly proposed dominant-multiscale-kurtosis- and dominant-multiscale-central-frequency-based criterion. The signal is reconstructed using selected subbands, and it is emphasized using the weights derived from the proposed relative squared dominant multiscale kurtosis. The Shannon energy followed by autocorrelation coefficients is computed for systole envelope construction. Finally, AO peaks are detected by a Gaussian-derivative-filtering-based scheme. The robustness of the proposed method is tested using clean and noisy SCG signals from the combined measurement of ECG, breathing, and SCG database. Evaluation results show that the method can achieve an average sensitivity of 94%, a prediction rate of 90%, and a detection accuracy of 86% approximately over 4585 analyzed beats. Tilendra Choudhary, L. N. Sharma, Manas Kamal Bhuyan |
IEEE J. Biomed. Health Informatics | 3 |
| 2018 | Defect Classification of Electronic Board Using Dense SIFT and CNNabstractThis paper proposes a new defect classification method of electronic board using Dense SIFT and CNN which can represent the effective features to the gray scale image. Proposed method does not use any reference image and effective keypoints are detected using Dense SIFT on the defect candidate region. Removing the feature points except defect region and Bag of Features are used to represent the histogram features. Dense SIFT and SVM are used to judge defect or not. CNN is further introduced to classify true or pseudo defect. Classification accuracy was evaluated and effectiveness of the proposed method is shown. Yuji Iwahori, Yohei Takada, Tokiko Shiina, Yoshinori Adachi, Manas Kamal Bhuyan, Boonserm Kijsirikul |
KES | 5 |
| 2018 | Review of constraints on vision-based gesture recognition for human-computer interactionabstractThe ability of computers to recognise hand gestures visually is essential for progress in human–computer interaction. Gesture recognition has applications ranging from sign language to medical assistance to virtual reality. However, gesture recognition is extremely challenging not only because of its diverse contexts, multiple interpretations, and spatio‐temporal variations but also because of the complex non‐rigid properties of the hand. This study surveys major constraints on vision‐based gesture recognition occurring in detection and pre‐processing, representation and feature extraction, and recognition. Current challenges are explored in detail. Biplab Ketan Chakraborty, Debajit Sarma, Manas Kamal Bhuyan, Karl F. MacDorman |
IET Comput. Vis. | 3 |
| 2018 | Hierarchical uncorrelated multiview discriminant locality preserving projection for multiview facial expression recognition
Manas Kamal Bhuyan, Brian C. Lovell, Yuji Iwahori |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | An optimized non-subsampled shearlet transform-based image fusion using Hessian features and unsharp masking
Amit Vishwakarma, Manas Kamal Bhuyan, Yuji Iwahori |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Non-subsampled shearlet transform-based image fusion using modified weighted saliency and local difference
Amit Vishwakarma, Manas Kamal Bhuyan, Yuji Iwahori |
Multim. Tools Appl. | 2 |
| 2018 | Heart Sound Extraction From Sternal Seismocardiographic SignalabstractThe phonocardiogram (PCG) signal indicates closing instants of atrio-ventricular and semilunar valves, and this information can also be extracted from two major profiles of a seismocardiographic (SCG) cycle. This letter presents a method to extract fundamental heart sounds (HSs) from a SCG signal. The proposed method employs discrete wavelet transform for signal decomposition, and subsequently, center of gravity (CoG) of power spectrum for each of the subbands is computed. Our method can optimally select wavelet subbands for any combination of sampling frequencies of the signal and number of decomposition levels. Based on proposed CoG criterion, the signal is reconstructed from a number of selected subbands, and instantaneous Hilbert envelope is constructed for localizing peaks of the signal. Finally, a simple decision rule is applied to annotate S1 and S2 sounds of the PCG signal. Experimental results show that the HS waves S1 and S2 can be well localized with the help of one SCG cycle without using a reference ECG cycle. Tilendra Choudhary, L. N. Sharma, Manas Kamal Bhuyan |
IEEE Signal Process. Lett. | 3 |
| 2017 | 3D Shape from SEM Image Using Improved Fast Marching Method
Yuji Iwahori, Aili Wang 0001, Manas Kamal Bhuyan |
ACIVS | 4 |
| 2017 | A Robust Method for Blood Vessel Extraction in Endoscopic Images with SVM-based Scene Classification
Mayank Golhar, Yuji Iwahori, Manas Kamal Bhuyan, Kenji Funahashi, Kunio Kasugai |
ICPRAM | 3 |
| 2017 | Automatic Polyp Detection from Endoscope Image using Likelihood Map based on Edge Information
Yuji Iwahori, Hiroaki Hagi, Hiroyasu Usami, Robert J. Woodham, Aili Wang 0001, Manas Kamal Bhuyan, Kunio Kasugai |
ICPRAM | 6 |
| 2017 | Tracking with Extraction of Moving Object under Moving Camera EnvironmentabstractThis paper proposes a new approach to archive the robust tracking of moving objects under moving camera environment where the similar moving objects cross each other. Tracking with moving camera sometimes fails to track the object with similar color objects or similar background. Proposed approach is a particle filter based approach. It introduces the likelihood calculated by probabilistic background model which is constructed using dense optical flow and fast density estimation. Proposed approach introduces SVM (Support Vector Machine)1 to judge the scene where it is difficult to construct the probabilistic background model with non-uniform optical flow. This SVM uses the degree histogram of optical flow. Usefulness of proposed approach is evaluated in the experiments using actual video and the performance is compared with recent tracking approaches by quantitative evaluations. Daimu Oiwa, Shinji Fukui 0001, Yuji Iwahori, Boonserm Kijsirikul, Tsuyoshi Nakamura, Manas Kamal Bhuyan |
KES | 6 |
| 2017 | Performance analysis of Gabor wavelet for extracting most informative and efficient features
T. Malathi, Manas Kamal Bhuyan |
Multim. Tools Appl. | 2 |
| 2017 | Combining image and global pixel distribution model for skin colour segmentation
Biplab Ketan Chakraborty, Manas Kamal Bhuyan |
Pattern Recognit. Lett. | 2 |
| 2016 | Tracking with probabilistic background model by density forestsabstractThis paper proposes an approach for a tracking method robust to the intersection with objects with appearances similar to a target object. The proposed method targets image sequences taken by a moving camera and is based on the particle filter. Tracking methods using color information tend to track mistakenly a background region or an object with color similar to the target object. The method constructs the probabilistic background model by the histogram of the optical flow and defines the likelihood function so that the likelihood in the region of the target object may become large. This causes increasing the accuracy of tracking. The probabilistic background model is made by the density forests. It can infer a probabilistic density fast. Results are demonstrated by experiments using the real videos of outdoor scenes. Daimu Oiwa, Shinji Fukui 0001, Yuji Iwahori, Tsuyoshi Nakamura, Manas Kamal Bhuyan |
ICIS | 5 |
| 2016 | Estimating Reflectance Parameter of Polyp using Medical Suture Information in Endoscope ImageabstractAn endoscope is a medical instrument that acquires images inside the human body. In this paper, a new 3-D reconstruction approach is proposed to estimate the size and shape of the polyp under conditions of both point light source illumination and perspective projection. Previous approaches could not know the size of polyp without assuming reflectance parameters as known constant. Even if it was possible to estimate the absolute size of polyp, it was assumed that the parameter of camera movement ∆Z is treated as a known along the depth direction. Here two images are used with a medical suture which is known size object to solve this problem and the proposed approach shows the parameter of camera movement can be estimated with robust accuracy with correspondence between two images taken via slight movement of Z. Experiments with endoscope images are demonstrated to evaluate the validity of proposed approach. Yuji Iwahori, Daiki Yamaguchi, Tsuyoshi Nakamura, Boonserm Kijsirikul, Manas Kamal Bhuyan, Kunio Kasugai |
ICPRAM | 5 |
| 2016 | Particle Filter Based Tracking with Image-based LocalizationabstractIn this paper, we propose a new method for object tracking robust to the intersection with other objects with similar appearance and to the great rotation of the camera. The method uses 3D information of the feature points and the camera position by the image-based localization method. The movement information of the camera is used by the prediction process and the calculation process of the likelihood. Furthermore, the method extracts foreground objects by the homography transformation. The result is used by the likelihood function and by the process judging whether the target object is occluded by a background object or not. The proposed method can track the target object robustly when the camera rotates greatly and when the target object is occluded by the background object or the other object. Results are demonstrated by experiments using real videos. Shinji Fukui 0001, So Hayakawa, Yuji Iwahori, Tsuyoshi Nakamura, Manas Kamal Bhuyan |
KES | 5 |
| 2016 | New Feature for Shadow Detection by Combination of Two Features Robust to Illumination ChangesabstractComputer vision methods need to deal with shadows explicitly because shadows often have a negative effect on the results computed. A new shadow detection method is proposed. The proposed method is a shadow model based method. A new feature for detecting shadows is introduced. The feature is obtained by L*a*b* components, Peripheral Increment Sign Correlation and Normalized Vector Distance. These features are robust to illumination changes. Shadows can be treated as local illumination changes. Using these features results in removing shadow effects, in part. The histogram is generated by the three features and is treated as the feature for detecting shadows. The SVM is used for the classifier. The SVM is trained in advance by shadow data and the trained SVM is used for detecting shadows. The proposed method can extract shadows with the accuracy similar to the previous approach in shorter time. Results are demonstrated by experiments using the real videos. Kota Higashi, Shinji Fukui 0001, Yuji Iwahori, Yoshinori Adachi, Manas Kamal Bhuyan |
KES | 5 |
| 2016 | Extraction of informative regions of a face for facial expression recognitionabstractThe aim of facial expression recognition (FER) algorithms is to extract discriminative features of a face. However, discriminative features for FER can only be obtained from the informative regions of a face. Also, each of the facial subregions have different impacts on different facial expressions. Local binary pattern (LBP) based FER techniques extract texture features from all the regions of a face, and subsequently the features are stacked sequentially. This process generates the correlated features among different expressions, and hence affects the accuracy. This research moves toward addressing these issues. The authors' approach entails extracting discriminative features from the informative regions of a face. In this view, they propose an informative region extraction model, which models the importance of facial regions based on the projection of the expressive face images onto the neural face images. However, in practical scenarios, neutral images may not be available, and therefore the authors propose to estimate a common reference image using Procrustes analysis. Subsequently, weighted‐projection‐based LBP feature is derived from the informative regions of the face and their associated weights. This feature extraction method reduces miss‐classification among different classes of expressions. Experimental results on standard datasets show the efficacy of the proposed method. Manas Kamal Bhuyan, Biplab Ketan Chakraborty |
IET Comput. Vis. | 2 |
| 2016 | Asymmetric occlusion detection using linear regression and weight-based filling for stereo disparity map estimationabstractStereo matching computes the disparity information from stereo image pairs. A number of stereo matching methods have been proposed to estimate a fine disparity map. However, objects present in the images are occluded on account of different camera viewpoints in a stereo vision setup, and hence it is quite difficult to get a fine disparity map. The methods which use disparity map information of two cameras (symmetric approach) to detect occluded pixels are computationally more complex. The authors approach entails to detect the occluded pixels only by using single disparity map information (asymmetric approach). The behaviour of reference and target pixels are analysed, and it is observed that the target matching pixels almost follow a linear pattern with respect to the reference image pixels. Hence, it is approximated by a linear regression model, and subsequently this model is used to detect the occluded pixels in the authors’ method. Finally, a fine disparity map is obtained by incorporating a novel occlusion filling method. Experimental results show that the proposed occlusion detection method gives almost similar performance as that of the methods which use two disparity maps for detection. For occlusion filling, the authors utilise support weights from both the stereo images, and hence their method can give better performance. T. Malathi, Manas Kamal Bhuyan |
IET Comput. Vis. | 2 |
| 2015 | Recovering size and shape of polyp from endoscope image by RBF-NN modificationabstractPrevious approaches have proposed to recover the poly shape but it is desired that absolute size of polyp can be obtained as a medical endoscope system. The VBW (Vogel-Breuss-Weickert) model is proposed as a method to recover 3-D shape under point light source illumination and perspective projection. However, the VBW model recovers relative, not absolute, shape. Here, shape modification is introduced to recover the exact shape. Modification is applied to the output of the VBW model. First, a local brightest point is used to estimate the reflectance parameter from two images obtained with movement of the endoscope camera in depth. After the reflectance parameter is estimated, a sphere image is generated and used for Radial Basis Function Neural Network (RBF-NN) learning. The NN implements the shape modification. NN input is the gradient parameters produced by the VBW model for the generated sphere. NN output is the true gradient parameters for the true values of the generated sphere. Depth can then be recovered using the modified gradient parameters. It was confirmed that NN gives better performance than the linear regression via computer simulation and real experiment. Seiya Tsuda, Yuji Iwahori, Yuki Hanai, Robert J. Woodham, Manas Kamal Bhuyan, Kunio Kasugai |
ICIP | 5 |
| 2015 | Improvement of Recovering Shape from Endoscope Images Using RBF Neural Network
Yuji Iwahori, Seiya Tsuda, Robert J. Woodham, Manas Kamal Bhuyan, Kunio Kasugai |
ICPRAM (2) | 4 |
| 2015 | Object Tracking with Improved Detector of Objects Similar to TargetabstractTracking methods based on the particle filter uses frequently the appearance information of the target object to calculate the likelihood. The method using it often fails in tracking when the target object intersects with objects similar to the target object. We propose a new approach for tracking an object in a video sequence taken by a moving camera. The proposed method is based on the particle filter. During tracking the target object, the method detects a similar object near the target object by the Mean-Shift tracker. After detecting the object, the size of it is recalculated and the similar object is tracked by the same way with the target object. The positions of the similar objects are used for calculating the likelihood and for judging situation under which the target object exists. These prevent the method from tracking the other object mistakenly. Results are demonstrated by experiments using real video sequences. Shinji Fukui 0001, Ryuji Nishiyama, Yuji Iwahori, Manas Kamal Bhuyan, Robert J. Woodham |
KES | 4 |
| 2015 | Automatic Detection of Polyp Using Hessian Filter and HOG FeaturesabstractAn endoscope is a medical instrument that acquires images inside the human body. This paper proposes a new approach for the automatic detection of polyp regions in an endoscope image using a Hessian Filter and machine learning approaches. The approach improves performance of automatic detection of polyp detection with higher accuracy. The approach uses HOG feature as a local feature since the polyp and non-polyp region often have similar color information. The approach also uses Real Adaboost and Random Forests as classifiers which works effciently even when the dimension of feature vector becomes large. It is suggested that Hessian filter can contribute to reducing the computational time in comparison with the case when only HOG features are used to detect the polyp region. K-means++ is introduced to integrate the detection results in the classification. It is shown that polyp detection with high accuracy is performed in the computer experiments with endoscope images. Yuji Iwahori, Akira Hattori, Yoshinori Adachi, Manas Kamal Bhuyan, Robert J. Woodham, Kunio Kasugai |
KES | 4 |
| 2015 | Estimation of disparity map of stereo image pairs using spatial domain local Gabor waveletabstractThe stereo matching problem takes two images captured by nearby cameras and attempts to recover quantitative disparity information. Most of the existing stereo matching algorithms find it difficult to estimate disparity in the occlusion, discontinuities and textureless regions in the images. In the last few decades, a number of stereo matching methods have been proposed to overcome some of these problems. In the same line of thought, the authors propose a new feature‐based stereo matching method, which consists of four basic steps – feature‐based stereo correspondence, two‐pass cost aggregation, disparity computation using winner‐takes‐all selection and finally, the disparity refinement. In the proposed method, local features of Gabor wavelet in spatial domain are used for matching cost computation and subsequently a cost aggregation step is implemented by combined use of the Kuwahara filter and the median filter. Experimental results on the Middlebury benchmark database shows that the proposed method outperforms many existing local stereo matching methods. T. Malathi, Manas Kamal Bhuyan |
IET Comput. Vis. | 2 |
| 2014 | Neural Network Based Image Modification for Shape from Observed SEM ImagesabstractA new approach to recover 3-D shape from a Scanning Electron Microscope (SEM) image is described. With an ideal SEM image, 3-D shape can be recovered using the Fast Marching Method (FMM) applied to the Eikonal equation. However, when the light source direction is oblique, the correct shape cannot be obtained by the usual one-pass FMM. The new approach modifies the intensities in the original SEM image using an additional SEM image of a sphere and Neural Network (NN) training. Image modification is a two degree-of-freedom (DOF) rotation. No assumption is made about the specific functional form for intensity in an SEM image. The correct 3-D shape can be obtained using the FMM and NN learning, without iteration. The approach is demonstrated through computer simulation and validated through real experiment. Yuji Iwahori, Kenji Funahashi, Robert J. Woodham, Manas Kamal Bhuyan |
ICPR | 4 |
| 2014 | Automatic Polyp Detection using DSC Edge Detector and HOG FeaturesabstractEndoscopy is a very powerful technology to examine the intestinal tract and to detect the presence of any
possible abnormalities like polyps, the main cause of cancer. This paper presents an edge based method for
polyp detection in endoscopic video images. It utilizes discrete singular convolution (DSC) algorithm for edge
detection/segmentation scheme, then by using conic fitting techniques (ellipse and hyperbola) potential candidates
are determined. These candidates are first rotated so as to make major axis in the x-axis direction, and
then classified as polyp or non-polyp by SVM classifier which is trained separately for ellipse and hyperbola
with HOG features. Himanshu Agrahari, Yuji Iwahori, Manas Kamal Bhuyan, Somnath Ghorai, Himanshu Kohli, Robert J. Woodham, Kunio Kasugai |
ICPRAM | 3 |
| 2014 | Defect Classification of Electronic Circuit Board Using SVM based on Random SamplingabstractThis paper proposes a new approach to improve the classification accuracy of true defect and pseudo defect of electronic circuit board. The proposed approach introduces the defect detection which corresponds to the color image and concept of random sampling using multiple SVMs. The approach first detects the defect candidate region with high accuracy based on the difference between test image and reference image, then extracts the features to recognize true or pseudo defect. Data and features for multiple subsets based on random sampling and feature selection is applied to find the effective combination of features. Selected combination of features are used for the recognition by each SVM and weighted voting process is applied to determine the final discrimination. Computer experiments were demonstrated and the usefulness of the proposed approach is evaluated with the accuracy of defect classification. Hiroaki Hagi, Yuji Iwahori, Shinji Fukui 0001, Yoshinori Adachi, Manas Kamal Bhuyan |
KES | 5 |
| 2014 | Shadow Detection by Three Shadow Models with Features Robust to Illumination ChangesabstractComputer vision methods need to deal with shadows explicitly because shadows often have a negative effect on the results computed. A new shadow detection method is proposed. The new method constructs three shadow models. Three features robust to illumination changes are used to construct the models. The method uses color information, Peripheral Increment Sign Correlation image and edge information. Each of these features removes shadow effects, in part. The overall method can construct an effective shadow model by using all of the features. The result is improved further by region based analysis and by online update of the shadow model. The proposed method extracts shadows accurately. Results are demonstrated by experiments using the real videos of outdoor scenes. Shuya Ishida, Shinji Fukui 0001, Yuji Iwahori, Manas Kamal Bhuyan, Robert J. Woodham |
KES | 4 |
| 2013 | Tracking Method in Consideration of Existence of Similar Object around Target ObjectabstractTracking methods based on the particle filter uses frequently the appearance information of the target object to calculate the likelihood. The method using it often fails in tracking when the target object intersects with other objects with similar appearances. We propose a new approach for tracking objects with similar patterns in a video sequence taken by a moving camera. The proposed method based on the particle filter is robust to the intersection with other objects. Two state transition functions are defined for robust tracking. The method changes the function depending on the situation. In addition, the likelihood is calculated by using four factors which are the information of the color, the velocity, the distance between the objects and the values calculated by the probability background model. The method detects objects which are similar to the target object and which exist around the target object. This prevents the method from tracking other object mistakenly. Results are demonstrated by experiments using real video sequences. Gaku Watanabe, Shinji Fukui 0001, Yuji Iwahori, Manas Kamal Bhuyan, Robert J. Woodham, Yoshinori Adachi |
KES | 4 |
| 2011 | An Integrated Approach to the Recognition of a Wide Class of Continuous Hand GesturesabstractThe gesture segmentation is a method that distinguishes meaningful gestures from unintentional movements. Gesture segmentation is a prerequisite stage to continuous gesture recognition which locates the start and end points of a gesture in an input sequence. Yet, this is an extremely difficult task due to both the multitude of possible gesture variations in spatio-temporal space and the co-articulation/movement epenthesis of successive gestures. In this paper, we focus our attention on coping with this problem associated with continuous gesture recognition. This requires gesture spotting that distinguishes meaningful gestures from co-articulation and unintentional movements. In our method, we first segment the input video stream by detecting gesture boundaries at which the hand pauses for a while during gesturing. Next, every segment is checked for movement epenthesis and co-articulation via finite state machine (FSM) matching or by using hand motion information. Thus, movement epenthesis phases are detected and eliminated from the sequence and we are left with a set of isolated gestures. Finally, we apply different recognition schemes to identify each individual gesture in the sequence. Our experimental results show that the proposed scheme is suitable for recognition of continuous gestures having different spatio-temporal behavior. Manas Kamal Bhuyan, Prabin Kumar Bora |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2008 | Tracking of persons for video surveillance of unattended environmentsabstractThis paper describes a visual surveillance system for remote monitoring of unattended environments. For the purpose of efficiently tracking multiple people in the presence of occlusions, we propose: (i) to combine blob matching with particle filtering, and (ii) to augment these tracking algorithms with a novel colour appearance model. The proposed system efficiently counteracts the shortcomings of the two algorithms by switching from one to the other during occlusions. Results on public datasets as well as real surveillance videos from a metropolitan railway station demonstrate the efficacy of the proposed system. Suyu Kong, Manas Kamal Bhuyan, Conrad Sanderson, Brian C. Lovell |
ICPR | 2 |
| 2006 | Hand motion tracking and trajectory matching for dynamic hand gesture recognitionabstractHand gesture recognition finds applications in areas like human computer interaction, machine vision, virtual reality and so on. In this article, we present a vision-based method for recognizing dynamic hand gestures via hand motion tracking and trajectory matching. A model-based approach based on Hausdorff distance is used for tracking hand motion thereby estimating the gesture trajectories. Dynamic Time Warping technique is employed for gesture trajectory time alignment and normalization. Recognition is done by extracting trajectory information like trajectory length, location, orientation and hand velocity from the estimated trajectory. Experimental results confirm the appropriateness of our proposed trajectory features and demonstrate that our proposed trajectory estimator and trajectory matching-based gesture classifier are efficient enough for use in Human Computer Interaction system. Manas Kamal Bhuyan, Prabin Kumar Bora |
J. Exp. Theor. Artif. Intell. | 1 |