EDBT 2026 Demo / reviewers in the wild / expert
Subhajit Chaudhury
dblp:143/3842
· DBLP profile ↗
30ranked-venue papers
12as first author
21since 2021 · last 2026
0000-0003-3435-2584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMsabstractAbinitha Gourabathina, Inkit Padhi, Manish Nagireddy, Subhajit Chaudhury, Prasanna Sattigeri. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Abinitha Gourabathina, Inkit Padhi, Manish Nagireddy, Subhajit Chaudhury, Prasanna Sattigeri |
ACL (1) | 4 |
| 2026 | ImReasoner: Improving Memory-based Language Models for Reasoning-in-a-Haystack TasksabstractChing-Yun Ko, Payel Das, Sihui Dai, Georgios Kollias, Subhajit Chaudhury, Aurelie C. Lozano, Pin-Yu Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ching Yun Ko, Sihui Dai, Georgios Kollias, Subhajit Chaudhury, Aurélie C. Lozano |
ACL (1) | 5 |
| 2026 | ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value RetrievalabstractDavid H. Yang, Yuxuan Zhu, Mohammad Mohammadi Amiri, Keerthiram Murugesan, Tejaswini Pedapati, Subhajit Chaudhury, Pin-Yu Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. David H. Yang, Yuxuan Zhu 0004, Mohammad Mohammadi Amiri, Keerthiram Murugesan, Tejaswini Pedapati, Subhajit Chaudhury |
ACL (1) | 6 |
| 2025 | EpMAN: Episodic Memory AttentioN for Generalizing to Longer ContextsabstractRecent advances in Large Language Models (LLMs) have yielded impressive successes on many language tasks. However, efficient processing of long contexts using LLMs remains a significant challenge. We introduce EpMAN – a method for processing long contexts in an episodic memory module while holistically attending to semantically-relevant context chunks. Output from episodic attention is then used to reweigh the decoder’s self-attention to the stored KV cache of the context during training and generation. When an LLM decoder is trained using EpMAN, its performance on multiple challenging single-hop long-context recall and question-answering benchmarks is found to be stronger and more robust across the range from 16k to 256k tokens than baseline decoders trained with self-attention, and popular retrieval-augmented generation frameworks. Subhajit Chaudhury, Sarathkrishna Swaminathan, Georgios Kollias, Elliot Nelson, Khushbu Pahwa, Tejaswini Pedapati, Igor Melnyk, Matthew Riemer |
ACL (1) | 1 |
| 2025 | On the Effects of Fine-tuning Language Models for Text-Based Reinforcement LearningabstractText-based reinforcement learning involves an agent interacting with a fictional environment using observed text and admissible actions in natural language to complete a task. Previous works have shown that agents can succeed in text-based interactive environments even in the complete absence of semantic understanding or other linguistic capabilities. The success of these agents in playing such games suggests that semantic understanding may not be important for the task. This raises an important question about the benefits of LMs in guiding the agents through the game states. In this work, we show that rich semantic understanding leads to efficient training of text-based RL agents. Moreover, we describe the occurrence of semantic degeneration as a consequence of inappropriate fine-tuning of language models in text-based reinforcement learning (TBRL). Specifically, we describe the shift in the semantic representation of words in the LM, as well as how it affects the performance of the agent in tasks that are semantically similar to the training games. These results may help develop better strategies to fine-tune agents in text-based RL scenarios. Maurício Gruppi, Soham Dan, Keerthiram Murugesan, Subhajit Chaudhury |
COLING | 4 |
| 2025 | TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery Over Data LakesabstractEnterprises have a growing need to identify relevant tables in data lakes; e.g. tables that are unionable, joinable, or subsets of each other. Tabular neural models can be help-ful for such data discovery tasks. In this paper, we present TabSketchFM, a neural tabular model for data discovery over data lakes. First, we propose novel pre-training: a sketch-based approach to enhance the effectiveness of data discovery in neural tabular models. Second, we finetune the pretrained model for identifying unionable, joinable, and subset table pairs and show significant improvement over previous tabular neural models. Third, we present a detailed ablation study to highlight which sketches are crucial for which tasks. Fourth, we use these finetuned models to perform table search; i.e., given a query table, find other tables in a corpus that are unionable, joinable, or that are subsets of the query. Our results demonstrate significant improvements in F1 scores for search compared to state-of-the-art techniques. Finally, we show significant transfer across datasets and tasks establishing that our model can generalize across different tasks and over different data lakes. Aamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury, Julian Dolby, Oktie Hassanzadeh, Zhenhan Huang, Tejaswini Pedapati, Horst Samulowitz, Kavitha Srinivas |
ICDE | 4 |
| 2025 | Large Language Models can Become Strong Self-DetoxifiersabstractReducing the likelihood of generating harmful and toxic output is an essential task when aligning large language models (LLMs). Existing methods mainly rely on training an external reward model (i.e., another language model) or fine-tuning the LLM using self-generated data to influence the outcome. In this paper, we show that LLMs have the capability of self-detoxification without external reward model learning or retraining of the LM. We propose \textit{Self-disciplined Autoregressive Sampling (SASA)}, a lightweight controlled decoding algorithm for toxicity reduction of LLMs. SASA leverages the contextual representations from an LLM to learn linear subspaces from labeled data characterizing toxic v.s. non-toxic output in analytical forms. When auto-completing a response token-by-token, SASA dynamically tracks the margin of the current output to steer the generation away from the toxic subspace, by adjusting the autoregressive sampling strategy. Evaluated on LLMs of different scale and nature, namely Llama-3.1-Instruct (8B), Llama-2 (7B), and GPT2-L models with the RealToxicityPrompts, BOLD, and AttaQ benchmarks, SASA markedly enhances the quality of the generated sentences relative to the original models and attains comparable performance to state-of-the-art detoxification techniques, significantly reducing the toxicity level by only using the LLM's internal representations. Ching Yun Ko, Youssef Mroueh, Soham Dan, Georgios Kollias, Subhajit Chaudhury, Tejaswini Pedapati, Luca Daniel |
ICLR | 7 |
| 2024 | API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMsabstractKinjal Basu, Ibrahim Abdelaziz, Subhajit Chaudhury, Soham Dan, Maxwell Crouse, Asim Munawar, Vernon Austel, Sadhana Kumaravel, Vinod Muthusamy, Pavan Kapanipathi, Luis Lastras. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kinjal Basu 0002, Ibrahim Abdelaziz, Subhajit Chaudhury, Soham Dan, Maxwell Crouse, Asim Munawar, Vernon Austel, Sadhana Kumaravel, Vinod Muthusamy, Pavan Kapanipathi, Luis A. Lastras |
ACL (1) | 3 |
| 2024 | EXPLORER: Exploration-guided Reasoning for Textual Reinforcement LearningabstractKinjal Basu, Keerthiram Murugesan, Subhajit Chaudhury, Murray Campbell, Kartik Talamadupula, Tim Klinger. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Kinjal Basu 0002, Keerthiram Murugesan, Subhajit Chaudhury, Murray Campbell, Kartik Talamadupula, Tim Klinger |
EACL (1) | 3 |
| 2024 | Leveraging Visual Handicaps for Text-Based Reinforcement LearningabstractWe introduce VisualHandicaps, a novel benchmark environment for the systematic analysis of interactive text-based reinforcement learning (TBRL) agents by providing visual handicaps. Unlike previous TBRL environments, which focus on providing additional textual information to measure agent understanding of sequential natural language information, VisualHandicaps seeks to improve the generalization ability of RL agents using varying details of maps and textual information, allowing for the study and demonstration of robust planning and self-localization. We provide automatically generated variations and difficulty levels in our environment and show that an agent using our systematic visual handicaps along with textual observation generally outperforms previous methods (that use only textual handicaps) in terms of success rate and the number of steps required to reach the goal. We also provide a detailed analysis of each handicap, which we believe to be important findings for driving future improvements in RL agents on text-based applications. Subhajit Chaudhury, Keerthiram Murugesan, Thomas Carta, Kartik Talamadupula, Michiaki Tatsubori |
ICASSP | 1 |
| 2024 | Adversarial Robustness of Convolutional Models Learned in the Frequency DomainabstractThis paper presents an extensive comparison of the noise robustness of standard Convolutional Neural Networks (CNNs) trained on image inputs and those trained in the frequency domain. We investigate the robustness of CNNs to small adversarial noise in the RGB input space and show that CNNs trained on Discrete Cosine Transform (DCT) inputs exhibit significantly better noise robustness to both adversarial and common spatial transformations compared to standard CNNs learned on RGB/Grayscale input. Our results suggest that frequency-domain learning of convolutional models may disentangle frequencies corresponding to semantic and adversarial features, resulting in improved adversarial robustness. This research highlights the potential of frequency domain learning to improve neural network robustness to test-time noise and warrants further investigation in this area. Subhajit Chaudhury, Toshihiko Yamasaki |
ICASSP | 1 |
| 2024 | Variance Reduction Can Improve Trade-Off in Multi-Objective LearningabstractMany machine learning problems today have multiple objective functions, which are often tackled by the multi-objective learning (MOL) framework. Albeit many encouraging results are obtained by MOL algorithms, a recent theoretical study [1] revealed that these gradient-based MOL methods (e.g., MGDA, CAGrad) all reflect an inherent trade-off between optimization convergence speeds and conflict-avoidance abilities. To this end, we develop an improved stochastic variance-reduced multi-objective gradient correction method for MOL, achieving the ${\mathcal{O}}\left({{\varepsilon ^{ - 1.5}}}\right)$ sample complexity. In addition, our proposed method simultaneously improves the theoretical guarantees for conflict avoidance and convergence rate compared to prior stochastic gradient-based MOL methods in the non-convex setting. We further validate the effectiveness of the proposed method empirically using popular multi-task learning (MTL) benchmarks. Heshan Devaka Fernando, Lisha Chen, Songtao Lu, Miao Liu 0001, Subhajit Chaudhury, Keerthiram Murugesan, Gaowen Liu, Meng Wang 0003, Tianyi Chen 0002 |
ICASSP | 6 |
| 2024 | Larimar: Large Language Models with Episodic Memory ControlabstractEfficient and accurate updating of knowledge stored in Large Language Models (LLMs) is one of the most pressing research challenges today. This paper presents Larimar - a novel, brain-inspired architecture for enhancing LLMs with a distributed episodic memory. Larimar’s memory allows for dynamic, one-shot updates of knowledge without the need for computationally expensive re-training or fine-tuning. Experimental results on multiple fact editing benchmarks demonstrate that Larimar attains accuracy comparable to most competitive baselines, even in the challenging sequential editing setup, but also excels in speed—yielding speed-ups of 8-10x depending on the base LLM —as well as flexibility due to the proposed architecture being simple, LLM-agnostic, and hence general. We further provide mechanisms for selective fact forgetting, information leakage prevention, and input context length generalization with Larimar and show their effectiveness. Our code is available at https://github.com/IBM/larimar. Subhajit Chaudhury, Elliot Nelson, Igor Melnyk, Sarathkrishna Swaminathan, Sihui Dai, Aurélie C. Lozano, Georgios Kollias, Vijil Chenthamarakshan, Jirí Navrátil 0001, Soham Dan |
ICML | 2 |
| 2023 | Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement LearningabstractSubhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen, Keerthiram Murugesan, Rosario Uceda-Sosa, Michiaki Tatsubori, Achille Fokoue, Pavan Kapanipathi, Asim Munawar, Alexander Gray. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Subhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen, Keerthiram Murugesan, Rosario Uceda-Sosa, Michiaki Tatsubori, Achille Fokoue, Pavan Kapanipathi, Asim Munawar, Alexander G. Gray |
ACL (1) | 1 |
| 2023 | Laziness Is a Virtue When It Comes to Compositionality in Neural Semantic ParsingabstractMaxwell Crouse, Pavan Kapanipathi, Subhajit Chaudhury, Tahira Naseem, Ramon Fernandez Astudillo, Achille Fokoue, Tim Klinger. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Maxwell Crouse, Pavan Kapanipathi, Subhajit Chaudhury, Tahira Naseem, Ramón Fernandez Astudillo, Achille Fokoue, Tim Klinger |
ACL (1) | 3 |
| 2023 | Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Approach
Heshan Devaka Fernando, Miao Liu 0001, Subhajit Chaudhury, Keerthiram Murugesan, Tianyi Chen 0002 |
ICLR | 4 |
| 2023 | On the Convergence and Sample Complexity Analysis of Deep Q-Networks with ε-Greedy Exploration
Shuai Zhang 0015, Hongkang Li, Meng Wang 0003, Miao Liu 0001, Songtao Lu, Sijia Liu 0001, Keerthiram Murugesan, Subhajit Chaudhury |
NeurIPS | 9 |
| 2022 | Eye of the Beholder: Improved Relation Generalization for Text-Based Reinforcement Learning AgentsabstractText-based games (TBGs) have become a popular proving ground for the demonstration of learning-based agents that make decisions in quasi real-world settings. The crux of the problem for a reinforcement learning agent in such TBGs is identifying the objects in the world, and those objects' relations with that world. While the recent use of text-based resources for increasing an agent's knowledge and improving its generalization have shown promise, we posit in this paper that there is much yet to be learned from visual representations of these same worlds. Specifically, we propose to retrieve images that represent specific instances of text observations from the world and train our agents on such images. This improves the agent's overall understanding of the game scene and objects' relationships to the world around them, and the variety of visual representations on offer allow the agent to generate a better generalization of a relationship. We show that incorporating such images improves the performance of agents in various TBG settings. Keerthiram Murugesan, Subhajit Chaudhury, Kartik Talamadupula |
AAAI | 2 |
| 2022 | X-FACTOR: A Cross-metric Evaluation of Factual Correctness in Abstractive SummarizationabstractSubhajit Chaudhury, Sarathkrishna Swaminathan, Chulaka Gunasekara, Maxwell Crouse, Srinivas Ravishankar, Daiki Kimura, Keerthiram Murugesan, Ramón Fernandez Astudillo, Tahira Naseem, Pavan Kapanipathi, Alexander Gray. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Subhajit Chaudhury, Sarathkrishna Swaminathan, R. Chulaka Gunasekara, Maxwell Crouse, Srinivas Ravishankar, Daiki Kimura, Keerthiram Murugesan, Ramón Fernandez Astudillo, Tahira Naseem, Pavan Kapanipathi, Alexander G. Gray |
EMNLP | 1 |
| 2021 | Neuro-Symbolic Approaches for Text-Based Policy LearningabstractText-Based Games (TBGs) have emerged as important testbeds for reinforcement learning (RL) in the natural language domain.Previous methods using LSTM-based action policies are uninterpretable and often overfit the training games showing poor performance to unseen test games.We present SymboLic Action policy for Textual Environments (SLATE), that learns interpretable action policy rules from symbolic abstractions of textual observations for improved generalization.We outline a method for end-to-end differentiable symbolic rule learning and show that such symbolic policies outperform previous stateof-the-art methods in text-based RL for the coin collector environment from 5 -10x fewer training games.Additionally, our method provides human-understandable policy rules that can be readily verified for their logical consistency and can be easily debugged.1 Subhajit Chaudhury, Prithviraj Sen, Masaki Ono, Daiki Kimura, Michiaki Tatsubori, Asim Munawar |
EMNLP (1) | 1 |
| 2021 | Neuro-Symbolic Reinforcement Learning with First-Order LogicabstractDaiki Kimura, Masaki Ono, Subhajit Chaudhury, Ryosuke Kohita, Akifumi Wachi, Don Joven Agravante, Michiaki Tatsubori, Asim Munawar, Alexander Gray. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Daiki Kimura, Masaki Ono, Subhajit Chaudhury, Ryosuke Kohita, Akifumi Wachi, Don Joven Agravante, Michiaki Tatsubori, Asim Munawar, Alexander G. Gray |
EMNLP (1) | 3 |
| 2020 | Understanding Generalization in Neural Networks for Robustness against Adversarial VulnerabilitiesabstractNeural networks have contributed to tremendous progress in the domains of computer vision, speech processing, and other real-world applications. However, recent studies have shown that these state-of-the-art models can be easily compromised by adding small imperceptible perturbations. My thesis summary frames the problem of adversarial robustness as an equivalent problem of learning suitable features that leads to good generalization in neural networks. This is motivated from learning in humans which is not trivially fooled by such perturbations due to robust feature learning which shows good out-of-sample generalization. Subhajit Chaudhury |
AAAI | 1 |
| 2020 | Bootstrapped Q-learning with Context Relevant Observation Pruning to Generalize in Text-based GamesabstractSubhajit Chaudhury, Daiki Kimura, Kartik Talamadupula, Michiaki Tatsubori, Asim Munawar, Ryuki Tachibana. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Subhajit Chaudhury, Daiki Kimura, Kartik Talamadupula, Michiaki Tatsubori, Asim Munawar, Ryuki Tachibana |
EMNLP (1) | 1 |
| 2020 | Investigating Generalization in Neural Networks Under Optimally Evolved Training PerturbationsabstractIn this paper, we study the generalization properties of neural networks under input perturbations and show that minimal training data corruption by a few pixel modifications can cause drastic overfitting. We propose an evolutionary algorithm to search for optimal pixel perturbations using novel cost function inspired from literature in domain adaptation that explicitly maximizes the generalization gap and domain divergence between clean and corrupted images. Our method outperforms previous pixel-based data distribution shift methods on state-of-the-art Convolutional Neural Networks (CNNs) architectures. Interestingly, we find that the choice of optimization plays an important role in generalization robustness due to the empirical observation that SGD is resilient to such training data corruption unlike adaptive optimization techniques (ADAM). Subhajit Chaudhury, Toshihiko Yamasaki |
ICASSP | 1 |
| 2020 | Adversarial Discriminative Attention for Robust Anomaly DetectionabstractExisting methods for visual anomaly detection predominantly rely on global level pixel comparisons for anomaly score computation without emphasizing on unique local features. However, images from real-world applications are susceptible to unwanted noise and distractions, that might jeopardize the robustness of such anomaly score. To alleviate this problem, we propose a self-supervised masking method that specifically focuses on discriminative parts of images to enable robust anomaly detection. Our experiments reveal that discriminator's class activation map in adversarial training evolves in three stages and finally fixates on the foreground location in the images. Using this property of the activation map, we construct a mask that suppresses spurious signals from the background thus enabling robust anomaly detection by focusing on local discriminative attributes. Additionally, our method can further improve the accuracy by learning a semi-supervised discriminative classifier in cases where a few samples from anomaly classes are available during the training. Experimental evaluations on four different types of datasets demonstrate that our method outperforms previous state-of-the-art methods for each condition and in all domains. Daiki Kimura, Subhajit Chaudhury, Minori Narita, Asim Munawar, Ryuki Tachibana |
WACV | 2 |
| 2019 | Unsupervised Temporal Feature Aggregation for Event Detection in Unstructured Sports VideosabstractImage-based sports analytics enable automatic retrieval of key events in a game to speed up the analytics process for human experts. However, most existing methods focus on structured television broadcast video datasets with a straight and fixed camera having minimum variability in the capturing pose. In this paper, we study the case of event detection in sports videos for unstructured environments with arbitrary camera angles. The transition from structured to unstructured video analysis produces multiple challenges that we address in our paper. Specifically, we identify and solve two major problems: unsupervised identification of players in an unstructured setting and generalization of the trained models to pose variations due to arbitrary shooting angles. For the first problem, we propose a temporal feature aggregation algorithm using person re-identification features to obtain high player retrieval precision by boosting a weak heuristic scoring method. Additionally, we propose a data augmentation technique, based on multi-modal image translation model, to reduce bias in the appearance of training samples. Experimental evaluations show that our proposed method improves precision for player retrieval from 0.78 to 0.86 for obliquely angled videos. Additionally, we obtain an improvement in F1 score for rally detection in table tennis videos from 0.79 in case of global frame-level features to 0.89 using our proposed player-level features. Please see the supplementary video submission at https://ibm.biz/BdzeZA. Subhajit Chaudhury, Hiroki Ozaki, Daiki Kimura, Phongtharin Vinayavekhin, Asim Munawar, Ryuki Tachibana, Koji Ito, Yuki Inaba, Minoru Matsumoto, Shuji Kidokoro |
ISM | 1 |
| 2019 | Injective State-Image Mapping facilitates Visual Adversarial Imitation LearningabstractThe growing use of virtual autonomous agents in applications like games and entertainment demands better control policies for natural-looking movements and actions. Unlike the conventional approach of hard-coding motion routines, we propose a deep learning method for obtaining control policies by directly mimicking raw video demonstrations. Previous methods in this domain rely on extracting low-dimensional features from expert videos followed by a separate hand-crafted reward estimation step. We propose an imitation learning framework that reduces the dependence on hand-engineered reward functions by jointly learning the feature extraction and reward estimation steps using Generative Adversarial Networks (GANs). Our main contribution in this paper is to show that under injective mapping between low-level joint state (angles and velocities) trajectories and corresponding raw video stream, performing adversarial imitation learning on video demonstrations is equivalent to learning from the state trajectories. Experimental results show that the proposed adversarial learning method from raw videos produces a similar performance to state-of-the-art imitation learning techniques while frequently outperforming existing hand-crafted video imitation methods. Furthermore, we show that our method can learn action policies by imitating video demonstrations on YouTube with similar performance to learned agents from true reward signal. Please see the supplementary video submission at https://ibm.biz/BdzzNA. Subhajit Chaudhury, Daiki Kimura, Asim Munawar, Ryuki Tachibana |
MMSP | 1 |
| 2018 | Transfer Learning from Synthetic to Real Images Using Variational Autoencoders for Precise Position DetectionabstractCapturing and labeling camera images in the real world is an expensive task, whereas synthesizing labeled images in a simulation environment is easy for collecting large-scale image data. However, learning from only synthetic images may not achieve the desired performance in the real world due to a gap between synthetic and real images. We propose a method that transfers learned detection of an object position from a simulation environment to the real world. This method uses only a significantly limited dataset of real images while leveraging a large dataset of synthetic images using variational autoen-coders. Additionally, the proposed method consistently performed well in different lighting conditions, in the presence of other distractor objects, and on different backgrounds. Experimental results showed that it achieved accuracy of 1.5 mm to 3.5 mm on average. Furthermore, we showed how the method can be used in a real-world scenario like a “pick-and-place” robotic task. Tadanobu Inoue, Subhajit Chaudhury, Giovanni De Magistris, Sakyasingha Dasgupta |
ICIP | 2 |
| 2018 | Focusing on What is Relevant: Time-Series Learning and Understanding using AttentionabstractThis paper is a contribution towards interpretability of the deep learning models in different applications of time-series. We propose a temporal attention layer that is capable of selecting the relevant information to perform various tasks, including data completion, key-frame detection and classification. The method uses the whole input sequence to calculate an attention value for each time step. This results in more focused attention values and more plausible visualisation than previous methods. We apply the proposed method to three different tasks. Experimental results show that the proposed network produces comparable results to a state of the art. In addition, the network provides better interpretability of the decision, that is, it generates more significant attention weight to related frames compared to similar techniques attempted in the past. Phongtharin Vinayavekhin, Subhajit Chaudhury, Asim Munawar, Don Joven Agravante, Giovanni De Magistris, Daiki Kimura, Ryuki Tachibana |
ICPR | 2 |
| 2017 | Spatial-Temporal Motion Field Analysis for Pixelwise Crack Detection on Concrete SurfacesabstractCrack development in concrete structures starts at the micro-crack stage and proceeds to the macro-crack stage due to repeated cyclic loading, like ongoing vehicles on bridges. Automatic detection of early stage cracks is required for both safety and economic reasons. We present an automatic crack detection method that scans a captured concrete area and provides a pixel-wise localization of both visible macro-cracks and early stage micro-cracks from video sequences. The key component in the proposed method is a spatial-temporal non-linear filtering on framewise dense 2D motion field combined with Conditional Random Fields based crack localization refinement. We evaluate our method against labeled ground truth data provided by an expert crack inspector. Experimental results show that our method can produce high accuracy automatic crack localization having F1 score improvement of 0.14-0.22 compared to conventional image based detectors. The proposed method is also shown to detect cracks at an earlier stage which enables early preventive measures for repair operations. Subhajit Chaudhury, Gaku Nakano, Jun Takada, Akihiko Iketani |
WACV | 1 |