T. Sree Sharmila

dblp:139/6560 · also Sree Sharmila Thangaswamy · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0001-5744-9739ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Attention-enhanced dynamic graph neural networks for efficient video summarization
R. Deepa 0004, T. Sree Sharmila, Niruban Rathakrishnan
Expert Syst. Appl.2
2026 Integrating 3D convolutional neural network and Bidirectional Encoder Representations from Transformers for effective video summarization
abstract
Video summarisation, a critical task in multimedia analysis, aims to condense lengthy videos into concise representations while preserving essential content. As the volume of video data grows exponentially, efficient and informative summarisation becomes increasingly crucial. Striking the right balance between preserving essential content and adhering to length constraints remains an emerging field of study. In this context, we explore the fusion of 3D VGG16 (Visual Geometry Group 16-layer network) and BERT (Bidirectional Encoder Representations from Transformers) for video summarisation (3D-BERSUM). This method offers comprehensive coverage by leveraging 3D VGG16 to extract visual features and BERT to understand textual context, ensuring that the generated summaries capture both the visual elements of the footage and the key information conveyed through narration. BERT’s contextual understanding enhances the summarisation process by grasping nuances, relationships, and sentiment in the text, enriching the summary with a deeper understanding of the narrative. By integrating visual and textual information, we generate summaries that convey more than just visual scenes. To assess the performance of 3D-BERSUM, we compared it with existing methods using the content coverage measures (recall, precision, and F1-score) and textual quality metrics (ROUGE and BLEU). The findings indicate that the 3D-BERSUM yields enhancements in recall, precision, and F1-score by 5.6%, 3.9%, and 5.2% for the SumMe dataset, and 10.1%, 7.2%, and 6.1% for the TVSum dataset, respectively. Finally, comparison is performed with other state-of-the-art methods, and the result shows that the superiority of 3D-BERSUM. This underscores the efficacy and potential of our method in achieving enhanced video summarisation performance across diverse datasets and scenarios.
R. Deepa 0004, T. Sree Sharmila, Niruban Rathakrishnan
J. Exp. Theor. Artif. Intell.2
2024 Dynamic graph neural network-based computational paradigm for video summarization
R. Deepa 0004, T. Sree Sharmila, R. Niruban
Multim. Tools Appl.2
2024 An efficient Moving object, Encryption, Compression and Interpolation technique for video steganography
R. Roselin Kiruba, T. Sree Sharmila, J. K. Josephine Julina
Multim. Tools Appl.2
2023 A novel data hiding by image interpolation using edge quad-tree block complexity
R. Roselin Kiruba, T. Sree Sharmila
Vis. Comput.2
2022 An efficient hand gesture recognition based on optimal deep embedded hybrid convolutional neural network-long short term memory network model
abstract
Abstract Hand gestures are the nonverbal communication done by individuals who cannot represent their thoughts in form of words. It is mainly used during human‐computer interaction (HCI), deaf and mute people interaction, and other robotic interface applications. Gesture recognition is a field of computer science mainly focused on improving the HCI via touch screens, cameras, and kinetic devices. The state‐of‐art systems mainly used computer vision‐based techniques that utilize both the motion sensor and camera to capture the hand gestures in real‐time and interprets them via the usage of the machine learning algorithms. Conventional machine learning algorithms often suffer from the different complexities present in the visible hand gesture images such as skin color, distance, light, hand direction, position, and background. In this article, an adaptive weighted multi‐scale resolution (AWMSR) network with a deep embedded hybrid convolutional neural network and long short term memory network (hybrid CNN‐LSTM) is proposed for identifying the different hand gesture signs with higher recognition accuracy. The proposed methodology is formulated using three steps: input preprocessing, feature extraction, and classification. To improve the complex visual effects present in the input images, a histogram equalization technique is used which improves the size of the gray level pixel in the image and also their occurrence probability. The multi‐block local binary pattern (MB‐LBP) algorithm is employed for feature extraction which extracts the crucial features present in the image such as hand shape structure feature, curvature feature, and invariant movements. The AWMSR with the deep embedded hybrid CNN–LSTM network is applied in the two‐benchmark datasets namely Jochen Triesch static hand posture and NUS hand posture dataset‐II to detect its stability in identifying different hand gestures. The weight function of the deep embedded CNN‐LSTM architecture is optimized using the puzzle optimization algorithm. The efficiency of the proposed methodology is verified in terms of different performance evaluation metrics such as accuracy, loss, confusion matrix, Intersection over the union, and execution time. The proposed methodology offers recognition accuracy of 97.86% and 98.32% for both datasets.
Gajalakshmi Palanisamy, T. Sree Sharmila
Concurr. Comput. Pract. Exp.2
2022 Optimum anamorphic image generation using image rotation and relative entropy
Pavithra Latha Kumaresan, T. Sree Sharmila
Multim. Tools Appl.3
2022 Degenerative disc disease diagnosis from lumbar MR images using hybrid features
Beulah A, T. Sree Sharmila, V. K. Pramod
Vis. Comput.2
2020 An Improved Seed Point Selection-Based Unsupervised Color Clustering for Content-Based Image Retrieval Application
abstract
Abstract The images involved in the content-based image retrieval (CBIR) applications are collectively represented by features such as color, texture and shape. The precision of the CBIR application relies on the key features used in image representation and its similarity measure. In CBIR, dominant color feature extraction is affected by the predefined intervals used in color quantization. The proposed work mainly concentrates on extracting the dominant color information of the image using the clustering process. The clustering process is initiated by the proposed seed point’s selection approach. This approach derives the number of seed points using the first order statistical measure and maximum range of the distributed pixel values. Moreover, this work gives equal priority to dominant color and its occurrence information in calculating the similarity between query and database images. Finally, the standard databases such as SIMPLIcity, Corel-10k, OT-scene, Oxford flower and GHIM are taken to investigate the performance of the proposed dominant color based image retrieval application.
Pavithra Latha Kumaresan, T. Sree Sharmila
Comput. J.2
2019 DICENet: Fine-Grained Recognition via Dilated Iterative Contextual Encoding
abstract
Material Recognition is an intriguing problem in Computer Vision. While traditional approaches prefer an ensemble of networks to capture essential properties such as texture, more recent approaches leverage the power of Deep Learning to design end-to-end models. We do the same, and propose Dilated Iterative Contextual Encoding Network, a novel end-to-end framework for material recognition. As a result of gathering extensive knowledge on various characteristics of materials, our approach combines different components on the base network to address specific properties, which also helps in general recognition tasks. The traditional ResNet is replaced by a Dilated Residual Network to help capture fine-grained material information. Iterative deep aggregation helps capture and fuse global homogeneous material properties across multiple resolutions and scales. To enhance the discriminatory power of the learnt latent representation, we propose gramedial loss which is intuitively applied on a texture vector space. Spatial similarity loss is applied on strategic intermediate feature maps to effectively capture local non-homogeneous texture features from a global context, crucial for the primary classification task. Extensive experiments conducted on golden material datasets such as the FMD, MINC-2500, KTH-TIPS-2b, DTD and GTOS indicate improved performances over state of the art approaches on large datasets and two small datasets, while achieving compatible accuracies on the challenging FMD. Furthermore, our architecture also performed convincingly while categorizing general indoor and object classification datasets such as MITIndoor and CalTech-101.
Abhishek Pal, Gautham Krishnan, Manav Rajiv Moorthy, Narasimha Yadav, Adithya R. Ganesh, T. Sree Sharmila
IJCNN6
2018 Disc bulge diagnostic model in axial lumbar MR images using Intervertebral disc Descriptor (IdD)
Beulah A, T. Sree Sharmila, V. K. Pramod
Multim. Tools Appl.2
2018 Real time blink recognition from various head pose using single eye
Sofia Jennifer John, T. Sree Sharmila
Multim. Tools Appl.2
2018 Acoustic image enhancement using Gaussian and laplacian pyramid - a multiresolution based technique
Priyadharsini Ravisankar, T. Sree Sharmila, V. Rajendran
Multim. Tools Appl.2