EDBT 2026 Demo / reviewers in the wild / expert
Satadisha Saha Bhowmick
dblp:312/6480
· DBLP profile ↗
4ranked-venue papers
4as first author
4since 2021 · last 2023
0000-0002-7442-0839ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Globally Aware Contextual Embeddings for Named Entity Recognition in Social Media StreamsabstractAn important task for Information Extraction from Microblogs is Named Entity Recognition (NER) that extracts mentions of real-world entities from microblog messages and meta-information like entity type for better entity characterization. A lot of microblog NER systems have rightly sought to prioritize modeling the non-literary nature of microblog text. These systems are trained on offline static datasets and extract a combination of surface-level features – orthographic, lexical, and semantic – from individual messages for noisy text modeling and entity extraction. But given the constantly evolving nature of microblog streams, detecting all entity mentions from such varying yet limited context in short messages remains a difficult problem to generalize. In this paper, we propose the NER Globalizer pipeline better suited for NER on microblog streams. It characterizes the isolated message processing by existing NER systems as modeling local contextual embeddings, where learned knowledge from the immediate context of a message is used to suggest seed entity candidates. Additionally, it recognizes that messages within a microblog stream are topically related and often repeat mentions of the same entity. This suggests building NER systems that go beyond localized processing. By leveraging occurrence mining, the proposed system therefore follows up traditional NER modeling by extracting additional mentions of seed entity candidates that were previously missed. Candidate mentions are separated into well-defined clusters which are then used to generate a pooled global embedding drawn from the collective context of the candidate within a stream. The global embeddings are utilized to separate false positives from entities whose mentions are produced in the final NER output. Our experiments show that the proposed NER system exhibits superior effectiveness on multiple NER datasets with an average Macro F1 improvement of 47.04% over the best NER baseline while adding only a small computational overhead. Satadisha Saha Bhowmick, Eduard C. Dragut, Weiyi Meng |
ICDE | 1 |
| 2023 | TwiCS: Lightweight Entity Mention Detection in Targeted Twitter StreamsabstractMicroblogging sites, like Twitter, continuously generate a large volume of streaming data. This streaming environment creates new challenges for two concomitant Information Extraction tasks: Entity Mention Detection (EMD) and Entity Detection (ED). The new challenges include (1) continuously evolving topics, which may deprecate model-based approaches quickly; (2) non-literary nature of posts, which makes traditional NLP techniques less effective; and (3) huge volume of streaming data, which makes computationally expensive approaches less suitable. In this paper, we propose an approach for EMD/ED whose creation is guided by the constraints specific to streaming environments from the ground up. Our system TwiCS implements this approach. TwiCS employs a computationally light two-phase process. In the first phase, it exploits simple (low computation) syntactic cues to suggest Entity Mention (EM) candidates. In the second phase, it uses occurrence mining to classify candidates according to their likelihood of being true EMs. Our experiments show that TwiCS achieves an average effectiveness improvement of 14.6%, while maintaining at least 2.64 times higher throughput, when compared to several state-of-the-art systems. Satadisha Saha Bhowmick, Eduard C. Dragut, Weiyi Meng |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Boosting Entity Mention Detection for Targetted Twitter Streams with Global Contextual EmbeddingsabstractMicroblogging sites, like Twitter, have emerged as ubiquitous sources of information. Two important tasks related to the automatic extraction and analysis of information in Microblogs are Entity Mention Detection (EMD) and Entity Detection (ED). The state-of-the-art EMD systems aim to model the non-literary nature of microblog text by training upon offline static datasets. They extract a combination of surface-level features - orthographic, lexical, and semantic - from individual messages for noisy text modeling and entity extraction. But given the constantly evolving nature of microblog streams, detecting all entity mentions from such varying yet limited context of short messages remains a difficult problem. To this end, we propose a framework named EMD Globalizer, better suited for the execution of EMD learners on microblog streams. It deviates from the processing of isolated microblog messages by existing EMD systems, where learned knowledge from the immediate context of a message is used to suggest entities. Instead, it recognizes that messages within a microblog stream are topically related and often repeat entity mentions, thereby leaving the scope for EMD systems to go beyond the localized processing of individual messages. After an initial extraction of entity candidates by an EMD system, the proposed framework leverages occurrence mining to find additional candidate mentions that are missed during this first detection. Aggregating the local contextual representations of these mentions, a global embedding is drawn from the collective context of an entity candidate within a stream. The global embeddings are then utilized to separate entities within the candidates from false positives. All mentions of said entities from the stream are produced in the framework's final outputs. Our experiments show that EMD Globalizer can enhance the effectiveness of all existing EMD systems that we tested (on average by 25.61 %) with a small additional computational overhead. Satadisha Saha Bhowmick, Eduard C. Dragut, Weiyi Meng |
ICDE | 1 |
| 2022 | TwiCS: Twitter Stream Entity Mention Detection (Extended Abstract)abstractIn this paper, we propose a system TwiCS for Entity Mention Detection (EMD) and Entity Detection (ED) in streaming environments. TwiCS employs a computationally light two-phase process: (1) exploit simple (low computation) syntactic cues to suggest Entity Mention (EM) candidates and (2) use occurrence mining to classify candidates according to their likelihood of being true EMs. Our experiments show that on average TwiCS improves effectiveness by 14.6%, while achieving at least 2.64 times higher throughput, when compared to several state-of-the-art systems. Satadisha Saha Bhowmick, Eduard C. Dragut, Weiyi Meng |
ICDE | 1 |