EDBT 2026 Demo / reviewers in the wild / expert
Cara Jones
dblp:200/7933
· DBLP profile ↗
5ranked-venue papers
0as first author
3since 2021 · last 2024
0000-0001-9112-4751ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Web and social media mining · 66% Information retrieval · 26% Data mining · 8% | |
| Artificial intelligence
2 papers |
Graph learning · 72% Information extraction and text analysis · 28% |
Topics — the 3 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Web and social media mining
bot detection |
0.5 | 1 | 2021 | INFOSHIELD: Generalizable Information-Theoretic Human-Trafficking Detection · ICDE 2021 |
Information retrieval › similarity search
near-duplicate detection |
0.5 | 1 | 2021 | INFOSHIELD: Generalizable Information-Theoretic Human-Trafficking Detection · ICDE 2021 |
Data mining › clustering
document clustering |
0.1 | 1 | 2021 | INFOSHIELD: Generalizable Information-Theoretic Human-Trafficking Detection · ICDE 2021 |
Methods — techniques the papers use, named apart from their topics
weak supervision · 1.5graph neural network · 1.5text and image classification · 0.6deep multimodal model · 0.6parameter-free clustering · 0.5information-theoretic clustering · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | T-NET: Weakly Supervised Graph Learning for Combatting Human TraffickingabstractHuman trafficking (HT) for forced sexual exploitation, often described as modern-day slavery, is a pervasive problem that affects millions of people worldwide. Perpetrators of this crime post advertisements (ads) on behalf of their victims on adult service websites (ASW). These websites typically contain hundreds of thousands of ads including those posted by independent escorts, massage parlor agencies and spammers (fake ads). Detecting suspicious activity in these ads is difficult and developing data-driven methods is challenging due to the hard-to-label, complex and sensitive nature of the data. In this paper, we propose T-Net, which unlike previous solutions, formulates this problem as weakly supervised classification. Since it takes several months to years to investigate a case and obtain a single definitive label, we design domain-specific signals or indicators that provide weak labels. T-Net also looks into connections between ads and models the problem as a graph learning task instead of classifying ads independently. We show that T-Net outperforms all baselines on a real-world dataset of ads by 7% average weighted F1 score. Given that this data contains personally identifiable information, we also present a realistic data generator and provide the first publicly available dataset in this domain which may be leveraged by the wider research community. Pratheeksha Nair, Javin Liu, Catalina Vajiac, Andreas M. Olligschlaeger, Polo Chau, Mirela Teixeira Cazzolato, Cara Jones, Christos Faloutsos, Reihaneh Rabbany |
AAAI | 7 |
| 2023 | DeltaShield: Information Theory for Human- Trafficking DetectionabstractGiven a million escort advertisements, how can we spot near-duplicates? Such micro-clusters of ads are usually signals of human trafficking (HT). How can we summarize them to convince law enforcement to act? Spotting micro-clusters of near-duplicate documents is useful in multiple, additional settings, including spam-bot detection in Twitter ads, plagiarism, and more. We present InfoShield , which makes the following contributions: practical , being scalable and effective on real data; parameter-free and principled , requiring no user-defined parameters; interpretable , finding a document to be the cluster representative, highlighting all the common phrases, and automatically detecting “slots” (i.e., phrases that differ in every document); and generalizable , beating or matching domain-specific methods in Twitter bot detection and HT detection, respectively, as well as being language independent. Interpretability is particularly important for the anti-HT domain, where law enforcement must visually inspect ads. Our experiments on real data show that InfoShield correctly identifies Twitter bots with an F1 score over 90% and detects HT ads with 84% precision. Moreover, it is scalable, requiring about 8 hours for 4 million documents on a stock laptop. Our incremental version, DeltaShield , allows for fast, incremental updates, with minor loss of accuracy. Catalina Vajiac, Meng-Chieh Lee, Aayushi Kulshrestha, Sacha Levy, Namyong Park 0001, Andreas M. Olligschlaeger, Cara Jones, Reihaneh Rabbany, Christos Faloutsos |
ACM Trans. Knowl. Discov. Data | 7 |
| 2021 | INFOSHIELD: Generalizable Information-Theoretic Human-Trafficking DetectionabstractGiven a million escort advertisements, how can we spot near-duplicates? Such micro-clusters of ads are usually signals of human trafficking. How can we summarize them, visually, to convince law enforcement to act? Can we build a general tool that works for different languages? Spotting micro-clusters of near-duplicate documents is useful in multiple, additional settings, including spam-bot detection in Twitter ads, plagiarism, and more.We present INFOSHIELD, which makes the following contributions: (a) Practical, being scalable and effective on real data, (b) Parameter-free and Principled, requiring no user-defined parameters, (c) Interpretable, finding a document to be the cluster representative, highlighting all the common phrases, and automatically detecting "slots", i.e. phrases that differ in every document; and (d) Generalizable, beating or matching domain-specific methods in Twitter bot detection and human trafficking detection respectively, as well as being language-independent finding clusters in Spanish, Italian, and Japanese. Interpretability is particularly important for the anti human-trafficking domain, where law enforcement must visually inspect ads.Our experiments on real data show that INFOSHIELD correctly identifies Twitter bots with an F1 score over 90% and detects human-trafficking ads with 84% precision. Moreover, it is scalable, requiring about 8 hours for 4 million documents on a stock laptop. Meng-Chieh Lee, Catalina Vajiac, Aayushi Kulshrestha, Sacha Levy, Namyong Park 0001, Cara Jones, Reihaneh Rabbany, Christos Faloutsos |
ICDE | 6 |
| 2018 | Detection and Characterization of Human Trafficking Networks Using Unsupervised Scalable Text Template MatchingabstractHuman trafficking is a form of modern-day slavery affecting an estimated 40 million victims worldwide, primarily through the commercial sexual exploitation of women and children. In the last decade, the advertising of victims has moved from the streets to websites on the Internet, providing greater efficiency and anonymity for sex traffickers. This shift has allowed traffickers to list their victims in multiple geographic areas simultaneously, while also improving operational security by using multiple methods of electronic communication with buyers; complicating the ability of law enforcement to disrupt these illicit organizations. In this paper, we address this issue and present a novel unsupervised and scalable template matching algorithm for analyzing and detecting complex organizations operating on adult service websites. The algorithm uses only the advertisement content to uncover signature patterns in text that are indicative of organized activities and organizational structure. We apply this method to a large corpus of adult service advertisements retrieved from backpage.com, and show that the networks identified through the algorithm match well with surrogate truth data derived from phone number networks in the same corpus. Further exploration of the results show that the proposed method provides deeper insights into the complex structures of sex trafficking organizations, not possible through networks derived from phone numbers alone. This method provides a powerful new capability for law enforcement to more completely identify and gather evidence about trafficking networks and their operations. Lin Li 0005, Olga Simek, Angela Lai, Matthew P. Daggett, Charlie K. Dagli, Cara Jones |
IEEE BigData | 6 |
| 2017 | Combating Human Trafficking with Multimodal Deep ModelsabstractHuman trafficking is a global epidemic affecting millions of people across the planet.Sex trafficking, the dominant form of human trafficking, has seen a significant rise mostly due to the abundance of escort websites, where human traffickers can openly advertise among at-will escort advertisements.In this paper, we take a major step in the automatic detection of advertisements suspected to pertain to human trafficking.We present a novel dataset called Trafficking-10k, with more than 10,000 advertisements annotated for this task.The dataset contains two sources of information per advertisement: text and images.For the accurate detection of trafficking advertisements, we designed and trained a deep multimodal model called the Human Trafficking Deep Network (HTDN). Edmund Tong, Amir Zadeh 0001, Cara Jones, Louis-Philippe Morency |
ACL (1) | 3 |