VLDB 2026 Research / reviewers in the wild / expert
Lanyu Shang
dblp:234/2928
· DBLP profile ↗
27ranked-venue papers in the field
10as first author
22since 2021 · last 2025
0000-0002-7480-6889ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 13 (6 first)Big Data, Cloud & Distributed Data Systems · 8 (2 first)Data Mining & Knowledge Discovery · 6 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Structure-Aware Content Classification on Social Networks via Diffusion Graph Learning and LLMs
Cameron Hajaliloo, Cameron Scolari, Sophie Kadifa, Lanyu Shang, Junyuan Lin |
IEEE Big Data | 4 |
| 2024 | ClimateMiSt: Climate Change Misinformation and Stance Detection Dataset
YeonJung Choi, Lanyu Shang, Dong Wang 0002 |
ASONAM (2) | 2 |
| 2024 | A Domain Adaptive Graph Learning Framework to Early Detection of Emergent Healthcare Misinformation on Social MediaabstractA fundamental issue in healthcare misinformation detection is the lack of timely resources (e.g., medical knowledge, annotated data), making it challenging to accurately detect emergent healthcare misinformation at an early stage. In this paper, we develop a crowdsourcing-based early healthcare misinformation detection framework that jointly exploits the medical expertise of expert crowd workers and adapts the medical knowledge from a source domain (e.g., COVID-19) to detect misleading posts in an emergent target domain (e.g., Mpox, Polio). Two important challenges exist in developing our solution: (i) How to leverage the complex and noisy knowledge from the source domain to facilitate the detection of misinformation in the target domain? (ii) How to effectively utilize the limited amount of expert workers to correct the inapplicable knowledge facts in the source domain and adapt the corrected facts to examine the truthfulness of the posts in the emergent target domain? To address these challenges, we develop CrowdAdapt, a crowdsourcing-based domain adaptive approach that effectively identifies and adapts relevant knowledge facts from the source domain to accurately detect misinformation in the target domain. Evaluation results from two real-world case studies demonstrate the superiority of CrowdAdapt over state-of-the-art baselines in accurately detecting emergent healthcare misinformation. Lanyu Shang, Yang Zhang 0031, Zhenrui Yue, YeonJung Choi, Huimin Zeng 0001, Dong Wang 0002 |
ICWSM | 1 |
| 2024 | SocialDrought: A Social and News Media Driven Dataset and Analytical Platform towards Understanding Societal Impact of DroughtabstractDrought poses significant challenges to sustainability across various sectors in our society, leading to substantial consequences on agriculture, environments, ecosystems, public health, and socioeconomic stability. While prior work has studied the impacts of drought using professionally measured data sources, the societal perspectives of drought impacts remain largely under-explored. In this work, we present SocialDrought, a novel and comprehensive dataset to facilitate research on the societal impacts of drought. In particular, SocialDrought consists of three major components: 1) over 1.5 million social media posts, 2) over 1,400 news articles collected and verified by domain experts, and 3) over 31,000 meteorological records from the U.S. Drought Monitor about drought severity. In addition, we also introduce an online analytical platform that enables interactive and real-time data exploration to gain timely insights into the societal impacts of drought. Our interdisciplinary dataset integrates both conventional meteorological data and unconventional social and news media data to provide a holistic understanding of drought impacts. SocialDrought opens new opportunities to study the societal impacts of drought through the lens of social and news media. Lanyu Shang, Bozhang Chen, Anav Vora, Yang Zhang 0031, Ximing Cai, Dong Wang 0002 |
ICWSM | 1 |
| 2024 | MMAdapt: A Knowledge-guided Multi-source Multi-class Domain Adaptive Framework for Early Health Misinformation DetectionabstractThis paper studies a critical problem of emergent health misinformation detection, aiming to mitigate the spread of misinformation in emergent health domains to support well-informed healthcare decisions towards a Web for good health. Our work is motivated by the lack of timely resources (e.g., medical knowledge, annotated data) during the initial phases of an emergent health event or topic. In this paper, we develop a multi-source domain adaptive framework that jointly exploits medical knowledge and annotated data from different high-resource source domains (e.g., cancer, COVID-19) to detect misleading posts in an emergent target domain (e.g., mpox, polio). Two important challenges exist in developing our solution: 1) how to accurately detect the partially misleading and unverifiable content in an emergent target domain? 2) How to identify the conflicting knowledge facts from different source domains to accurately detect emergent misinformation in the target domain? To address these challenges, we develop MMAdapt, a multi-source multi-class domain adaptive misinformation detection framework that effectively explores diverse knowledge facts from different source domains to accurately detect not only the outright misleading but also the partially misleading or unverifiable posts on the Web. Extensive experimental results on four real-world misinformation datasets demonstrate that MMAdapt substantially outperforms state-of-the-art baselines in accurately detecting misinformation in an emergent health domain. Lanyu Shang, Yang Zhang 0031, Bozhang Chen, Ruohan Zong, Zhenrui Yue, Huimin Zeng 0001, Na Wei 0001, Dong Wang 0002 |
WWW | 1 |
| 2024 | SymLearn: A Symbiotic Crowd-AI Collective Learning Framework to Web-based Healthcare Policy Adherence AssessmentabstractThis paper develops a symbiotic human-AI collective learning framework that explores the complementary strengths of both AI and crowdsourced human intelligence to address a novel Web-based healthcare-policy-adherence assessment (WebHA) problem. In particular, the objective of the WebHA problem is to automatically assess people's public health policy adherence during emergent global health crisis events (e.g., COVID-19, MonkeyPox) by exploring massive social media imagery data. Recent advances in human-AI systems exhibit a significant potential in addressing the intricate imagery-based classification problems like WebHA by leveraging the collective intelligence of both humans and AI. This paper aims to address the limitation of existing human-AI systems that often rely heavily on human intelligence to improve AI model performance while overlooking the fact that humans themselves can be fallible and prone to errors. To address the above limitation, this paper develops SymLearn, a symbiotic human-AI co-learning framework that leverages human intelligence to troubleshoot and fine-tune the AI model while using AI models to guide human crowd workers to reduce the inherent human errors in their labels. Extensive experiments on two real-world WebHA applications show that SymLearn clearly outperforms the state-of-the-art baselines by improving WebHA performance and reducing crowd response delay. Yang Zhang 0031, Ruohan Zong, Lanyu Shang, Huimin Zeng 0001, Zhenrui Yue, Dong Wang 0002 |
WWW | 3 |
| 2023 | Manipulating Out-Domain Uncertainty Estimation in Deep Neural Networks via Targeted Clean-Label PoisoningabstractRobust out-domain uncertainty estimation has gained growing attention for its capacity of providing adversary-resistant uncertainty estimates on out-domain samples. However, existing work on robust uncertainty estimation mainly focuses on evasion attacks that happen during test time. The threat of poisoning attacks against uncertainty models is largely unexplored. Compared to evasion attacks, poisoning attacks do not necessarily modify test data, and therefore, would be more practical in real-world applications. In this work, we systematically investigate the robustness of state-of-the-art uncertainty estimation algorithms against data poisoning attacks, with the ultimate objective of developing robust uncertainty training methods. In particular, we focus on attacking the out-domain uncertainty estimation. Under the proposed attack, the training process of models is affected. A fake high-confidence region is established around the targeted out-domain sample, which originally would have been rejected by the model due to low confidence. More fatally, our attack is clean-label and targeted: it leaves the poisoned data with clean labels and attacks a specific targeted test sample without degrading the overall model performance. We evaluate the proposed attack on several image benchmark datasets and a real-world application of COVID-19 misinformation detection. The extensive experimental results on different tasks suggest that the state-of-the-art uncertainty estimation methods could be extremely vulnerable and easily corrupted by our proposed attack. Huimin Zeng 0001, Zhenrui Yue, Yang Zhang 0031, Lanyu Shang, Dong Wang 0002 |
CIKM | 4 |
| 2023 | CollabEquality: A Crowd-AI Collaborative Learning Framework to Address Class-wise Inequality in Web-based Disaster ResponseabstractWeb-based disaster response (WebDR) is emerging as a pervasive approach to acquire real-time situation awareness of disaster events by collecting timely observations from the Web (e.g., social media). This paper studies a class-wise inequality problem in WebDR applications where the objective is to address the limitation of current WebDR solutions that often have imbalanced classification performance across different classes. To address such a limitation, this paper explores the collaborative strengths of the diversified yet complementary biases of AI and crowdsourced human intelligence to ensure a more balanced and accurate performance for WebDR applications. However, two critical challenges exist: 1) it is difficult to identify the imbalanced AI results without knowing the ground-truth WebDR labels a priori; ii) it is non-trivial to address the class-wise inequality problem using potentially imperfect crowd labels. To address the above challenges, we develop CollabEquality, an inequality-aware crowd-AI collaborative learning framework that carefully models the inequality bias of both AI and human intelligence from crowdsourcing systems into a principled learning framework. Extensive experiments on two real-world WebDR applications demonstrate that CollabEquality consistently outperforms the state-of-the-art baselines by significantly reducing class-wise inequality while improving the WebDR classification accuracy. Yang Zhang 0031, Lanyu Shang, Ruohan Zong, Huimin Zeng 0001, Zhenrui Yue, Dong Wang 0002 |
WWW | 2 |
| 2023 | ContrastFaux: Sparse Semi-supervised Fauxtography Detection on the Web using Multi-view Contrastive LearningabstractThe widespread misinformation on the Web has raised many concerns with serious societal consequences. In this paper, we study a critical type of online misinformation, namely fauxtography, where the image and associated text of a social media post jointly convey a questionable or false sense. In particular, we focus on a sparse semi-supervised fauxtography detection problem, which aims to accurately identify fauxtography by only using the sparsely annotated ground truth labels of social media posts. Our problem is motivated by the key limitation of current fauxtography detection approaches that often require a large amount of expensive and inefficient manual annotations to train an effective fauxtography detection model. We identify two key technical challenges in solving the problem: 1) it is non-trivial to train an accurate detection model given the sparse fauxtography annotations, and 2) it is difficult to extract the heterogeneous and complicated fauxtography features from the multi-modal social media posts for accurate fauxtography detection. To address the above challenges, we propose ContrastFaux, a multi-view contrastive learning framework that jointly explores the sparse fauxtography annotations and the cross-modal fauxtography feature similarity between the image and text in multi-modal posts to accurately detect fauxtography on social media. Evaluation results on two social media datasets demonstrate that ContrastFaux consistently outperforms state-of-the-art deep learning and semi-supervised learning fauxtography detection baselines by achieving the highest fauxtography detection accuracy. Ruohan Zong, Yang Zhang 0031, Lanyu Shang, Dong Wang 0002 |
WWW | 3 |
| 2022 | A Knowledge-driven Domain Adaptive Approach to Early Misinformation Detection in an Emergent Health Domain on Social MediaabstractThis paper focuses on an important problem of early misinformation detection in an emergent health domain on social media. Current misinformation detection solutions often suffer from the lack of resources (e.g., labeled datasets, sufficient medical knowledge) in the emerging health domain to accurately identify online misinformation at an early stage. To address such a limitation, we develop a knowledge-driven domain adaptive approach that explores a good set of annotated data and reliable knowledge facts in a source domain (e.g., COVID-19) to learn the domain-invariant features that can be adapted to detect misinformation in the emergent target domain with little ground truth labels (e.g., Monkeypox). Two critical challenges exist in developing our solution: i) how to leverage the noisy knowledge facts in the source domain to obtain the medical knowledge related to the target domain? ii) How to adapt the domain discrepancy between the source and target domains to accurately assess the truthfulness of the social media posts in the target domain? To address the above challenges, we develop KAdapt, a knowledge-driven domain adaptive early misinformation detection framework that explicitly extracts rel-evant knowledge facts from the source domain and jointly learns the domain-invariant representation of the social media posts and their relevant knowledge facts to accurately identify misleading posts in the target domain. Evaluation results on five real-world datasets demonstrate that KAdapt significantly outperforms state-of-the-art baselines in terms of accurately detecting misleading Monkeypox posts on social media. Lanyu Shang, Yang Zhang 0031, Zhenrui Yue, YeonJung Choi, Huimin Zeng 0001, Dong Wang 0002 |
ASONAM | 1 |
| 2022 | Unsupervised Domain Adaptation for COVID-19 Information Service with Contrastive Adversarial Domain MixupabstractIn the real-world application of COVID-19 misinformation detection, a fundamental challenge is the lack of the labeled COVID data to enable supervised end-to-end training of the models, especially at the early stage of the pandemic. To address this challenge, we propose an unsupervised domain adaptation framework using contrastive learning and adversarial domain mixup to transfer the knowledge from an existing source data domain to the target COVID-19 data domain. In particular, to bridge the gap between the source domain and the target domain, our method reduces a radial basis function (RBF) based discrepancy between these two domains. Moreover, we leverage the power of domain adversarial examples to establish an intermediate domain mixup, where the latent representations of the input text from both domains could be mixed during the training process. Extensive experiments on multiple real-world datasets suggest that our method can effectively adapt misinformation detection systems to the unseen COVID-19 target domain with significant improvements compared to the state-of-the-art baselines. Huimin Zeng 0001, Zhenrui Yue, Ziyi Kou, Lanyu Shang, Yang Zhang 0031, Dong Wang 0002 |
ASONAM | 4 |
| 2022 | Boosting Demographic Fairness of Face Attribute Classifiers via Latent Adversarial RepresentationsabstractModern machine learning (ML) is one of the prevailing tools for big data applications of face attribute recognition. However, due to the commonly observed imbalanced distribution of the training data, well-trained models could suffer severely from undesired performance bias across different demographic groups. Motivated by the fact that neural networks could be extremely sensitive to adversarial examples, we argue that there exists the possibility of properly leveraging adversarial examples to address the imbalanced data distribution, and guiding the training convergence towards the direction of improved fairness. That is, we propose to use adversarial examples to alleviate the performance bias issue from the origin: the data source. In this paper, we present a novel adversarial training framework that generates adversarial features in the latent space to automatically balance the distribution of training features and adjust the deep classification layers of the face attribute classifiers to be more fair. Extensive experimental results on the CelebA face dataset show that our method is able to boost the model fairness more effectively compared to the state-of-the-art adversarial debiasing algorithms. Huimin Zeng 0001, Zhenrui Yue, Lanyu Shang, Yang Zhang 0031, Dong Wang 0002 |
IEEE Big Data | 3 |
| 2022 | Contrastive Domain Adaptation for Early Misinformation Detection: A Case Study on COVID-19abstractDespite recent progress in improving the performance of misinformation detection systems, classifying misinformation in an unseen domain remains an elusive challenge. To address this issue, a common approach is to introduce a domain critic and encourage domain-invariant input features. However, early misinformation often demonstrates both conditional and label shifts against existing misinformation data (e.g., class imbalance in COVID-19 datasets), rendering such methods less effective for detecting early misinformation. In this paper, we propose contrastive adaptation network for early misinformation detection (CANMD). Specifically, we leverage pseudo labeling to generate high-confidence target examples for joint training with source data. We additionally design a label correction component to estimate and correct the label shifts (i.e., class priors) between the source and target domains. Moreover, a contrastive adaptation loss is integrated in the objective function to reduce the intra-class discrepancy and enlarge the inter-class discrepancy. As such, the adapted model learns corrected class priors and an invariant conditional distribution across both domains for improved estimation of the target data distribution. To demonstrate the effectiveness of the proposed CANMD, we study the case of COVID-19 early misinformation detection and perform extensive experiments using multiple real-world datasets. The results suggest that CANMD can effectively adapt misinformation detection systems to the unseen COVID-19 target domain with significant improvements compared to the state-of-the-art baselines. Zhenrui Yue, Huimin Zeng 0001, Ziyi Kou, Lanyu Shang, Dong Wang 0002 |
CIKM | 4 |
| 2022 | Defending Substitution-Based Profile Pollution Attacks on Sequential RecommendersabstractWhile sequential recommender systems achieve significant improvements on capturing user dynamics, we argue that sequential recommenders are vulnerable against substitution-based profile pollution attacks. To demonstrate our hypothesis, we propose a substitution-based adversarial attack algorithm, which modifies the input sequence by selecting certain vulnerable elements and substituting them with adversarial items. In both untargeted and targeted attack scenarios, we observe significant performance deterioration using the proposed profile pollution algorithm. Motivated by such observations, we design an efficient adversarial defense method called Dirichlet neighborhood sampling. Specifically, we sample item embeddings from a convex hull constructed by multi-hop neighbors to replace the original items in input sequences. During sampling, a Dirichlet distribution is used to approximate the probability distribution in the neighborhood such that the recommender learns to combat local perturbations. Additionally, we design an adversarial training method tailored for sequential recommender systems. In particular, we represent selected items with one-hot encodings and perform gradient ascent on the encodings to search for the worst case linear combination of item embeddings in training. As such, the embedding function learns robust item representations and the trained recommender is resistant to test-time adversarial examples. Extensive experiments show the effectiveness of both our attack and defense methods, which consistently outperform baselines by a significant margin across model architectures and datasets. Zhenrui Yue, Huimin Zeng 0001, Ziyi Kou, Lanyu Shang, Dong Wang 0002 |
RecSys | 4 |
| 2022 | Can I only share my eyes? A Web Crowdsourcing based Face Partition Approach Towards Privacy-Aware Face RecognitionabstractHuman face images represent a rich set of visual information for online social media platforms to optimize the machine learning (ML)/AI models in their data-driven facial applications (e.g., face detection, face recognition). However, there exists a growing privacy concern from social media users to share their online face images that will be annotated by unknown crowd workers and analyzed by ML/AI researchers in the model training and optimization process. In this paper, we focus on a privacy-aware face recognition problem where the goal is to empower the facial applications to train their face recognition models with images shared by social media users while protecting the identity of the users. Our problem is motivated by the limitation of current privacy-aware face recognition approaches that mainly prevent algorithmic attacks by manipulating face images but largely ignore the potential privacy leakage related to human activities (e.g., crowdsourcing annotation). To address such limitations, we develop FaceCrowd, a web crowdsourcing based face partition approach to improve the performance of current face recognition models by designing a novel crowdsourced partial face graph generated from privacy-preserved social media face images. We evaluate the performance of FaceCrowd using two real-world human face datasets that consist of large-scale human face images. The results show that FaceCrowd not only improves the accuracy of the face recognition models but also effectively protects the identity information of the social media users who share their face images. Ziyi Kou, Lanyu Shang, Yang Zhang 0031, Siyu Duan, Dong Wang 0002 |
WWW | 2 |
| 2022 | A Duo-generative Approach to Explainable Multimodal COVID-19 Misinformation DetectionabstractThis paper focuses on a critical problem of explainable multimodal COVID-19 misinformation detection where the goal is to accurately detect misleading information in multimodal COVID-19 news articles and provide the reason or evidence that can explain the detection results. Our work is motivated by the lack of judicious study of the association between different modalities (e.g., text and image) of the COVID-19 news content in current solutions. In this paper, we present a generative approach to detect multimodal COVID-19 misinformation by investigating the cross-modal association between the visual and textual content that is deeply embedded in the multimodal news content. Two critical challenges exist in developing our solution: 1) how to accurately assess the consistency between the visual and textual content of a multimodal COVID-19 news article? 2) How to effectively retrieve useful information from the unreliable user comments to explain the misinformation detection results? To address the above challenges, we develop a duo-generative explainable misinformation detection (DGExplain) framework that explicitly explores the cross-modal association between the news content in different modalities and effectively exploits user comments to detect and explain misinformation in multimodal COVID-19 news articles. We evaluate DGExplain on two real-world multimodal COVID-19 news datasets. Evaluation results demonstrate that DGExplain significantly outperforms state-of-the-art baselines in terms of the accuracy of multimodal COVID-19 misinformation detection and the explainability of detection explanations. Lanyu Shang, Ziyi Kou, Yang Zhang 0031, Dong Wang 0002 |
WWW | 1 |
| 2022 | SAT-Geo: A social sensing based content-only approach to geolocating abnormal traffic events using syntax-based probabilistic learning
Lanyu Shang, Yang Zhang 0031, Christina Youn, Dong Wang 0002 |
Inf. Process. Manag. | 1 |
| 2021 | A deep contrastive learning approach to extremely-sparse disaster damage assessment in social sensingabstractSocial sensing has emerged as a pervasive and scalable sensing paradigm to obtain timely information of the physical world from "human sensors". In this paper, we study a new extremely-sparse disaster damage assessment (DBA) problem in social sensing. The objective is to automatically assess the damage severity of affected areas in a disaster event by leveraging the imagery data reported on online social media with extremely sparse training data (e.g., only 1% of the data samples have labels). Our problem is motivated by the limitation of current DDA solutions that often require a significant amount of high-quality training data to learn an effective DDA model. We identify two critical challenges in solving our problem: i) it remains to be a fundamental challenge on how to effectively train a reliable DDA model given the lack of sufficient damage severity labels; ii) it is a difficult task to capture the excessive and fine-grained damage-related features in each image for accurate damage assessment. In this paper, we propose ContrastDDA, a deep contrastive learning approach to address the extremely-sparse DDA problem by designing an integrated contrastive and augmentative neural network architecture for accurate disaster damage assessment using the extremely sparse training samples. The evaluation results on two real-world DDA applications demonstrate that ContrastDDA clearly outperforms state-of-the-art deep learning and semi-supervised learning baselines with the highest DDA accuracy under different application scenarios. Yang Zhang 0031, Ruohan Zong, Lanyu Shang, Ziyi Kou, Dong Wang 0002 |
ASONAM | 3 |
| 2021 | ExgFair: A Crowdsourcing Data Exchange Approach To Fair Human Face Datasets AugmentationabstractHuman face images represent a rich set of visual data information that is utilized by various big data driven human facial applications. However, the performance of these applications is usually biased towards the majority demographic group due to the data imbalance issue. In this paper, we focus on a fair human face data exchange problem where the goal is to exchange visual features of human face images between different human face datasets and obtain a set of augmented datasets that improve the fairness and performance of human facial applications. Our problem is motivated by the limitations of current fairness approaches that only focus on a single human face dataset from a particular application and require a large amount of pre-annotated demographic attribute labels to develop fair human facial models. To address these limitations, we develop ExgFair, a crowdsourcing-based fair data exchange framework to generate a set of augmented fair face image datasets by leveraging the crowdsourced demographic attribute labels of human face images. We evaluate ExgFair using a set of real-world human face image datasets with different demographic distributions. The results show that ExgFair not only reduces demographic biases of the datasets but also improves the accuracy of human facial applications trained on the augmented fair datasets. Ziyi Kou, Lanyu Shang, Huimin Zeng 0001, Yang Zhang 0031, Dong Wang 0002 |
IEEE BigData | 2 |
| 2021 | A Multimodal Misinformation Detector for COVID-19 Short Videos on TikTokabstractThis paper studies an emerging and important problem of identifying misleading COVID-19 short videos where the misleading content is jointly expressed in the visual, audio, and textual content of videos. Existing solutions for misleading video detection mainly focus on the authenticity of videos or audios against AI algorithms (e.g., deepfake) or video manipulation, and are insufficient to address our problem where most videos are user-generated and intentionally edited. Two critical challenges exist in solving our problem: i) how to effectively extract information from the distractive and manipulated visual content in TikTok videos? ii) How to efficiently aggregate heterogeneous information across different modalities in short videos? To address the above challenges, we develop TikTec, a multimodal misinformation detection framework that explicitly exploits the captions to accurately capture the key information from the distractive video content, and effectively learns the composed misinformation that is jointly conveyed by the visual and audio content. We evaluate TikTec on a real-world COVID- 19 video dataset collected from TikTok. Evaluation results show that TikTec achieves significant performance gains compared to state-of-the-art baselines in accurately detecting misleading COVID-19 short videos. Lanyu Shang, Ziyi Kou, Yang Zhang 0031, Dong Wang 0002 |
IEEE BigData | 1 |
| 2021 | StreamCollab: A Streaming Crowd-AI Collaborative System to Smart Urban Infrastructure Monitoring in Social SensingabstractSocial sensing has emerged as a pervasive and scalable sensing paradigm to collect observations of the physical world from human sensors. A key advantage of social sensing is its infrastructure-free nature. In this paper, we focus on a streaming urban infrastructure monitoring (Streaming UIM) problem in social sensing. The goal is to automatically detect the urban infrastructure damages from the streaming imagery data posted on social media by exploring the collective power of both AI and human intelligence from crowdsourcing systems. Our work is motivated by the limitation of current AI and crowdsourcing solutions that either fail in many critical time-sensitive UIM application scenarios or are not easily generalizable to monitor the damage of different types of urban infrastructures. We identify two critical challenges in solving our problem: i) it is difficult to dynamically integrate AI and crowd intelligence to effectively identify and fix the failure cases of AI solutions; ii) it is non-trivial to obtain accurate human intelligence from unreliable crowd workers in streaming UIM applications. In this paper, we propose StreamCollab, a streaming crowd-AI collaborative system that explores the collaborative intelligence from AI and crowd to solve the streaming UIM problem. The evaluation results on a real-world urban infrastructure imagery dataset collected from social media demonstrate that StreamCollab consistently outperforms both state-of-the-art AI and crowd-AI baselines in UIM accuracy while maintaining the lowest computational cost. Yang Zhang 0031, Lanyu Shang, Ruohan Zong, Ziyi Kou, Dong Wang 0002 |
HCOMP | 2 |
| 2021 | AOMD: An analogy-aware approach to offensive meme detection on social media
Lanyu Shang, Yang Zhang 0031, Yuheng Zha, Yingxi Chen, Christina Youn, Dong Wang 0002 |
Inf. Process. Manag. | 1 |
| 2020 | CaMR: Towards Connotation-aware Music Retrieval on Social Media with Visual InputsabstractWith the ubiquitous network connectivity and the proliferation of mobile devices, people are increasingly consuming digital contents from social media driven music sharing platforms (e.g., YouTube, Soundcloud). In this paper, we study a novel problem of connotation-aware music retrieval that focuses on the connotation which expresses the implicit feeling or emotion beyond the explicit content in artworks. Our goal is to automatically retrieve relevant music on social media based on the connotation of visual inputs (e.g., images, photos) provided by the users. The problem is challenging as it requires the accurate identification of the implicit connotation from both images and music pieces, and the precise matching of the identified connotation across different data modalities. We develop a connotation-aware music retrieval (CaMR) framework to address the above challenges. Evaluation results from a real-world social media dataset demonstrate that the CaMR framework can retrieve music that is highly relevant to the connotation of the input image. Lanyu Shang, Daniel Yue Zhang, Siamul Karim Khan, Jialie Shen 0001, Dong Wang 0002 |
ASONAM | 1 |
| 2020 | ExFaux: A Weakly Supervised Approach to Explainable Fauxtography DetectionabstractFauxtography is a category of multi-modal posts that spreads misleading information on various online social platforms (e.g., Facebook, Twitter, Reddit). A fauxtography post usually consists of an image, a text description and comments from its readers. In this paper, we focus on an explainable fauxtography detection problem where the goal is to explain which a specific component of a post leads to the fauxtography decision. This problem is motivated by the limitations of current fauxtography detection solutions that only focus on the detection but ignore the important explanation aspect of their results. Two critical challenges exist in solving our problem: i) it is difficult to accurately identify the "guilty" component of a fauxtography post given the fact that different components of the post and their associations could all lead to the fauxtography; ii) it is expensive and time-consuming to obtain a good training set with fine-grained labels of fauxtography posts in terms of explainability, making the corresponding solutions weakly supervised in nature. To address the above challenges, we develop ExFaux, an end-to-end graph-based fauxtography explanation framework, to effectively explain which part of the post contributes to its fauxtography. We evaluate the ExFaux by creating a real-world dataset from online social media (Twitter and Reddit). The results show that ExFaux not only detects the fauxtography posts more accurately than the state-of-the-arts but also provides well-justified explanations to its results. Ziyi Kou, Daniel Yue Zhang, Lanyu Shang, Dong Wang 0002 |
IEEE BigData | 3 |
| 2019 | VulnerCheck: A Content-Agnostic Detector for Online Hatred-Vulnerable VideosabstractWith the increasing popularity of online video platforms (e.g., YouTube, Vimeo), the spread of hateful videos and the lack of rigorous hateful content control have become a critical issue. This paper focuses on the problem of identifying online hatred-vulnerable videos where the videos themselves do not contain any hateful content but unexpectedly trigger hateful comments from the audience. It is suboptimal to simply treat the hatred-vulnerable videos as hateful ones and remove them from the sharing platforms. This will discourage the uploaders of such videos from sharing valid and informative videos in the future. However, treating these hatred-vulnerable videos as hatred-free ones will provide undesirable opportunities for hateful users to spread their toxic comments and extreme ideology. In this paper, we develop VulnerCheck, an end-to-end supervised learning approach to effectively classify hatred-vulnerable videos from hateful and hatred-free ones by exploring the structure and semantics features of audience's comment networks. VulnerCheck is content-agnostic in the sense that it does not analyze the content of the video and is therefore robust against sophisticated content creators who craft hateful videos to bypass the current content censorship. We evaluate VulnerCheck on a real-world dataset collected from YouTube. Results demonstrate that our scheme is both effective and efficient in identifying hatred-vulnerable videos and significantly outperforms the state-of-the-art baselines. Lanyu Shang, Daniel Yue Zhang, Dong Wang 0002 |
IEEE BigData | 1 |
| 2018 | RiskSens: A Multi-view Learning Approach to Identifying Risky Traffic Locations in Intelligent Transportation Systems Using Social and Remote SensingabstractWith the ever-increasing number of road traffic accidents worldwide, the road traffic safety has become a critical problem in intelligent transportation systems. A key step towards improving the road traffic safety is to identify the locations where severe traffic accidents happen with a high probability so the precautions can be applied effectively. We refer to this problem as risky traffic location identification. While previous efforts have been made to address similar problems, two important limitations exist: i) data availability: many cities (especially in developing countries) do not maintain a publicly accessible database for the traffic accident records in a city, which makes it difficult to accurately estimate the accidents in the city; ii) location accuracy: many self-reported traffic accidents (e.g., social media posts from common citizens) are not associated with the exact GPS locations due to the privacy concerns. To address these limitations, this paper develops the RiskSens, a multi-view learning approach to identifying the risky traffic locations in a city by jointly exploring the social and remote sensing data. We evaluate RiskSens using a real world dataset from New York. The evaluation results show that RiskSens significantly outperforms the state-of-the- art baselines in identifying risky traffic locations in a city. Yang Zhang 0031, Daniel Yue Zhang, Lanyu Shang, Dong Wang 0002 |
IEEE BigData | 4 |
| 2018 | FauxBuster: A Content-free Fauxtography Detector Using Social Media CommentsabstractWith the increasing popularity of online social media (e.g., Facebook, Twitter, Reddit), the detection of misleading content on social media has become a critical undertaking. This paper focuses on an important but largely unsolved problem: detecting fauxtography (i.e., social media posts with misleading images). We found that the existing literature falls short in solving this problem. In particular, current solutions either focus on the detection of fake images or misinformed texts of a social media post. However, they cannot solve our problem because the detection of fauxtography depends not only on the truthfulness of the images and the texts but also on the information they deliver together on the posts. In this paper, we develop the FauxBuster, an end-to-end supervised learning scheme that can effectively track down fauxtography by exploring the valuable clues from user's comments of a post on social media. The FauxBuster is content-free in that it does not rely on the analysis of the actual content of the images, and hence is robust against malicious uploaders who can intentionally modify the presentation and description of the images. We evaluate FauxBuster on real-world data collected from two mainstream social media platforms - Reddit and Twitter. Results show that our scheme is both effective and efficient in addressing the fauxtography problem. Daniel Yue Zhang, Lanyu Shang, Biao Geng, Shuyue Lai, Hongmin Zhu, Md. Tanvir Al Amin, Dong Wang 0002 |
IEEE BigData | 2 |