EDBT 2026 Demo / reviewers in the wild / expert
Chongyang Bai
dblp:241/6973
· DBLP profile ↗
12ranked-venue papers
8as first author
8since 2021 · last 2023
0000-0002-1245-9877ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Representation and self-supervised learning · 35% Graph learning · 21% Face, body and person analysis · 20% | |
| Network and information security
1 paper |
Malware analysis · 77% Web and mobile security · 23% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
User interface design and tools · 50% Accessibility and assistive technology · 50% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
dynamic graph learning |
0.5 | 1 | 2021 | TEDIC: Neural Modeling of Behavioral Patterns in Dynamic Social Interaction Networks · WWW 2021 |
Computer vision › Vision and language › multimodal understanding
GUI understanding |
0.5 | 1 | 2021 | UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021 |
Machine learning › Representation and self-supervised learning › pre-training
multimodal pretraining |
0.5 | 1 | 2021 | UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.5 | 1 | 2021 | UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021 |
Machine learning › Representation and self-supervised learning
pre-training |
0.5 | 1 | 2021 | UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021 |
Web and social media mining › user behavior analysis
social interaction analysis |
0.5 | 1 | 2021 | TEDIC: Neural Modeling of Behavioral Patterns in Dynamic Social Interaction Networks · WWW 2021 |
Malware analysis › mobile malware detection
android malware detection |
0.5 | 1 | 2021 | $\sf {DBank}$DBank: Predictive Behavioral Analysis of Recent Android Banking Trojans · IEEE Trans. Dependable Secur. Comput. 2021 |
Machine learning › Graph learning › graph neural network › node classification
collective classification |
0.4 | 1 | 2019 | Predicting the Visual Focus of Attention in Multi-Person Discussion Videos · IJCAI 2019 |
Computer vision › Video understanding and tracking › video analytics › behavior analysis
multi-person video analysis |
0.4 | 1 | 2019 | Predicting dominance in multi-person videos · IJCAI 2019 |
Computer vision › Face, body and person analysis › gaze analysis
visual focus of attention |
0.4 | 1 | 2019 | Predicting the Visual Focus of Attention in Multi-Person Discussion Videos · IJCAI 2019 |
Multimedia analysis and retrieval
multimodal fusion |
0.2 | 1 | 2023 | M2P2: Multimodal Persuasion Prediction Using Adaptive Fusion · IEEE Trans. Multim. 2023 |
Accessibility and assistive technology
accessibility |
0.1 | 1 | 2021 | UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021 |
User interface design and tools
UI understanding |
0.1 | 1 | 2021 | UIBert: Learning Generic Multimodal Representations for UI Understanding · IJCAI 2021 |
Web and mobile security › mobile security
android security |
0.1 | 1 | 2021 | $\sf {DBank}$DBank: Predictive Behavioral Analysis of Recent Android Banking Trojans · IEEE Trans. Dependable Secur. Comput. 2021 |
Computer vision › 3D vision
multi-person interaction |
0.1 | 1 | 2019 | Predicting the Visual Focus of Attention in Multi-Person Discussion Videos · IJCAI 2019 |
Computer vision › Face, body and person analysis
nonverbal behavior analysis |
0.1 | 1 | 2019 | Predicting the Visual Focus of Attention in Multi-Person Discussion Videos · IJCAI 2019 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.0temporal convolution network · 1.0set pooling · 1.0pre-training tasks · 1.0diffusion convolution · 1.0attention mechanism · 0.7adaptive fusion · 0.7static and dynamic feature analysis · 0.5pagerank generalization · 0.5lightly supervised learning · 0.4facial action units · 0.4ensemble learning · 0.4dominance rank features · 0.4collective classification · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | M2P2: Multimodal Persuasion Prediction Using Adaptive FusionabstractIdentifying persuasive speakers in an adversarial environment is a critical task. In a national election, politicians would like to have persuasive speakers campaign on their behalf. When a company faces adverse publicity, they would like to engage persuasive advocates for their position in the presence of adversaries who are critical of them. Debates represent a common platform for these forms of adversarial persuasion. This paper solves two problems: the Debate Outcome Prediction (DOP) problem predicts who wins a debate while the Intensity of Persuasion Prediction (IPP) problem predicts the change in the number of votes before and after a speaker speaks. Though DOP has been previously studied, we are the first to study IPP. Past studies on DOP fail to leverage two important aspects of multimodal data: 1) multiple modalities are often semantically aligned, and 2) different modalities may provide diverse information for prediction. Our$\mathsf{M2P2}$(Multimodal Persuasion Prediction) framework is the first to use multimodal (acoustic, visual, language) data to solve the IPP problem. To leverage the alignment of different modalities while maintaining the diversity of the cues they provide,$\mathsf{M2P2}$devises a novel adaptive fusion learning framework which fuses embeddings obtained from two modules – analignmentmodule that extracts shared information between modalities and aheterogeneitymodule that learns the weights of different modalities with guidance from three separately trained unimodal reference models. We test$\mathsf{M2P2}$on the popular IQ2US dataset designed for DOP. We also introduce a new dataset called QPS (from Qipashuo, a popular Chinese debate TV show) for IPP.$\mathsf{M2P2}$significantly outperforms 4 recent baselines on both datasets. Chongyang Bai, Haipeng Chen 0001, Srijan Kumar, Jure Leskovec, V. S. Subrahmanian |
IEEE Trans. Multim. | 1 |
| 2023 | DIPS: A Dyadic Impression Prediction System for Group Interaction VideosabstractWe consider the problem of predicting the impression that one subject has of another in a video clip showing a group of interacting people. Our novel Dyadic Impression Prediction System ( DIPS ) contains two major innovations. First, we develop a novel method to align the facial expressions of subjects p i and p j as well as account for the temporal delay that might be involved in p i reacting to p j ’s facial expressions. Second, we propose the concept of a multilayered stochastic network for impression prediction on top of which we build a novel Temporal Delayed Network graph neural network architecture. Our overall DIPS architecture predicts six dependent variables relating to the impression p i has of p j . Our experiments show that DIPS beats eight baselines from the literature, yielding statistically significant improvements of 19.9% to 30.8% in AUC and 12.6% to 47.2% in F1-score. We further conduct ablation studies showing that our novel features contribute to the overall quality of the predictions made by DIPS . Chongyang Bai, Maksim Bolonkin, Viney Regunath, V. S. Subrahmanian |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | POLLY: A Multimodal Cross-Cultural Context-Sensitive Framework to Predict Political Lying from VideosabstractPoliticians lie. Frequently. Depending on the country they are from, politicians may lie more frequently on some topics than others. We develop the novel concept of a tripartite “VAT” graph (Video-Article-Topic) with three types of nodes: videos (with a politician featured in each), news articles that mention the politician, and topics discussed in the videos or articles. We develop several novel types of audio and video deception scores for each audio/video, as well as a topic deception score and an edge deception score for each edge in the graph. Our POLLY (POLitical LYing) system builds upon past work by others to generate predictions for whether a politician is lying or not. We test POLLY on a novel dataset (which will be made publicly available upon publication of this paper) consisting of 146 videos and 6337 news articles involving 73 politicians from 18 countries from all major continents. We show that POLLY achieves AUC and F1 scores over 77%, beating out several baselines. We further show that POLLY is robust to translation errors made by Google Translate. Chongyang Bai, Maksim Bolonkin, Viney Regunath, V. S. Subrahmanian |
ICMI | 1 |
| 2021 | Unimodal Face Classification with Multimodal TrainingabstractFace recognition is a crucial task in various multimedia applications such as security check, credential access and motion sensing games. However, the task is challenging when an input face is noisy (e.g. poor-condition RGB image) or lacks certain information (e.g. 3D face without color). In this work, we propose a Multimodal Training Unimodal Test (MTUT) framework for robust face classification, which exploits the cross-modality relationship during training and applies it as a complementary of the imperfect single modality input during testing. Technically, during training, the framework (1) builds both intra-modality and cross-modality autoencoders with the aid of facial attributes to learn latent embeddings as multimodal descriptors, (2) proposes a novel multimodal embedding divergence loss to align the heterogeneous features from different modalities, which also adaptively avoids the useless modality (if any) from confusing the model. This way, the learned autoencoders can generate robust embeddings in single-modality face classification on test stage. We evaluate our framework in two face classification datasets and two kinds of testing input: (1) poor-condition image and (2) point cloud or 3D face mesh, when both 2D and 3D modalities are available for training. We experimentally show that our MTUT framework consistently outperforms ten baselines on 2D and 3D settings of both datasets11Code can be found in the following url: https://github.com/wbteng9526/mtut_fr. Wenbin Teng, Chongyang Bai |
FG | 2 |
| 2021 | Deception Detection in Group Video Conversations using Dynamic Interaction Networks
Srijan Kumar, Chongyang Bai, V. S. Subrahmanian, Jure Leskovec |
ICWSM | 2 |
| 2021 | UIBert: Learning Generic Multimodal Representations for UI UnderstandingabstractTo improve the accessibility of smart devices and to simplify their usage, building models which understand user interfaces (UIs) and assist users to complete their tasks is critical. However, unique challenges are proposed by UI-specific characteristics, such as how to effectively leverage multimodal UI features that involve image, text, and structural metadata and how to achieve good performance when high-quality labeled data is unavailable. To address such challenges we introduce UIBert, a transformer-based joint image-text model trained through novel pre-training tasks on large-scale unlabeled UI data to learn generic feature representations for a UI and its components. Our key intuition is that the heterogeneous features in a UI are self-aligned, i.e., the image and text features of UI components, are predictive of each other. We propose five pretraining tasks utilizing this self-alignment among different features of a UI component and across various components in the same UI. We evaluate our method on nine real-world downstream UI tasks where UIBert outperforms strong multimodal baselines by up to 9.26% accuracy. Chongyang Bai, Xiaoxue Zang, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, Blaise Agüera y Arcas |
IJCAI | 1 |
| 2021 | TEDIC: Neural Modeling of Behavioral Patterns in Dynamic Social Interaction NetworksabstractDynamic social interaction networks are an important abstraction to model time-stamped social interactions such as eye contact, speaking and listening between people. These networks typically contain informative while subtle patterns that reflect people’s social characters and relationship, and therefore attract the attentions of a lot of social scientists and computer scientists. Previous approaches on extracting those patterns primarily rely on sophisticated expert knowledge of psychology and social science, and the obtained features are often overly task-specific. More generic models based on representation learning of dynamic networks may be applied, but the unique properties of social interactions cause severe model mismatch and degenerate the quality of the obtained representations. Here we fill this gap by proposing a novel framework, termed TEmporal network-DIffusion Convolutional networks (TEDIC), for generic representation learning on dynamic social interaction networks. We make TEDIC a good fit by designing two components: 1) Adopt diffusion of node attributes over a combination of the original network and its complement to capture long-hop interactive patterns embedded in the behaviors of people making or avoiding contact; 2) Leverage temporal convolution networks with hierarchical set-pooling operation to flexibly extract patterns from different-length interactions scattered over a long time span. The design also endows TEDIC with certain self-explaining power. We evaluate TEDIC over five real datasets for four different social character prediction tasks including deception detection, dominance identification, nervousness detection and community detection. TEDIC not only consistently outperforms previous SOTA’s, but also provides two important pieces of social insight. In addition, it exhibits favorable societal characteristics by remaining unbiased to people from different regions. Our project website is: http://snap.stanford.edu/tedic/. Yanbang Wang, Pan Li 0005, Chongyang Bai, Jure Leskovec |
WWW | 3 |
| 2021 | $\sf {DBank}$DBank: Predictive Behavioral Analysis of Recent Android Banking TrojansabstractUsing a novel dataset of Android banking trojans (ABTs), other Android malware, and goodware, we develop the$\sf {DBank}$system to predict whether a given Android APK is a banking trojan or not. We introduce the novel concept of aTriadic Suspicion Graph(TSG for short) which contains three kinds of nodes: goodware, banking trojans, and API packages. We develop a novel feature space based on two classes of scores derived from TSGs:suspicion scores(SUS) andsuspicion ranks(SR)—the latter yields a family of features that generalize PageRank. While TSG features (based on SUS/SR scores) provide very high predictive accuracy on their own in predicting recent (2016-2017) ABTs, we show that the combination of TSG features with previously studied lightweight static and dynamic features in the literature yields the highest accuracy in distinguishing ABTs from goodware, while preserving the same accuracy of prior feature combinations in distinguishing ABTs from other Android malware. In particular,$\sf {DBank}$’s overall accuracy in predicting whether an APK is a banking trojan or not is up to 99.9% AUC with 0.3% false positive rate. Moreover, we have already reported two unlabeled APKs from VirusTotal (which$\sf {DBank}$has detected as ABTs) to the Google Android Security Team—in one case, we discovered it before any of the 63 anti-virus products on VirusTotal did, and in the other case, we beat 62 of 63 anti-viruses on VirusTotal. This suggests that$\sf {DBank}$is capable of making new discoveries in the wild before other established vendors. We also show that our novel TSG features have some interesting defensive properties as they are robust to knowledge of the training set by an adversary: even if the adversary uses 90% of our training set and uses the exact TSG features that we use, it is difficult for him to infer$\sf {DBank}$’s predictions on APKs. We additionally identify the features that best separate and characterize ABTs from goodware as well as from other Android malware. Finally, we develop a detailed data-driven analysis of five major recent ABT families:FakeToken,Svpeng,Asacub,BankBot, andMarcher, and identify the features that best separate them from goodware and other-malware. Chongyang Bai, Qian Han, Ghita Mezzour, Fabio Pierazzi, V. S. Subrahmanian |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2019 | Attention-based Facial Behavior Analytics inSocial Communication
Lezi Wang, Chongyang Bai, Maksim Bolonkin, Judee K. Burgoon, Norah E. Dunbar, V. S. Subrahmanian, Dimitris N. Metaxas |
BMVC | 2 |
| 2019 | Automatic Long-Term Deception Detection in Group Interaction VideosabstractMost work on automated deception detection (ADD) in video has two restrictions: (i) it focuses on a video of one person, and (ii) it focuses on a single act of deception in a one or two minute video. In this paper, we propose a new ADD framework which captures long term deception in a group setting. We study deception in the well-known Resistance game (like Mafia and Werewolf) which consists of 5-8 players of whom 2-3 are spies. Spies are deceptive throughout the game (typically 30-65 minutes) to keep their identity hidden. We develop an ensemble predictive model to identify spies in Resistance videos. We show that features from low-level and high-level video analysis are insufficient, but when combined with a new class of features that we call LiarRank, produce the best results. We achieve AUCs of over 0.70 in a fully automated setting. Chongyang Bai, Maksim Bolonkin, Judee K. Burgoon, Norah E. Dunbar, V. S. Subrahmanian, Zhe Wu 0001 |
ICME | 1 |
| 2019 | Predicting dominance in multi-person videosabstractWe consider the problems of predicting (i) the most dominant person in a group of people, and (ii) the more dominant of a pair of people, from videos depicting group interactions. We introduce a novel family of variables called Dominance Rank. We combine features not previously used for dominance prediction (e.g., facial action units, emotions), with a novel ensemble-based approach to solve these two problems. We test our models against four competing algorithms in the literature on two datasets and show that our results improve past performance. We show 2.4% to 16.7% improvement in AUC compared to baselines on one dataset, and a gain of 0.6% to 8.8% in accuracy on the other. Ablation testing shows that Dominance Rank features play a key role. Chongyang Bai, Maksim Bolonkin, Srijan Kumar, Jure Leskovec, Judee K. Burgoon, Norah E. Dunbar, V. S. Subrahmanian |
IJCAI | 1 |
| 2019 | Predicting the Visual Focus of Attention in Multi-Person Discussion VideosabstractVisual focus of attention in multi-person discussions is a crucial nonverbal indicator in tasks such as inter-personal relation inference, speech transcription, and deception detection. However, predicting the focus of attention remains a challenge because the focus changes rapidly, the discussions are highly dynamic, and the people's behaviors are inter-dependent. Here we propose ICAF (Iterative Collective Attention Focus), a collective classification model to jointly learn the visual focus of attention of all people. Every person is modeled using a separate classifier. ICAF models the people collectively---the predictions of all other people's classifiers are used as inputs to each person's classifier. This explicitly incorporates inter-dependencies between all people's behaviors. We evaluate ICAF on a novel dataset of 5 videos (35 people, 109 minutes, 7604 labels in all) of the popular Resistance game and a widely-studied meeting dataset with supervised prediction. See our demo at https://cs.dartmouth.edu/dsail/demos/icaf. ICAF outperforms the strongest baseline by 1%--5% accuracy in predicting the people's visual focus of attention. Further, we propose a lightly supervised technique to train models in the absence of training labels. We show that light-supervised ICAF performs similar to the supervised ICAF, thus showing its effectiveness and generality to previously unseen videos. Chongyang Bai, Srijan Kumar, Jure Leskovec, Miriam J. Metzger, Jay F. Nunamaker Jr., V. S. Subrahmanian |
IJCAI | 1 |