EDBT 2026 Demo / reviewers in the wild / expert
Bo Wu 0018
dblp:47/6534-18
· DBLP profile ↗
36ranked-venue papers
8as first author
18since 2021 · last 2025
0000-0001-6658-6452ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video SituationabstractVision-language temporal alignment is a crucial capability for human dynamic recognition and cognition in real-world scenarios. While existing research focuses on capturing vision-language relevance, it faces limitations due to biased temporal distributions, imprecise annotations, and insufficient compositionally. To achieve fair evaluation and comprehensive exploration, our objective is to investigate and evaluate the ability of models to achieve alignment from a temporal perspective, specifically focusing on their capacity to synchronize visual scenarios with linguistic context in a temporally coherent manner. As a preliminary step, we present the statistical analysis of existing benchmarks and reveal the existing challenges from a decomposed perspective. To this end, we introduce SVLTA, the Synthetic Vision-Language Temporal Alignment derived via a well-designed and feasible control generation method within a simulation environment. The approach considers commonsense knowledge, manipulable action, and constrained filtering, which generates reasonable, diverse, and balanced data distributions for diagnostic evaluations. Our experiments reveal diagnostic insights through the evaluations in temporal question answering, distributional shift sensitiveness, and temporal alignment adaptation. Bo Wu 0018, Yan Lu 0001, Zhendong Mao 0001 |
CVPR | 2 |
| 2025 | SMPV: Social Media Prediction for Videos
Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng |
ACM Multimedia | 1 |
| 2024 | SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World KnowledgeabstractLearning commonsense reasoning from visual contexts and scenes in real-world is a crucial step toward advanced artificial intelligence. However, existing video reasoning benchmarks are still inadequate since they were mainly designed for factual or situated reasoning and rarely involve broader knowledge in the real world. Our work aims to delve deeper into reasoning evaluations, specifically within dynamic, open-world, and structured context knowledge. We propose a new benchmark (SOK-Bench), consisting of 44K questions and 10K situations with instance-level annotations depicted in the videos. The reasoning process is required to understand and apply situated knowledge and general knowledge for problem-solving. To create such a dataset, we propose an automatic and scalable gener-ation method to generate question-answer pairs, knowledge graphs, and rationales by instructing the combinations of LLMs and MLLMs. Concretely, we first extract observable situated entities, relations, and processes from videos for situated knowledge and then extend to open-world knowledge beyond the visible content. The task generation is facilitated through multiple dialogues as iterations and subsequently corrected and refined by our designed self-promptings and demonstrations. With a corpus of both explicit situated facts and implicit commonsense, we generate associated question-answer pairs and reasoning processes, finally followed by manual reviews for quality assurance. We evaluated recent mainstream large vision-language models on the benchmark and found several in-sightful conclusions. For more information, please refer to our benchmark at www.bobbywu.com/SOKBench. Andong Wang, Bo Wu 0018, Sunli Chen, Zhenfang Chen, Haotian Guan, Wei-Ning Lee, Li Erran Li, Chuang Gan 0001 |
CVPR | 2 |
| 2024 | SMP Challenge Summary: Social Media Prediction Challenge
Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng |
ACM Multimedia | 1 |
| 2024 | Real-Time Decoding of Snapshot Compressive Imaging Using Tensor FISTA-NetabstractSnapshot compressive imaging (SCI) cameras compress high-speed videos or hyperspectral images into measurement frames. However, decoding the data frames from measurement frames is compute-intensive. Existing state-of-the-art decoding algorithms suffer from low decoding quality or heavy running time or both, which are not practical for real-time applications. In this article, we exploit the powerful learning ability of deep neural networks (DNN) and propose a novel tensor fast iterative shrinkage-thresholding algorithm net (Tensor FISTA-Net) as a real-time decoder for SCI cameras. Since SCI cameras have an accurate physical model, we can trade training time for the decoding time by generating abundant synthetic data and training a decoder on the cloud. Tensor FISTA-Net not only learns a sparse representation of the frames through convolution layers but also reduces the decoding time and memory consumption significantly through tensor operations, which makes Tensor FISTA-Net an appropriate approach for a real-time decoder. Our proposed Tensor FISTA-Net obtains an average PSNR improvement of 0.79-2.84 dB (video images) and 2.61-4.43 dB (hyperspectral images) over the state-of-the-art algorithms, along with more clear and detailed visual results on real SCI datasets, Hammer and Wheel, respectively. Our Tensor FISTA-Net reaches 45 frames per second in video datasets and 70 frames per second in hyperspectral datasets, meeting the real-time requirement. Besides, the trained model occupies only a 12 -MB memory footprint, making it applicable to real-time Internet of Things (IoT) applications. Xiao-Yang Liu, Qifan Huang, Xiaochen Han, Bo Wu 0018, Linghe Kong, Anwar Elwalid, Xiaodong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Personalized Dialogue Generation with Persona-Adaptive AttentionabstractPersona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, this requires a delicate weight balance between context and persona. To achieve that, in this paper, we propose an effective framework with Persona-Adaptive Attention (PAA), which adaptively integrates the weights from the persona and context information via our designed attention. In addition, a dynamic masking mechanism is applied to the PAA to not only drop redundant information in context and persona but also serve as a regularization mechanism to avoid overfitting. Experimental results demonstrate the superiority of the proposed PAA framework compared to the strong baselines in both automatic and human evaluation. Moreover, the proposed PAA approach can perform equivalently well in a low-resource regime compared to models trained in a full-data setting, which achieve a similar result with only 20% to 30% of data compared to the larger models trained in the full-data setting. To fully exploit the effectiveness of our design, we designed several variants for handling the weighted information in different ways, showing the necessity and sufficiency of our weighting and masking designs. Qiushi Huang, Yu Zhang 0006, Tom Ko, Xubo Liu 0001, Bo Wu 0018, Wenwu Wang 0001, Lilian Tang |
AAAI | 5 |
| 2023 | Learning Situation Hyper-Graphs for Video Question AnsweringabstractAnswering questions about complex situations in videos requires not only capturing the presence of actors, objects, and their relations but also the evolution of these relationships over time. A situation hyper-graph is a representation that describes situations as scene sub-graphs for video frames and hyper-edges for connected sub-graphs and has been proposed to capture all such information in a compact structured form. In this work, we propose an architecture for Video Question Answering (VQA) that enables answering questions related to video content by predicting situation hyper-graphs, coined Situation Hyper-Graph based Video Question Answering (SHG- VQA). To this end, we train a situation hyper-graph decoder to implicitly identify graph representations with actions and object/human-object relationships from the input video clip. and to use cross-attention between the predicted situation hyper-graphs and the question embedding to predict the correct answer. The proposed method is trained in an end-to-end manner and optimized by a VQA loss with the cross-entropy function and a Hungarian matching loss for the situation graph prediction. The effectiveness of the proposed architecture is extensively evaluated on two challenging benchmarks: AGQA and STAR. Our results show that learning the underlying situation hyper-graphs helps the system to significantly improve its performance for novel challenges of video question-answering tasks11Code will be available at https://github.com/aurooj/SHG-VQA. Aisha Urooj Khan, Hilde Kuehne, Bo Wu 0018, Kim Chheu, Walid Bousselham, Chuang Gan 0001, Niels da Vitoria Lobo, Mubarak Shah |
CVPR | 3 |
| 2023 | Reproducibility Companion Paper: MeTILDA - Platform for Melodic Transcription in Language Documentation and ApplicationabstractThis companion paper supports the replication of the development and evaluation of “MeTILDA - Platform for Melodic Transcription in Language Documentation and Application” that we presented in the ICMR 2021. MeTILDA aims to help document and analyze pitch patterns of endangered languages including Blackfoot, whose prosodic system is characterized by pitch movements. It develops a new form of audio analysis (termed MeT scale which is a perceptual scale) and automates the process of creating visual aids (Pitch Art) to provide more effective visuals of perceived changes in pitch movement. In this paper, we explain the file structure of the source code and publish the details of our data as well as system operations. Moreover, we provide a link to the demo video for facilitating the use of our platform. Mitchell Lee, Sanjay Penmetsa, Min Chen 0009, Mizuki Miyashita, Naatosi Fish, Bo Wu 0018, Omar Shahbaz Khan |
ICMR | 7 |
| 2023 | SMP Challenge: An Overview and Analysis of Social Media Prediction ChallengeabstractSocial Media Popularity Prediction (SMPP) is a crucial task that involves automatically predicting future popularity values of online posts, leveraging vast amounts of multimodal data available on social media platforms. Studying and investigating social media popularity becomes central to various online applications and requires novel methods of comprehensive analysis, multimodal comprehension, and accurate prediction. Bo Wu 0018, Peiye Liu, Wen-Huang Cheng, Bei Liu 0001, Zhaoyang Zeng, Jia Wang 0020, Qiushi Huang, Jiebo Luo 0001 |
ACM Multimedia | 1 |
| 2023 | Scene Graph Refinement Network for Visual Question AnsweringabstractVisual Question Answering aims to answer the free-form natural language question based on the visual clues in a given image. It is a difficult problem as it requires understanding the fine-grained structured information of both language and image for compositional reasoning. To establish the compositional reasoning, recent works attempt to introduce the scene graph in VQA. However, as the generated scene graphs are usually quite noisy, it greatly limits the performance of question answering. Therefore, this paper proposes to refine the scene graphs for improving the effectiveness. Specifically, we present a novelSceneGraphRefinement network (SGR), which introduces a transformer-based refinement network to enhance the object and relation features for better classification. Moreover, as the question provides valuable clues for distinguishing whether the$\left\langle \mathit{subject, predicate, object} \right\rangle$triplets are helpful or not, the SGR network exploits the semantic information presented in the questions to select the most relevant relations for question answering. Extensive experiments are conducted on the GQA benchmark demonstrate the effectiveness of our method. Tianwen Qian, Jingjing Chen 0001, Shaoxiang Chen 0001, Bo Wu 0018, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | DeepFD: Automated Fault Diagnosis and Localization for Deep Learning ProgramsabstractAs Deep Learning (DL) systems are widely deployed for mission-critical applications, debugging such systems becomes essential. Most existing works identify and repair suspicious neurons on the trained Deep Neural Network (DNN), which, unfortunately, might be a detour. Specifically, several existing studies have reported that many unsatisfactory behaviors are actually originated from the faults residing in DL programs. Besides, locating faulty neurons is not actionable for developers, while locating the faulty statements in DL programs can provide developers with more useful information for debugging. Though a few recent studies were proposed to pinpoint the faulty statements in DL programs or the training settings (e.g. too large learning rate), they were mainly designed based on predefined rules, leading to many false alarms or false negatives, especially when the faults are beyond their capabilities. Jialun Cao, Meiziniu Li, Xiao Chen 0026, Ming Wen 0001, Yongqiang Tian 0001, Bo Wu 0018, Shing-Chi Cheung |
ICSE | 6 |
| 2022 | Attention-guided transformation-invariant attack for black-box adversarial examplesabstractWith the development of media convergence, information acquisition is no longer limited to traditional media, such as newspapers and televisions, but more from digital media on the Internet, where media contents should be under supervision by platforms. At present, the media content analysis technology of Internet platforms relies on deep neural networks (DNNs). However, DNNs show vulnerability to adversarial examples, which results in security risks. Therefore, it is necessary to adequately study the internal mechanism of adversarial examples to build more effective supervision models. When coming to practical applications, supervision models are mostly faced with black-box attacks, where cross-model transferability of adversarial examples has attracted increasing attention. In this paper, to improve the transferability of adversarial examples, we propose an attention-guided transformation-invariant adversarial attack method, which incorporates an attention mechanism to disrupt the most distinctive features and simultaneously ensures adversarial attack invariance under different transformations. Specifically, we dynamically weight the latent features according to an attention mechanism and disrupt them accordingly. Meanwhile, considering the lack of semantics in low-level features, high-level semantics are introduced as spatial guidance to make low-level feature perturbations concentrate on the most discriminative regions. Moreover, since the attention heatmaps may vary significantly across different models, a transformation-invariant aggregated attack strategy is proposed to alleviate overfitting to the proxy model attention. Comprehensive experimental results show that the proposed method can significantly improve the transferability of adversarial examples. Lingyun Yu 0002, Hongtao Xie 0001, Bo Wu 0018, Yongdong Zhang 0001 |
Int. J. Intell. Syst. | 6 |
| 2022 | Text-instance graph: Exploring the relational semantics for text-based visual question answering
Bo Wu 0018, Jingkuan Song, Lianli Gao, Pengpeng Zeng, Chuang Gan 0001 |
Pattern Recognit. | 2 |
| 2022 | RL-Recruiter+: Mobility-Predictability-Aware Participant Selection Learning for From-Scratch Mobile CrowdsensingabstractParticipant selection is a fundamental research issue in Mobile Crowdsensing (MCS). Previous approaches commonly assume that adequately long periods of candidate participants’ historical mobility trajectories are available to model their patterns before the selection process, which is not realistic for some new MCS applications or platforms. The sparsity or even absence of mobility traces will incur inaccurate location prediction, thus undermining the deployment of new MCS applications. To this end, this paper investigates a novel problem called “From-Scratch MCS” (FS-MCS for short), in which we study how to intelligently select participants to minimize such a “cold-start” effect. Specifically, we propose a novel framework based on reinforcement learning, named RL-Recruiter+. With the gradual accumulation of mobility trajectories over time, RL-Recruiter+ is able to make a good sequence of participant selection decisions for each sensing slot. Compared to its previous version, RL-Recruiter, Re-Recruiter+ jointly considers both the previous coverage and current mobility predictability when training the participant selection decision model. We evaluate our approach experimentally based on two real-world mobility datasets. The results demonstrate that RL-Recruiter+ outperforms the baseline approaches, including RL-Recruiter under various settings. Yunfan Hu, Jiangtao Wang 0001, Bo Wu 0018, Abdelsalam Helal |
IEEE Trans. Mob. Comput. | 3 |
| 2021 | Bilateral Autotrading Framework for Stock PredictionabstractAs the core of quantitative trading, indicator effectiveness continuously plays a vital role in stock prediction. The majority of studies are currently dedicated to constructing indicators with high Pearson Correlation Coefficient (CORR) with returns. However, the pursuit of high CORR may ignore some indicators that produce high profits. Therefore, we propose a new Bilateral Correlation Coefficient (BCORR), which can detect some profitable indicators that are previously discarded due to low CORR. BCORR is a weighted correlation coefficient, and the weight is a variable related to the return so that the top and bottom ranked returns have a more significant impact on it. To generate an indicator that has high BCORR with the return, we propose a framework called the Bilateral Autotrading Framework (BAF) based on Bilateral Loss to forecast the cross-sectional rank of stock return, and the prediction is adopted as a bilateral indicator to select stocks to invest. Meanwhile, the positions of the selected stocks are optimized by the Sharpe-oriented optimization to reduce the risk and improve the return. Experiments on real-world stock market datasets show that the BAF can significantly improve performance to deep stock prediction methods, such as Transformer, LSTM, and NBEATS. Qifei Zhou, Hucheng Liu, Weiping Li 0002, Tong Mo, Bo Wu 0018 |
IJCNN | 5 |
| 2021 | Token-Level Supervised Contrastive Learning for Punctuation RestorationabstractPunctuation is critical in understanding natural language text. Currently, most automatic speech recognition (ASR) systems do not generate punctuation, which affects the performance of downstream tasks, such as intent detection and slot filling. This gives rise to the need for punctuation restoration. Recent work in punctuation restoration heavily utilizes pre-trained language models without considering data imbalance when predicting punctuation classes. In this work, we address this problem by proposing a token-level supervised contrastive learning method that aims at maximizing the distance of representation of different punctuation marks in the embedding space. The result shows that training with token-level supervised contrastive learning obtains up to 3.2% absolute F1 improvement on the test set. Qiushi Huang, Tom Ko, Lilian Tang, Xubo Liu 0001, Bo Wu 0018 |
Interspeech | 5 |
| 2021 | Counterfactual Debiasing Inference for Compositional Action RecognitionabstractCompositional action recognition is a novel challenge in the computer vision community and focuses on revealing the different combinations of verbs and nouns instead of treating subject-object interactions in videos as individual instances only. Existing methods tackle this challenging task by simply ignoring appearance information or fusing object appearances with dynamic instance tracklets. However, those strategies usually do not perform well for unseen action instances. For that, in this work we propose a novel learning framework called Counterfactual Debiasing Network (CDN) to improve the model generalization ability by removing the interference introduced by visual appearances of objects/subjects. It explicitly learns the appearance information in action representations and later removes the effect of such information in a causal inference manner. Specifically, we use tracklets and video content to model the factual inference by considering both appearance information and structure information. In contrast, only video content with appearance information is leveraged in the counterfactual inference. With the two inferences, we conduct a causal graph which captures and removes the bias introduced by the appearance information by subtracting the result of the counterfactual inference from that of the factual inference. By doing that, our proposed CDN method can better recognize unseen action instances by debiasing the effect of appearances. Extensive experiments on the Something-Else dataset clearly show the effectiveness of our proposed CDN over existing state-of-the-art methods. Pengzhan Sun 0001, Bo Wu 0018, Xunsong Li, Wen Li 0001, Lixin Duan, Chuang Gan 0001 |
ACM Multimedia | 2 |
| 2021 | STST: Spatial-Temporal Specialized Transformer for Skeleton-based Action RecognitionabstractSkeleton-based action recognition has been widely investigated considering their strong adaptability to dynamic circumstances and complicated backgrounds. To recognize different actions from skeleton sequences, it is essential and crucial to model the posture of the human represented by the skeleton and its changes in the temporal dimension. However, most of the existing works treat skeleton sequences in the temporal and spatial dimension in the same way, ignoring the difference between the temporal and spatial dimension in skeleton data which is not an optimal way to model skeleton sequences. The posture represented by the skeleton in each frame is proposed to be modeled individually. Meanwhile, capturing the movement of the entire skeleton in the temporal dimension is needed. So, we designed Spatial Transformer Block and Directional Temporal Transformer Block for modeling skeleton sequences in spatial and temporal dimensions respectively. Due to occlusion/sensor/raw video, etc., there are noises on both temporal and spatial dimensions in the extracted skeleton data reducing the recognition capabilities of models. To adapt to this imperfect information condition, we propose a multi-task self-supervised learning method by providing confusing samples in different situations to improve the robustness of our model. Combining the above design, we propose our Spatial-Temporal Specialized Transformer~(STST) and conduct experiments with our model on the SHREC, NTU-RGB+D, and Kinetics-Skeleton. Extensive experimental results demonstrate the improved performances and analysis of the proposed method. Bo Wu 0018, Wen Li 0001, Lixin Duan, Chuang Gan 0001 |
ACM Multimedia | 2 |
| 2020 | General Partial Label Learning via Dual Bipartite Graph AutoencoderabstractWe formulate a practical yet challenging problem: General Partial Label Learning (GPLL). Compared to the traditional Partial Label Learning (PLL) problem, GPLL relaxes the supervision assumption from instance-level — a label set partially labels an instance — to group-level: 1) a label set partially labels a group of instances, where the within-group instance-label link annotations are missing, and 2) cross-group links are allowed — instances in a group may be partially linked to the label set from another group. Such ambiguous group-level supervision is more practical in real-world scenarios as additional annotation on the instance-level is no longer required, e.g., face-naming in videos where the group consists of faces in a frame, labeled by a name set in the corresponding caption. In this paper, we propose a novel graph convolutional network (GCN) called Dual Bipartite Graph Autoencoder (DB-GAE) to tackle the label ambiguity challenge of GPLL. First, we exploit the cross-group correlations to represent the instance groups as dual bipartite graphs: within-group and cross-group, which reciprocally complements each other to resolve the linking ambiguities. Second, we design a GCN autoencoder to encode and decode them, where the decodings are considered as the refined results. It is worth noting that DB-GAE is self-supervised and transductive, as it only uses the group-level supervision without a separate offline training stage. Extensive experiments on two real-world datasets demonstrate that DB-GAE significantly outperforms the best baseline over absolute 0.159 F1-score and 24.8% accuracy. We further offer analysis on various levels of label ambiguities. Brian Chen 0001, Bo Wu 0018, Alireza Zareian, Hanwang Zhang, Shih-Fu Chang |
AAAI | 2 |
| 2020 | Tensor FISTA-Net for Real-Time Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) cameras capture high-speed videos by compressing multiple video frames into a measurement frame. However, reconstructing video frames from the compressed measurement frame is challenging. The existing state-of-the-art reconstruction algorithms suffer from low reconstruction quality or heavy time consumption, making them not suitable for real-time applications. In this paper, exploiting the powerful learning ability of deep neural networks (DNN), we propose a novel Tensor Fast Iterative Shrinkage-Thresholding Algorithm Net (Tensor FISTA-Net) as a decoder for SCI video cameras. Tensor FISTA-Net not only learns the sparsest representation of the video frames through convolution layers, but also reduces the reconstruction time significantly through tensor calculations. Experimental results on synthetic datasets show that the proposed Tensor FISTA-Net achieves average PSNR improvement of 1.63∼3.89dB over the state-of-the-art algorithms. Moreover, Tensor FISTA-Net takes less than 2 seconds running time and 12MB memory footprint, making it practical for real-time IoT applications. Xiaochen Han, Bo Wu 0018, Xiao-Yang Liu, Linghe Kong |
AAAI | 2 |
| 2020 | Enhancing Neural Models with Vulnerability via Adversarial AttackabstractNatural Language Sentence Matching (NLSM) serves as the core of many natural language processing tasks.1) Most previous work develops a single specific neural model for NLSM tasks.2) There is no previous work considering adversarial attack to improve the performance of NLSM tasks.3) Adversarial attack is usually used to generate adversarial samples that can fool neural models.In this paper, we first find a phenomenon that different categories of samples have different vulnerabilities.Vulnerability is the difficulty degree in changing the label of a sample.Considering the phenomenon, we propose a general two-stage training framework to enhance neural models with Vulnerability via Adversarial Attack (VAA).We design criteria to measure the vulnerability which is obtained by adversarial attack.VAA framework can be adapted to various neural models by incorporating the vulnerability.In addition, we prove a theorem and four corollaries to explain the factors influencing vulnerability effectiveness.Experimental results show that VAA significantly improves the performance of neural models on NLSM datasets.The results are also consistent with the theorem and corollaries.The code is released on https://github.com/rzhangpku/VAA. Qifei Zhou, Bo An 0002, Weiping Li 0002, Tong Mo, Bo Wu 0018 |
COLING | 6 |
| 2020 | MemNAS: Memory-Efficient Neural Architecture Search With Grow-Trim LearningabstractRecent studies on automatic neural architecture search techniques have demonstrated significant performance, competitive to or even better than hand-crafted neural architectures. However, most of the existing search approaches tend to use residual structures and a concatenation connection between shallow and deep features. A resulted neural network model, therefore, is non-trivial for resource-constraint devices to execute since such a model requires large memory to store network parameters and intermediate feature maps along with excessive computing complexity. To address this challenge, we propose MemNAS, a novel growing and trimming based neural architecture search framework that optimizes not only performance but also memory requirement of an inference network. Specifically, in the search process, we consider running memory use, including network parameters and the essential intermediate feature maps memory requirement, as an optimization objective along with performance. Besides, to improve the accuracy of the search, we extract the correlation information among multiple candidate architectures to rank them and then choose the candidates with desired performance and memory efficiency. On the ImageNet classification task, our MemNAS achieves 75.4% accuracy, 0.7% higher than MobileNetV2 with 42.1% less memory requirement. Additional experiments confirm that the proposed MemNAS can perform well across the different targets of the trade-off between accuracy and memory consumption. Peiye Liu, Bo Wu 0018, Huadong Ma, Mingoo Seok |
CVPR | 2 |
| 2020 | Detection by Attack: Detecting Adversarial Samples by Undercover Attack
Qifei Zhou, Bo Wu 0018, Weiping Li 0002, Tong Mo |
ESORICS (2) | 3 |
| 2020 | Learning the Compositional Visual Coherence for Complementary RecommendationsabstractComplementary recommendations, which aim at providing users product suggestions that are supplementary and compatible with their obtained items, have become a hot topic in both academia and industry in recent years. Existing work mainly focused on modeling the co-purchased relations between two items, but the compositional associations of item collections are largely unexplored. Actually, when a user chooses the complementary items for the purchased products, it is intuitive that she will consider the visual semantic coherence (such as color collocations, texture compatibilities) in addition to global impressions. Towards this end, in this paper, we propose a novel Content Attentive Neural Network (CANN) to model the comprehensive compositional coherence on both global contents and semantic contents. Specifically, we first propose a Global Coherence Learning (GCL) module based on multi-heads attention to model the global compositional coherence. Then, we generate the semantic-focal representations from different semantic regions and design a Focal Coherence Learning (FCL) module to learn the focal compositional coherence from different semantic-focal representations. Finally, we optimize the CANN in a novel compositional optimization strategy. Extensive experiments on the large-scale real-world data clearly demonstrate the effectiveness of CANN compared with several state-of-the-art methods. Zhi Li 0057, Bo Wu 0018, Qi Liu 0003, Likang Wu, Hongke Zhao, Tao Mei 0001 |
IJCAI | 2 |
| 2020 | Reproducibility Companion Paper: Outfit Compatibility Prediction and Diagnosis with Multi-Layered Comparison NetworkabstractThis companion paper supports the experimental replication of paper "Outfit Compatibility Prediction and Diagnosis with Multi-Layered Comparison Network", which is presented at ACM Multimedia 2019. We provide the software package for replicating the implementation of Multi-Layered Comparison Network (MCN), as well as the Polyvore-T dataset and baseline methods compared in the original paper. This paper contains the guides to reproduce the experiment results including outfit compatibility prediction, outfit diagnosis and automatic outfit revision. Xin Wang 0131, Bo Wu 0018, Yueqi Zhong, Wei Hu 0003, Jan Zahálka |
ACM Multimedia | 2 |
| 2020 | Video Synthesis via Transform-Based Tensor Neural NetworkabstractVideo frame synthesis is an important task in computer vision and has drawn great interests in wide applications. However, existing neural network methods do not explicitly impose tensor low-rankness of videos to capture the spatiotemporal correlations in a high-dimensional space, while existing iterative algorithms require hand-crafted parameters and take relatively long running time. In this paper, we propose a novel multi-phase deep neural network Transform-Based Tensor-Net that exploits the low-rank structure of video data in a learned transform domain, which unfolds an Iterative Shrinkage-Thresholding Algorithm (ISTA) for tensor signal recovery. Our design is based on two observations: (i) both linear and nonlinear transforms can be implemented by a neural network layer, and (ii) the soft-thresholding operator corresponds to an activation function. Further, such an unfolding design is able to achieve nearly real-time at the cost of training time and enjoys an interpretable nature as a byproduct. Experimental results on the KTH and UCF-101 datasets show that compared with the state-of-the-art methods, i.e., DVF and Super SloMo, the proposed scheme improves Peak Signal-to-Noise Ratio (PSNR) of video interpolation and prediction by 4.13 dB and 4.26 dB, respectively. Xiao-Yang Liu, Bo Wu 0018, Anwar Elwalid |
ACM Multimedia | 3 |
| 2020 | Participants Selection for From-Scratch Mobile Crowdsensing via Reinforcement LearningabstractParticipant selection is a major research challenge in Mobile Crowdsensing (MCS). Previous approaches commonly assume that adequately long and fixed periods of candidate participants’ historical mobility trajectories are available before the selection process. This enables the frameworks to accurately model mobility which enables the optimization of selection. However, this assumption may not be realistic for newly-released MCS applications or platforms because the candidates have just boarded without previous mobility profiles. The sparsity or even absence of mobility traces will incur inaccurate location prediction of the individual participant, thus imposing negative effects on the participant selection process and hindering the practical deployment of new MCS applications. To this end, this paper investigates a novel problem called "From-Scratch MCS" (FS-MCS for short), in which we study how to intelligently select participants to minimize such "cold-start" effect. Specifically, we propose a novel framework based on reinforcement learning, which we name RL-Recruiter. With the gradual accumulation of mobility trajectories over time, RL-Recruiter can make a good sequence of participant selection decisions for each sensing slot by incrementally extracting and utilizing the collective mobility patterns of all candidate participants, thus avoiding the prediction of individual participant’s location that is very inaccurate when the training data is sparse. We test our approach experimentally based on two real-world mobility datasets. Our experiment results demonstrate that RL-Recruiter outperforms the baseline approaches under various settings. Yunfan Hu, Jiangtao Wang 0001, Bo Wu 0018, Abdelsalam Helal |
PerCom | 3 |
| 2020 | What Do Questions Exactly Ask? MFAE: Duplicate Question Identification with Multi-Fusion Asking EmphasisabstractDuplicate Question Identification (DQI) improves the processing efficiency and accuracy of large-scale community question answering and automatic QA system. The purpose of DQI task is to identify whether the paired questions are semantically equivalent. However, how to distinguish the synonyms or homonyms in paired questions is still challenging. Most previous works focus on the word-level or phrase-level semantic differences. We firstly propose to explore the asking emphasis of a question as a key factor in DQI. Asking emphasis bridges semantic equivalence between two questions. In this paper, we propose an attention model with multi-fusion asking emphasis (MFAE) for DQI. At first, BERT is used to obtain the dynamic pre-trained word embeddings. Then we get inter- and intra-asking emphasis by summing inter-attention and self-attention, respectively; the idea is that, the more a word interacts with others, the more important the word is. Finally, we use eight-way combinations to generate multi-fusion asking emphasis and multi-fusion word representation. Experimental results demonstrate that our model achieves state-of-the-art performance on both Quora Question Pairs and CQADupStack data. In addition, our model can also improve the results for natural language inference task on SNLI and MultiNLI datasets. The code is available at https://github.com/rzhangpku/MFAE. Qifei Zhou, Bo Wu 0018, Weiping Li 0002, Tong Mo |
SDM | 3 |
| 2020 | Unlocking Author Power: On the Exploitation of Auxiliary Author-Retweeter Relations for Predicting Key RetweetersabstractRetweeting is a powerful driving force in information propagation on microblogging sites. However, identifying the most effective retweeters of a message (called the ”key retweeter prediction” problem) has become a significant research topic. Conventional approaches have addressed this topic from two main aspects: by analyzing either the personal attributes of microblogging users or the structures of user graph networks. However, according to sociological findings, author-retweeter dependencies also play a crucial role in influencing message propagation. In this paper, we propose a novel model to solve the key retweeter prediction problem by incorporating the auxiliary relations between a tweet author and potential retweeters. Without loss of generality, we formulate the relations from four relational factors: status relation, temporal relation, locational relation, and interactive relation. In addition, we propose a novel method, called “Relation-based Learning to Rank (RL2R),” to determine the key retweeters for a given tweet by ranking the potential retweeters in terms of their spreadability. The experimental results show that our method outperforms the state-of-the-art algorithms at top-k retweeter prediction, achieving a significant relative average improvement of 19.7-29.4 percent. These findings provide new insights for understanding user behaviors on social media for key retweeter prediction purposes. Bo Wu 0018, Wen-Huang Cheng, Yongdong Zhang 0001, Juan Cao 0001, Jintao Li 0001, Tao Mei 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Dressing for Attention: Outfit Based Fashion Popularity PredictionabstractAnalysis of fashion trends is crucial. However, existing predictive algorithms of fashion popularity are restricted to be feasible on the coarse style level but not a finer item level. That is, they are only predictive in the future popularity of a given type of fashion styles (e.g., Rocker), but cannot be precisely down to a particular outfit look chosen by individuals. This paper thus proposes the first solution directly aimed at predicting the fine-grained fashion popularity of an outfit look by taking social media as the learning source. Particularly, a deep temporal sequence learning framework is developed and the proposed framework is evaluated on a real dataset of 380,000 street fashion images collected from the fashion website lookbook.nu. The experimental results show that our proposed framework outperforms the state-of-the-art approaches, with a relative increase of 11.51% to 27.62% (MSE metric) and 7.02% to 32.61% (CSE metric) in the prediction accuracy. Ling Lo, Chia-Lin Liu, Rong-An Lin, Bo Wu 0018, Hong-Han Shuai, Wen-Huang Cheng |
ICIP | 4 |
| 2019 | SMP Challenge: An Overview of Social Media Prediction Challenge 2019abstract"SMP Challenge" aims to discover novel prediction tasks for numerous data on social multimedia and seek excellent research teams. Making predictions via social multimedia data (e.g. photos, videos or news) is not only helps us to make better strategic decisions for the future, but also explores advanced predictive learning and analytic methods on various problems and scenarios, such as multimedia recommendation, advertising system, fashion analysis etc. Bo Wu 0018, Wen-Huang Cheng, Peiye Liu, Bei Liu 0001, Zhaoyang Zeng, Jiebo Luo 0001 |
ACM Multimedia | 1 |
| 2019 | Outfit Compatibility Prediction and Diagnosis with Multi-Layered Comparison NetworkabstractExisting works about fashion outfit compatibility focus on predicting the overall compatibility of a set of fashion items with their information from different modalities. However, there are few works explore how to explain the prediction, which limits the persuasiveness and effectiveness of the model. In this work, we propose an approach to not only predict but also diagnose the outfit compatibility. We introduce an end-to-end framework for this goal, which features for: (1) The overall compatibility is learned from all type-specified pairwise similarities between items, and the backpropagation gradients are used to diagnose the incompatible factors. (2) We leverage the hierarchy of CNN and compare the features at different layers to take into account the compatibilities of different aspects from the low level (such as color, texture) to the high level (such as style). To support the proposed method, we build a new type-specified outfit dataset named Polyvore-T based on Polyvore dataset. We compare our method with the prior state-of-the-art in two tasks: outfit compatibility prediction and fill-in-the-blank. Experiments show that our approach has advantages in both prediction performance and diagnosis ability. Xin Wang 0131, Bo Wu 0018, Yueqi Zhong |
ACM Multimedia | 2 |
| 2017 | Sequential Prediction of Social Media Popularity with Deep Temporal Context NetworksabstractPrediction of popularity has profound impact for social media, since it offers opportunities to reveal individual preference and public attention from evolutionary social systems. Previous research, although achieves promising results, neglects one distinctive characteristic of social data, i.e., sequentiality. For example, the popularity of online content is generated over time with sequential post streams of social media. To investigate the sequential prediction of popularity, we propose a novel prediction framework called Deep Temporal Context Networks (DTCN) by incorporating both temporal context and temporal attention into account. Our DTCN contains three main components, from embedding, learning to predicting. With a joint embedding network, we obtain a unified deep representation of multi-modal user-post data in a common embedding space. Then, based on the embedded data sequence over time, temporal context learning attempts to recurrently learn two adaptive temporal contexts for sequential popularity. Finally, a novel temporal attention is designed to predict new popularity (the popularity of a new user-post pair) with temporal coherence across multiple time-scales. Experiments on our released image dataset with about 600K Flickr photos demonstrate that DTCN outperforms state-of-the-art deep prediction algorithms, with an average of 21.51% relative performance improvement in the popularity prediction (Spearman Ranking Correlation). Bo Wu 0018, Wen-Huang Cheng, Yongdong Zhang 0001, Qiushi Huang, Jintao Li 0001, Tao Mei 0001 |
IJCAI | 1 |
| 2017 | Fashion World Map: Understanding Cities Through Streetwear FashionabstractFashion is an integral part of life. Streets as a social center for people's interaction become the most important public stage to showcase the fashion culture of a metropolitan area. In this paper, therefore, we propose a novel framework based on deep neural networks (DNN) for depicting the street fashion of a city by automatically discovering fashion items (e.g., jackets) in a particular look that are most iconic for the city, directly from a large collection of geo-tagged street fashion photos. To obtain a reasonable collection of iconic items, our task is formulated as the prize-collecting Steiner tree (PCST) problem, whereby a visually intuitive summary of the world's iconic street fashion can be created. To the best of our knowledge, this is the first work devoted to investigate the world's fashion landscape in modern times through the visual analytics of big social data. It shows how the visual impression of local fashion cultures across the world can be depicted, modeled, analyzed, compared, and exploited. In the experiments, our approach achieves the best performance (43.19%) on our large collected GSFashion dataset (170K photos), with an average of two times higher than all the other algorithms (FII: 20.13%, AP: 18.76%, DC: 17.90%), in terms of the users' agreement ratio on the discovered iconic fashion items of a city. The potential of our proposed framework for advanced sociological understanding is also demonstrated via practical applications. Yu-Ting Chang, Wen-Huang Cheng, Bo Wu 0018, Kai-Lung Hua |
ACM Multimedia | 3 |
| 2016 | Unfolding Temporal Dynamics: Predicting Social Media Popularity Using Multi-scale Temporal DecompositionabstractTime information plays a crucial role on social media popularity. Existing research on popularity prediction, effective though, ignores temporal information which is highly related to user-item associations and thus often results in limited success. An essential way is to consider all these factors (user, item, and time), which capture the dynamic nature of photo popularity. In this paper, we present a novel approach to factorize the popularity into user-item context and time-sensitive context for exploring the mechanism of dynamic popularity. The user-item context provides a holistic view of popularity, while the time-sensitive context captures the temporal dynamics nature of popularity. Accordingly, we develop two kinds of time-sensitive features, including user activeness variability and photo prevalence variability. To predict photo popularity, we propose a novel framework named Multi-scale Temporal Decomposition (MTD), which decomposes the popularity matrix in latent spaces based on contextual associations. Specifically, the proposed MTD models time-sensitive context on different time scales, which is beneficial to automatically learn temporal patterns. Based on the experiments conducted on a real-world dataset with 1.29M photos from Flickr, our proposed MTD can achieve the prediction accuracy of 79.8% and outperform the best three state-of-the-art methods with a relative improvement of 9.6% on average. Bo Wu 0018, Tao Mei 0001, Wen-Huang Cheng, Yongdong Zhang 0001 |
AAAI | 1 |
| 2016 | Time Matters: Multi-scale Temporalization of Social Media PopularityabstractThe evolution of social media popularity exhibits rich temporality, i.e., popularities change over time at various levels of temporal granularity. This is influenced by temporal variations of public attentions or user activities. For example, popularity patterns of street snap on Flickr are observed to depict distinctive fashion styles at specific time scales, such as season-based periodic fluctuations for Trench Coat or one-off peak in days for Evening Dress. However, this fact is often overlooked by existing research of popularity modeling. We present the first study to incorporate multiple time-scale dynamics into predicting online popularity. We propose a novel computational framework in the paper, named Multi-scale Temporalization, for estimating popularity based on multi-scale decomposition and structural reconstruction in a tensor space of user, post, and time by joint low-rank constraints. By considering the noise caused by context inconsistency, we design a data rearrangement step based on context aggregation as preprocessing to enhance contextual relevance of neighboring data in the tensor space. As a result, our approach can leverage multiple levels of temporal characteristics and reduce the noise of data decomposition to improve modeling effectiveness. We evaluate our approach on two large-scale Flickr image datasets with over 1.8 million photos in total, for the task of popularity prediction. The results show that our approach significantly outperforms state-of-the-art popularity prediction techniques, with a relative improvement of 10.9%-47.5% in terms of prediction accuracy. Bo Wu 0018, Wen-Huang Cheng, Yongdong Zhang 0001, Tao Mei 0001 |
ACM Multimedia | 1 |