Da Cao

dblp:193/7279 · DBLP profile ↗
← Back
43ranked-venue papers
12as first author
20since 2021 · last 2026
0000-0002-2611-2559ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 6 first-author · 4 since 2021Computer networks · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 PelvicDiff: A novel diffusion model for high-fidelity synthetic CT generation from MRI in pelvic cancer radiotherapy
Can Hu, Xiayu Hang, Xiuhan Li, Da Cao
Expert Syst. Appl.4
2026 Automated data synthesis and retrieval-augmented generation for legal large language models
Wenqi Ren, Lixing Shen, Yinxia Hong, Jiawei Wang 0025, Da Cao
Knowl. Based Syst.7
2025 Neural Causal Graph for Interpretable and Intervenable Classification
abstract
Advancements in neural networks have significantly enhanced the performance of classification models, achieving remarkable accuracy across diverse datasets. However, these models often lack transparency and do not support interactive reasoning with human users, which are essential attributes for applications that require trust and user engagement. To overcome these limitations, we introduce an innovative framework, Neural Causal Graph (NCG), that integrates causal inference with neural networks to enable interpretable and intervenable reasoning. We then propose an intervention training method to model the intervention probability of the prediction, serving as a contextual prompt to facilitate the fine-grained reasoning and human-AI interaction abilities of NCG. Our experiments show that the proposed framework significantly enhances the performance of traditional classification baselines. Furthermore, NCG achieves nearly 95\% top-1 accuracy on the ImageNet dataset by employing a test-time intervention method. This framework not only supports sophisticated post-hoc interpretation but also enables dynamic human-AI interactions, significantly improving the model's transparency and applicability in real-world scenarios.
Jiawei Wang 0025, Shaofei Lu, Da Cao, Yuquan Le, Zhe Quan, Tat-Seng Chua
ICLR3
2025 Spatial-temporal video grounding with cross-modal understanding and enhancement
Shu Luo, Jingyu Pan, Da Cao, Jiawei Wang 0025, Yuquan Le, Meng Liu 0006
Expert Syst. Appl.3
2025 AutoVMR: An autonomous event generation and localization approach for video moment retrieval
Shu Luo, Qiwei Ma, Jiawei Wang 0025, Da Cao, Shaofei Lu
Inf. Sci.4
2025 Weakly-supervised spatial-temporal video grounding via spatial-temporal annotation on a single frame
Shu Luo, Shijie Jiang, Da Cao, Huangxiao Deng, Jiawei Wang 0025, Zheng Qin 0001
Knowl. Based Syst.3
2025 Graph Reasoning With Supervised Contrastive Learning for Legal Judgment Prediction
abstract
Given the fact descriptions of legal cases, the legal judgment prediction (LJP) problem aims to determine three judgment tasks of law articles, charges, and the term of penalty. Most existing studies have considered task dependencies while neglecting the prior dependencies of labels among different tasks. Therefore, how to make better use of the information on the relation dependencies among tasks and labels becomes a crucial issue. To this end, we transform the text classification problem into a node classification framework based on graph reasoning and supervised contrastive learning (SCL) techniques, named GraSCL. Specifically, we first design a graph reasoning network to model the potential dependency structures and facilitate relational learning under various graph topologies. Then, we introduce the SCL method for the LJP task to further leverage the label relation on the graph. To accommodate the node classification settings, we extend the traditional SCL method to novel variants for SCL at the node level, which allows the GraSCL framework to be trained efficiently even with small batches. Furthermore, to recognize the importance of hard negative samples in contrastive learning, we introduce a simple yet effective technique called online hard negative mining (OHNM) to enhance our SCL approach. This technique complements our SCL method and enables us to control the number and complexity of negative samples, leading to further improvements in the model's performance. Finally, extensive experiments are conducted on two well-known benchmarks, demonstrating the effectiveness and rationality of our proposed SCL approach as compared to the state-of-the-art competitors.
Jiawei Wang 0025, Yuquan Le, Da Cao, Shaofei Lu, Zhe Quan, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identification
abstract
The fine-grained attribute descriptions can significantly supplement the valuable semantic information for person image, which is vital to the success of person re-identification (ReID) task. However, current ReID algorithms typically failed to effectively leverage the rich contextual information available, primarily due to their reliance on simplistic and coarse utilization of image attributes. Recent advances in artificial intelligence generated content have made it possible to automatically generate plentiful fine-grained attribute descriptions and make full use of them. Thereby, this paper explores the potential of using the generated multiple person attributes as prompts in ReID tasks with off-the-shelf (large) models for more accurate retrieval results. To this end, we present a new framework called Multi-Prompts ReID (MP-ReID), based on prompt learning and language models, to fully dip fine attributes to assist ReID task. Specifically, MP-ReID first learns to hallucinate diverse, informative, and promptable sentences for describing the query images. This procedure includes (i) explicit prompts of which attributes a person has and furthermore (ii) implicit learnable prompts for adjusting/conditioning the criteria used towards this person identity matching. Explicit prompts are obtained by ensembling generation models, such as ChatGPT and VQA models. Moreover, an alignment module is designed to fuse multi-prompts (i.e., explicit and implicit ones) progressively and mitigate the cross-modal gap. Extensive experiments on the existing attribute-involved ReID datasets, namely, Market1501 and DukeMTMC-reID, demonstrate the effectiveness and rationality of the proposed MP-ReID solution.
Yajing Zhai, Yawen Zeng, Zheng Qin 0001, Xin Jin 0014, Da Cao
AAAI6
2024 Causal-driven Large Language Models with Faithful Reasoning for Knowledge Question Answering
abstract
In Large Language Models (LLMs), text generation that involves knowledge representation is often fraught with the risk of "hallucinations'', where models confidently produce erroneous or fabricated content. These inaccuracies often stem from intrinsic biases in the pre-training stage or from the incorporation of human preference biases during the fine-tuning process. To mitigate these issues, we take inspiration from Goldman's causal theory of knowledge, which asserts that knowledge is not merely about having a true belief but also involves a causal connection between the belief and the truth of the proposition. We instantiate this theory within the context of Knowledge Question Answering (KQA) by constructing a causal graph that delineates the pathways between the candidate knowledge and belief. Through the application of the do-calculus rules from structural causal models, we devise an unbiased estimation framework based on this causal graph, thereby establishing a methodology for knowledge modeling grounded in causal inference. The resulting CORE framework (short for "Causal knOwledge REasoning'') is comprised of four essential components: question answering, causal reasoning, belief scoring, and refinement. Together, they synergistically improve the KQA system by fostering faithful reasoning and introspection. Extensive experiments are conducted on ScienceQA and HotpotQA datasets, which demonstrate the effectiveness and rationality of the CORE framework.
Jiawei Wang 0025, Da Cao, Shaofei Lu, Zhanchang Ma, Junbin Xiao, Tat-Seng Chua
ACM Multimedia2
2024 $\boldsymbol{R}^{2}$: A Novel Recall & Ranking Framework for Legal Judgment Prediction
abstract
The legal judgment prediction (LJP) task is to automatically decide appropriate law articles, charges, and term of penalty for giving the fact description of a law case. It considerably influences many real legal applications and has thus attracted the attention of legal practitioners and AI researchers in recent years. In real scenarios, many confusing charges are encountered, which makes LJP challenging. Intuitively, for a controversial legal case, legal practitioners usually first obtain various possible judgment results as candidates based on the fact description of the case; then these candidates generally need to be carefully considered based on the facts and the rationality of the candidates. Inspired by this observation, this paper presents a novelRecall &Ranking framework, dubbed as$\boldsymbol{R}^{2}$, which attempts to formalize LJP as a two-stage problem. The recall stage is designed to collect high-likelihood judgment results for a given case; these results are regarded as candidates for the ranking stage. The ranking stage introduces a verification technique to learn the relationships between the fact description and the candidates. It treats the partially correct candidates as semi-negative samples, and thus has a certain ability to distinguish confusing candidates. Moreover, we devise a comprehensive judgment strategy to refine the final judgment results by comprehensively considering the rationality of multiple probable candidates. We carry out numerous experiments on two widely used benchmark datasets. The experimental results demonstrate our proposed approach's effectiveness compared to the other competitive baselines.
Yuquan Le, Zhe Quan, Jiawei Wang 0025, Da Cao, Kenli Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Universal Relocalizer for Weakly Supervised Referring Expression Grounding
abstract
This article introduces the Universal Relocalizer, a novel approach designed for weakly supervised referring expression grounding. Our method strives to pinpoint a target proposal that corresponds to a specific query, eliminating the need for region-level annotations during training. To bolster the localization precision and enrich the semantic understanding of the target proposal, we devise three key modules: the category module, the color module, and the spatial relationship module. The category and color modules assign respective category and color labels to region proposals, enabling the computation of category and color scores. Simultaneously, the spatial relationship module integrates spatial cues, yielding a spatial score for each proposal to enhance localization accuracy further. By adeptly amalgamating the category, color, and spatial scores, we derive a refined grounding score for every proposal. Comprehensive evaluations on the RefCOCO, RefCOCO+, and RefCOCOg datasets manifest the prowess of the Universal Relocalizer, showcasing its formidable performance across the board.
Meng Liu 0006, Xuemeng Song, Da Cao, Zan Gao 0001, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Deconfounded Multimodal Learning for Spatio-temporal Video Grounding
abstract
The task of spatio-temporal video grounding involves identifying the spatial and temporal regions in a video that correspond to the objects or actions described in a given textual description. However, current models used for spatio-temporal video grounding often rely heavily on spatio-temporal priors to make the predictions. As a result, they may suffer from spurious correlations and lack the ability to generalize well to new or diverse scenarios. To overcome this limitation, we introduce a deconfounded multimodal learning framework, which utilizes a structural causal model to treat dataset biases as a confounder and subsequently remove their confounding effect. Through this framework, we can perform causal intervention on the multimodal input and derive an unbiased estimation formula through the do-calculus technique. In order to tackle the challenge of diverse and often unobservable confounders, we further propose a novel retrieval-based approach with a causal mask mechanism. The proposed method leverages analogical reasoning to facilitate deconfounded learning and mitigate dataset biases, enabling unbiased spatio-temporal prediction without explicitly modeling the confounding factors. Extensive experiments on two challenging benchmarks have well verified the effectiveness and rationality of our proposed solution.
Jiawei Wang 0025, Zhanchang Ma, Da Cao, Yuquan Le, Junbin Xiao, Tat-Seng Chua
ACM Multimedia3
2023 Keyword-Based Diverse Image Retrieval With Variational Multiple Instance Graph
abstract
The task of cross-modal image retrieval has recently attracted considerable research attention. In real-world scenarios, keyword-based queries issued by users are usually short and have broad semantics. Therefore, semantic diversity is as important as retrieval accuracy in such user-oriented services, which improves user experience. However, most typical cross-modal image retrieval methods based on single point query embedding inevitably result in low semantic diversity, while existing diverse retrieval approaches frequently lead to low accuracy due to a lack of cross-modal understanding. To address this challenge, we introduce an end-to-end solution termed variational multiple instance graph (VMIG), in which a continuous semantic space is learned to capture diverse query semantics, and the retrieval task is formulated as a multiple instance learning problems to connect diverse features across modalities. Specifically, a query-guided variational autoencoder is employed to model the continuous semantic space instead of learning a single-point embedding. Afterward, multiple instances of the image and query are obtained by sampling in the continuous semantic space and applying multihead attention, respectively. Thereafter, an instance graph is constructed to remove noisy instances and align cross-modal semantics. Finally, heterogeneous modalities are robustly fused under multiple losses. Extensive experiments on two real-world datasets have well verified the effectiveness of our proposed solution in both retrieval accuracy and semantic diversity.
Yawen Zeng, Dongliang Liao, Gongfu Li, Jin Xu 0014, Da Cao, Hong Man
IEEE Trans. Neural Networks Learn. Syst.7
2022 TriReID: Towards Multi-Modal Person Re-Identification via Descriptive Fusion Model
abstract
The cross-modal person re-identification (ReID) aims to retrieve one person from one modality to the other single modality, such as text-based and sketch-based ReID tasks. However, for these different modalities of describing a person, combining multiple aspects can obviously make full use of complementary information and improve the identification performance. Therefore, to explore how to comprehensively consider multi-modal information, we advance a novel multi-modal person re-identification task, which utilizes both text and sketch as a descriptive query to retrieve desired images. In fact, the textual description and the visual description are understood together to retrieve the person in the database to be more aligned with real-world scenarios, which is promising but seldom considered. Besides, based on an existing sketch-based ReID dataset, we construct a new dataset, TriReID, to support this challenging task in a semi-automated way. Particularly, we implement an image captioning model under the active learning paradigm to generate sentences suitable for ReID, in which the quality scores of the three levels are customized. Moreover, we propose a novel framework named Descriptive Fusion Model (DFM) to solve the multi-modal ReID issue. Specifically, we first develop a flexible descriptive embedding function to fuse the text and sketch modalities. Further, the fused descriptive semantic feature is jointly optimized under the generative adversarial paradigm to mitigate the cross-modal semantic gap. Extensive experiments on the TriReID dataset demonstrate the effectiveness and rationality of our proposed solution.
Yajing Zhai, Yawen Zeng, Da Cao, Shaofei Lu
ICMR3
2022 Vision talks: Visual relationship-enhanced transformer for video-guided machine translation
Yawen Zeng, Da Cao, Shaofei Lu
Expert Syst. Appl.3
2022 Video-guided machine translation via dual-level back-translation
Yawen Zeng, Da Cao, Shaofei Lu
Knowl. Based Syst.3
2022 Moment is Important: Language-Based Video Moment Retrieval via Adversarial Learning
abstract
The newly emerging language-based video moment retrieval task aims at retrieving a target video moment from an untrimmed video given a natural language as the query. It is more applicable in reality since it is able to accurately localize a specific video moment, as compared to traditional whole video retrieval. In this work, we propose a novel solution to thoroughly investigate the language-based video moment retrieval issue under the adversarial learning. The key of our solution is to formulate the language-based video moment retrieval task as an adversarial learning problem with two tightly connected components. Specifically, a reinforcement learning is employed as a generator to produce a set of possible video moments. Meanwhile, a multi-task learning is utilized as a discriminator, which integrates inter-modal and intra-modal in a unified framework by employing a sequential update strategy. Finally, the generator and the discriminator are mutually reinforced in the adversarial learning, which is able to jointly optimize the performance of both video moment ranking and video moment localization. Extensive experimental results on two challenging benchmarks, i.e., Charades-STA and TACoS datasets, have well demonstrated the effectiveness and rationality of our proposed solution. Meanwhile, on the larger and unbiased datasets, i.e., ActivityNet Captions and ActivityNet-CD, our proposed framework exhibits excellent robustness.
Yawen Zeng, Da Cao, Shaofei Lu, Hanling Zhang, Jiao Xu 0001, Zheng Qin 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Multi-Modal Relational Graph for Cross-Modal Video Moment Retrieval
abstract
Given an untrimmed video and a query sentence, cross-modal video moment retrieval aims to rank a video moment from pre-segmented video moment candidates that best matches the query sentence. Pioneering work typically learns the representations of the textual and visual content separately and then obtains the interactions or alignments between different modalities. However, the task of cross-modal video moment retrieval is not yet thoroughly addressed as it needs to further identify the fine-grained differences of video moment candidates with high repeatability and similarity. Moveover, the relation among objects in both video and sentence is intuitive and efficient for understanding semantics but is rarely considered.Toward this end, we contribute a multi-modal relational graph to capture the interactions among objects from the visual and textual content to identify the differences among similar video moment candidates. Specifically, we first introduce a visual relational graph and a textual relational graph to form relation-aware representations via message propagation. Thereafter, a multi-task pre-training is designed to capture domain-specific knowledge about objects and relations, enhancing the structured visual representation after explicitly defined relation. Finally, the graph matching and boundary regression are employed to perform the cross-modal retrieval. We conduct extensive experiments on two datasets about daily activities and cooking activities, demonstrating significant improvements over state-of-the-art solutions.
Yawen Zeng, Da Cao, Xiaochi Wei, Meng Liu 0006, Zhou Zhao 0001, Zheng Qin 0001
CVPR2
2021 M2GUDA: Multi-Metrics Graph-Based Unsupervised Domain Adaptation for Cross-Modal Hashing
abstract
Cross-modal hashing is a critical but very challenging task that is to retrieve similar samples of one modality via queries of other modalities. To improve the unsupervised cross-modal hashing, domain adaptation techniques can be used to support unsupervised hashing learning by transferring semantic knowledge from labeled source domain to unlabeled target domain. However, there are two problems that cannot be ignored: (1) most of domain adaptation based researches mainly focused on unimodal hashing or cross-modal real value-based retrieval but the study for cross-modal hashing is limited; (2) most existing studies only consider one or two consistency constraints during the domain adaptation learning. To this end, this paper propose a novel end-to-end framework to realize unsupervised domain adaptation for cross-modal hashing. This method, dubbed M$^2$GUDA, including four different consistency constraints: structure consistency, domain consistency, semantic consistency and modality consistency for domain adaptation learning. Besides, to enhance the structure consistency learning, we develop a multi-metrics graph modeling method to capture structure information comprehensively. Extensive experiments are performed on three common used benchmarks to evaluate the effectivity of our method. The results show that our method outperforms several state-of-the-art cross-modal hashing methods.
Chengyuan Zhang 0001, Lei Zhu 0005, Shichao Zhang 0001, Da Cao
ICMR5
2021 Social-Enhanced Attentive Group Recommendation
abstract
With the proliferation of social networks, group activities have become an essential ingredient of our daily life. A growing number of users share their group activities online and invite their friends to join in. This imposes the need of an in-depth study on the group recommendation task, i.e., recommending items to a group of users. Despite its value and significance, group recommendation remains an unsolved problem due to 1) the weights of group members are crucial to the recommendation performance but are rarely learnt from data; 2) social followee information is beneficial to understand users' preferences but is rarely considered; and 3) user-item interactions are helpful to reinforce the performance of group recommendation but are seldom investigated. Toward this end, we devise neural network-based solutions by utilizing the recent developments of attention network and neural collaborative filtering (NCF). First of all, we adopt an attention network to form the representation of a group by aggregating the group members' embeddings, which allows the attention weights of group members to be dynamically learnt from data. Second, the social followee information is incorporated via another attention network to enhance the representation of individual user, which is helpful to capture users' personal preferences. Third, considering that many online group systems also have abundant interactions of individual users on items, we further integrate the modeling of user-item interactions into our method. Through this way, the recommendation for groups and users can be mutually reinforced. Extensive experiments on the scope of both macro-level performance comparison and micro-level analyses justify the effectiveness and rationality of our proposed approaches.
Da Cao, Xiangnan He 0001, Lianhai Miao, Guangyi Xiao 0001, Hao Chen 0051, Jia Xu 0005
IEEE Trans. Knowl. Data Eng.1
2020 STRONG: Spatio-Temporal Reinforcement Learning for Cross-Modal Video Moment Localization
abstract
In this article, we tackle the cross-modal video moment localization issue, namely, localizing the most relevant video moment in an untrimmed video given a sentence as the query. The majority of existing methods focus on generating video moment candidates with the help of multi-scale sliding window segmentation. They hence inevitably suffer from numerous candidates, which result in the less effective retrieval process. In addition, the spatial scene tracking is crucial for realizing the video moment localization process, but it is rarely considered in traditional techniques. To this end, we innovatively contribute a spatial-temporal reinforcement learning framework. Specifically, we first exploit a temporal-level reinforcement learning to dynamically adjust the boundary of localized video moment instead of the traditional window segmentation strategy, which is able to accelerate the localization process. Thereafter, a spatial-level reinforcement learning is proposed to track the scene on consecutive image frames, therefore filtering out less relevant information. Lastly, an alternative optimization strategy is proposed to jointly optimize the temporal- and spatial-level reinforcement learning. Thereinto, the two tasks of temporal boundary localization and spatial scene tracking are mutually reinforced. By experimenting on two real-world datasets, we demonstrate the effectiveness and rationality of our proposed solution.
Da Cao, Yawen Zeng, Meng Liu 0006, Xiangnan He 0001, Meng Wang 0001, Zheng Qin 0001
ACM Multimedia1
2020 Adversarial Video Moment Retrieval by Jointly Modeling Ranking and Localization
abstract
Retrieving video moments from an untrimmed video given a natural language as the query is a challenging task in both academia and industry. Although much effort has been made to address this issue, traditional video moment ranking methods are unable to generate reasonable video moment candidates and video moment localization approaches are not applicable to large-scale retrieval scenario. How to combine ranking and localization into a unified framework to overcome their drawbacks and reinforce each other is rarely considered. Toward this end, we contribute a novel solution to thoroughly investigate the video moment retrieval issue under the adversarial learning paradigm. The key of our solution is to formulate the video moment retrieval task as an adversarial learning problem with two tightly connected components. Specifically, a reinforcement learning is employed as a generator to produce a set of possible video moments. Meanwhile, a pairwise ranking model is utilized as a discriminator to rank the generated video moments and the ground truth. Finally, the generator and the discriminator are mutually reinforced in the adversarial learning framework, which is able to jointly optimize the performance of both video moment ranking and video moment localization. Extensive experiments on two well-known datasets have well verified the effectiveness and rationality of our proposed solution.
Da Cao, Yawen Zeng, Xiaochi Wei, Liqiang Nie, Richang Hong, Zheng Qin 0001
ACM Multimedia1
2020 Context-Aware Multi-View Summarization Network for Image-Text Matching
abstract
Image-text matching is a vital yet challenging task in the field of multimedia analysis. Over the past decades, great efforts have been made to bridge the semantic gap between the visual and textual modalities. Despite the significance and value, most prior work is still confronted with a multi-view description challenge, i.e., how to align an image to multiple textual descriptions with semantic diversity. Toward this end, we present a novel context-aware multi-view summarization network to summarize context-enhanced visual region information from multiple views. To be more specific, we design an adaptive gating self-attention module to extract representations of visual regions and words. By controlling the internal information flow, we are able to adaptively capture context information. Afterwards, we introduce a summarization module with a diversity regularization to aggregate region-level features into image-level ones from different perspectives. Ultimately, we devise a multi-view matching scheme to match multi-view image features with corresponding text ones. To justify our work, we have conducted extensive experiments on two benchmark datasets, i.e., Flickr30K and MS-COCO, which demonstrates the superiority of our model as compared to several state-of-the-art baselines.
Leigang Qu, Meng Liu 0006, Da Cao, Liqiang Nie, Qi Tian 0001
ACM Multimedia3
2020 Multi-modal product title compression
Lianhai Miao, Da Cao, Weili Guan
Inf. Process. Manag.2
2020 Video-based recipe retrieval
Da Cao, Ning Han 0005, Hao Chen 0051, Xiaochi Wei, Xiangnan He 0001
Inf. Sci.1
2020 Cross-modal recipe retrieval via parallel- and cross-attention networks learning
Da Cao, Jingjing Chu, Ningbo Zhu, Liqiang Nie
Knowl. Based Syst.1
2020 Hashtag our stories: Hashtag recommendation for micro-videos via harnessing multiple modalities
Da Cao, Lianhai Miao, Huigui Rong, Zheng Qin 0001, Liqiang Nie
Knowl. Based Syst.1
2020 Gated and attentive neural collaborative filtering for user generated list recommendation
Chao Yang 0015, Lianhai Miao, Bin Jiang 0006, Dongsheng Li 0002, Da Cao
Knowl. Based Syst.5
2020 A Deep Transfer Learning Solution for Food Material Recognition Using Electronic Scales
abstract
In this article, we present a novel solution to automating the procurement of food materials by using electronic scales, which can automatically identify the food materials along weighing them. Although the CNN model is regarded as one of the most effective solutions to image recognition, the traditional techniques cannot handle the mismatch problem between the lab training data and the real world data. To solve the problem, we propose to embed a partial-and-imbalanced domain adaptation technique (tree adaptation network) in the deep learning model, which can borrow knowledge from sibling classes, to overcome the imbalance problem, and transfer knowledge from the source domain to the target domain, to fight the mismatch problem between the lab training data and the real world data. Experiments show that the proposed approach outperforms state-of-the-art algorithms. Furthermore, the proposed techniques have already been used in practice.
Guangyi Xiao 0001, Hao Chen 0051, Da Cao, Jingzhi Guo, Zhiguo Gong
IEEE Trans. Ind. Informatics4
2019 Video-Based Cross-Modal Recipe Retrieval
abstract
As a natural extension of image-based cross-modal recipe retrieval, retrieving a specific video given a recipe as the query is seldom explored. There are various temporal and spatial elements hidden in cooking videos. In addition, current image-based cross-modal recipe retrieval approaches mostly emphasize the understanding of textual and visual content independently. Such methods overlook the interaction between textual and visual content. In this work, we innovatively propose a new problem of video-based cross-modal recipe retrieval and thoroughly investigate this issue under the attention paradigm. In particular, we firstly exploit a parallel-attention network to independently learn the representations of videos and recipes. Next, a co-attention network is proposed to explicitly emphasize the cross-modal interactive features between videos and recipes. Meanwhile, a cross-modal fusion sub-network is proposed to learn both the independent and collaborative dynamics, which can enhance the associated representation of videos and recipes. Last but not the least, the embedding vectors of videos and recipes stemming from joint network are optimized with a pairwise ranking loss. Extensive experiments on a self-collected dataset have verified the effectiveness and rationality of our proposed solution.
Da Cao, Zhiwang Yu, Hanling Zhang, Jiansheng Fang, Liqiang Nie, Qi Tian 0001
ACM Multimedia1
2019 MOC: Measuring the Originality of Courseware in Online Education Systems
abstract
In online education systems, the courseware plays a pivotal role in helping educators present and impart knowledge to students. The originality of courseware heavily impacts the choice of educators, because the teaching content evolves and so does courseware. However, how to measure the originality of a courseware is a challenging task, due to the lack of labels and the difficulty of quantification. To this end, we contribute a similarity ranking-based unsupervised approach to measure the originality of a courseware. In particular, we first exploit a pre-trained deep visual-text embedding to obtain the representations of images and texts in a local manner. Next, inspired by the design of capsule neural network, a vector-based pooling network is proposed to learn multimodal representations of images and texts. Finally, we propose a Discriminator to optimize the model by maximizing the mutual information between local features and global features in an unsupervised manner. To evaluate the performance of our proposed model, we further subtly collect a dataset for evaluating the originality of courseware by treating sequential versions of each courseware as ranking lists. Therefore, the learning-to-rank scheme can be utilized to evaluate the similarity-based ranking performance. Extensive experimental results have demonstrated the superiority of our proposed framework as compared to other state-of-the-art competitors.
Jiawei Wang 0025, Jiansheng Fang, Jiao Xu 0001, Da Cao, Ming Yang 0039
ACM Multimedia5
2019 The Retrieval of the Beautiful: Self-Supervised Salient Object Detection for Beauty Product Retrieval
abstract
Beauty product retrieval is a challenging task due to the severe image variation issue in real-world scenes. In this work, to mitigate the data variation problem, we contribute a background-agnostic feature extractor, which is trained by a self-supervised salient object detection method. In particular, we first propose a foreground augmentation technique to acquire the augmentation image with its foreground mask. Next, a feature extractor with an attention pooling layer is proposed to learn background-agnostic representations by performing the salient object detection in a self-supervised manner. Finally, we ensemble the background-agnostic features of multiple models to perform the beauty product retrieval. Extensive experimental results have demonstrated the superiority of our proposed framework.
Jiawei Wang 0025, Shuai Zhu, Jiao Xu 0001, Da Cao
ACM Multimedia4
2019 Virtually Trying on New Clothing with Arbitrary Poses
abstract
Thanks to the recent advance in the multimedia techniques, increasing research attention has been paid to the virtual try-on task, especially with the 2D image modeling. The traditional try-on task aims to align the target clothing item naturally to the given person's body and hence present a try-on look of the person. However, in practice, people may also be interested in their try-on looks with different poses. Therefore, in this work, we introduce a new try-on setting, which enables the changes of both the clothing item and the person's pose. Towards this end, we propose a pose-guided virtual try-on scheme based on the generative adversarial networks (GANs) with a bi-stage strategy. In particular, in the first stage, we propose a shape enhanced clothing deformation model for deforming the clothing item, where the user body shape is incorporated as the intermediate guidance. For the second stage, we present an attentive bidirectional GAN, which jointly models the attentive clothing-person alignment and bidirectional generation consistency. For evaluation, we create a large-scale dataset, FashionTryOn, comprising $28,714$ triplets with each consisting of a clothing item image and two model images in different poses. Extensive experiments on FashionTryOn validate the superiority of our model over the state-of-the-art methods.
Xuemeng Song, Zhaozheng Chen, Linmei Hu, Da Cao, Liqiang Nie
ACM Multimedia5
2019 Multi-criteria active deep learning for image classification
Jin Yuan 0002, Xingxing Hou, Yaoqiang Xiao, Da Cao, Weili Guan, Liqiang Nie
Knowl. Based Syst.4
2018 Learner Behavioral Feature Refinement and Augmentation Using GANs
Da Cao, Andrew S. Lan, Christopher G. Brinton, Mung Chiang
AIED (2)1
2018 Principles for Assessing Adaptive Online Courses
Carlee Joe-Wong, Christopher G. Brinton, Liang Zheng 0002, Da Cao
EDM5
2018 Behavioral Analysis at Scale: Learning Course Prerequisite Structures from Learner Clickstreams
Andrew S. Lan, Da Cao, Christopher G. Brinton, Mung Chiang
EDM3
2018 Attentive Group Recommendation
abstract
Due to the prevalence of group activities in people's daily life, recommending content to a group of users becomes an important task in many information systems. A fundamental problem in group recommendation is how to aggregate the preferences of group members to infer the decision of a group. Toward this end, we contribute a novel solution, namely AGREE (short for ''Attentive Group REcommEndation''), to address the preference aggregation problem by learning the aggregation strategy from data, which is based on the recent developments of attention network and neural collaborative filtering (NCF). Specifically, we adopt an attention mechanism to adapt the representation of a group, and learn the interaction between groups and items from data under the NCF framework. Moreover, since many group recommender systems also have abundant interactions of individual users on items, we further integrate the modeling of user-item interactions into our method. Through this way, we can reinforce the two tasks of recommending items for both groups and users. By experimenting on two real-world datasets, we demonstrate that our AGREE model not only improves the group recommendation performance but also enhances the recommendation for users, especially for cold-start users that have no historical interactions individually.
Da Cao, Xiangnan He 0001, Lianhai Miao, Yahui An, Chao Yang 0015, Richang Hong
SIGIR1
2018 On the Efficiency of Online Social Learning Networks
Christopher G. Brinton, Swapna Buccapatnam, Liang Zheng 0002, Da Cao, Andrew S. Lan, Felix Ming Fai Wong, Sangtae Ha, Mung Chiang, H. Vincent Poor
IEEE/ACM Trans. Netw.4
2017 Behavior in social learning networks: Early detection for online short-courses
abstract
We study learning outcome prediction for online courses. Whereas prior work has focused on semester-long courses with frequent student assessments, we focus on short-courses that have single outcomes assigned by instructors at the end. The lack of performance data makes the behavior of learners, captured as they interact with course content and with one another in Social Learning Networks (SLN), essential for prediction. Our method defines several (machine) learning features based on behaviors collected on the modes of (human) learning in a course, and uses them in appropriate classifiers. Through evaluation on data captured from three two-week courses hosted through our delivery platforms, we make three key observations: (i) behavioral data is predictive of learning outcomes in short-courses (our classifiers achieving AUCs ≥ 0.8 after the two weeks), (ii) it has an early detection capability (AUCs ≥ 0.7 with the first week of data), and (iii) the content features have an “earliest” detection capability (with higher AUC in the first few days), while the SLN features become the more predictive set over time, as the network matures. We also discuss how our method can generate behavioral analytics for instructors.
Christopher G. Brinton, Da Cao, Mung Chiang
INFOCOM3
2017 Embedding Factorization Models for Jointly Recommending Items and User Generated Lists
abstract
Existing recommender algorithms mainly focused on recommending individual items by utilizing user-item interactions. However, little attention has been paid to recommend user generated lists (e.g., playlists and booklists). On one hand, user generated lists contain rich signal about item co-occurrence, as items within a list are usually gathered based on a specific theme. On the other hand, a user's preference over a list also indicate her preference over items within the list. We believe that 1) if the rich relevance signal within user generated lists can be properly leveraged, an enhanced recommendation for individual items can be provided, and 2) if user-item and user-list interactions are properly utilized, and the relationship between a list and its contained items is discovered, the performance of user-item and user-list recommendations can be mutually reinforced.
Da Cao, Liqiang Nie, Xiangnan He 0001, Xiaochi Wei, Shunzhi Zhu, Tat-Seng Chua
SIGIR1
2017 Version-sensitive mobile App recommendation
Da Cao, Liqiang Nie, Xiangnan He 0001, Xiaochi Wei, Jialie Shen 0001, Shunxiang Wu, Tat-Seng Chua
Inf. Sci.1
2017 Cross-Platform App Recommendation by Jointly Modeling Ratings and Texts
abstract
Over the last decade, the renaissance of Web technologies has transformed the online world into an application (App) driven society. While the abundant Apps have provided great convenience, their sheer number also leads to severe information overload, making it difficult for users to identify desired Apps. To alleviate the information overloading issue, recommender systems have been proposed and deployed for the App domain. However, existing work on App recommendation has largely focused on one single platform (e.g., smartphones), while it ignores the rich data of other relevant platforms (e.g., tablets and computers). In this article, we tackle the problem of cross-platform App recommendation, aiming at leveraging users’ and Apps’ data on multiple platforms to enhance the recommendation accuracy. The key advantage of our proposal is that by leveraging multiplatform data, the perpetual issues in personalized recommender systems—data sparsity and cold-start—can be largely alleviated. To this end, we propose a hybrid solution, STAR (short for “croSs-plaTform App Recommendation”) that integrates both numerical ratings and textual content from multiple platforms. In STAR, we innovatively represent an App as an aggregation of common features across platforms (e.g., App’s functionalities) and specific features that are dependent on the resided platform. In light of this, STAR can discriminate a user’s preference on an App by separating the user’s interest into two parts (either in the App’s inherent factors or platform-aware features). To evaluate our proposal, we construct two real-world datasets that are crawled from the App stores of iPhone, iPad, and iMac. Through extensive experiments, we show that our STAR method consistently outperforms highly competitive recommendation methods, justifying the rationality of our cross-platform App recommendation proposal and the effectiveness of our solution.
Da Cao, Xiangnan He 0001, Liqiang Nie, Xiaochi Wei, Xia Ben Hu, Shunxiang Wu, Tat-Seng Chua
ACM Trans. Inf. Syst.1