Jianmin Wu

dblp:23/2490 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
8since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 4Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2023 ReProMSig: an integrative platform for development and application of reproducible multivariable models for cancer prognosis supporting guideline-based transparent reporting
abstract
Adequate reporting is essential for evaluating the performance and clinical utility of a prognostic prediction model. Previous studies indicated a prevalence of incomplete or suboptimal reporting in translational and clinical studies involving development of multivariable prediction models for prognosis, which limited the potential applications of these models. While reporting templates introduced by the established guidelines provide an invaluable framework for reporting prognostic studies uniformly, there is a widespread lack of qualified adherence, which may be due to miscellaneous challenges in manual reporting of extensive model details, especially in the era of precision medicine. Here, we present ReProMSig (Reproducible Prognosis Molecular Signature), a web-based integrative platform providing the analysis framework for development, validation and application of a multivariable prediction model for cancer prognosis, using clinicopathological features and/or molecular profiles. ReProMSig platform supports transparent reporting by presenting both methodology details and analysis results in a strictly structured reporting file, following the guideline checklist with minimal manual input needed. The generated reporting file can be published together with a developed prediction model, to allow thorough interrogation and external validation, as well as online application for prospective cases. We demonstrated the utilities of ReProMSig by development of prognostic molecular signatures for stage II and III colorectal cancer respectively, in comparison with a published signature reproduced by ReProMSig. Together, ReProMSig provides an integrated framework for development, evaluation and application of prognostic/predictive biomarkers for cancer in a more transparent and reproducible way, which would be a useful resource for health care professionals and biomedical researchers.
Lihua Cao, Jiafu Ji, David K. Chang, Jianmin Wu
Briefings Bioinform.5
2023 Real-time object tracking in the wild with Siamese network
Shaokui Jiang, Jianmin Wu, Baile Xu, Jian Zhao 0013, Furao Shen
Multim. Tools Appl.3
2023 Correction to: Real-time object tracking in the wild with Siamese network
Shaokui Jiang, Jianmin Wu, Baile Xu, Jian Zhao 0013, Furao Shen
Multim. Tools Appl.3
2023 Switch and Refine: A Long-Term Tracking and Segmentation Framework
abstract
In long-term video object tracking (VOT) tasks, most long-term trackers are modified from short-term trackers, which contain more and more machine learning modules to improve their performance. However, we empirically find that more modules do not necessarily lead to better results. In this paper, we make the long-term tracking framework simple by carefully selecting the cutting-edge trackers. Specifically, we propose a new long-term VOT framework that combines the benefits of two mainstream short-term tracking pipelines, i.e., the discriminative online tracker and the one-shot Siamese tracker, with a global re-detector awakened when the target is lost. Such a framework fully exploits existing advanced works from three complementary perspectives. Experimental results show that by exploiting the capabilities of existing methods instead of designing new neural networks, we can still achieve remarkable results on seven long-term VOT datasets. By introducing a continuous adjustable speed control parameter, our tracker reaches 20+FPS with only a small performance loss. The refine module not only improves the bounding box estimations but also outputs segmentation masks, so that our framework can handle the video object segmentation (VOS) tasks by using only VOT trackers. We obtain a trade-off between time and accuracy on two representative VOS datasets by only using bounding boxes as the initial input.
Jian Zhao 0013, Jianmin Wu, Furao Shen
IEEE Trans. Circuits Syst. Video Technol.3
2022 Deeply Exploit Visual and Language Information for Social Media Popularity Prediction
abstract
Social media popularity prediction task is to predict future attractiveness of new posts, which could be applied for online advertising, social recommendation, and demand prediction. Existing methods have explored multiple feature types to model the popularity prediction, including user profile, tag, space-time, category, and others. However, images and texts of social media posts, as important and primary information, are usually used by simple or insufficient processing. In this paper, we propose a method to deeply exploit visual and language information to explore the attractiveness of posts. Specifically, images are parsed from multiple perspectives including multi-modal semantic representation, perceptual image quality, and scene analysis. Different word-level and sentence-level semantic embedding are extracted from all available language texts including title, tags, concept and category. It makes social media popularity modeling more reliable with the powerful visual and language representation. Experimental results demonstrate the effectiveness of exploiting visual and language information by the proposed method, and we achieve new state-of-the-art results on the SMP Challenge at ACM Multimedia 2022.
Jianmin Wu, Dangwei Li, Chen-Wei Xie, Siyang Sun
ACM Multimedia1
2022 Deep Video Understanding with a Unified Multi-Modal Retrieval Framework
abstract
In this paper, we propose a unified multi-modal retrieval framework to tackle two typical video understanding tasks, i.e., matching movie scenes and text descriptions, and scene sentiment classification. For the task of matching movie scenes and text descriptions, it is a natural multi-modal retrieval problem, while for the task of scene sentiment classification, the proposed framework aims at finding most related sentiment tag for each movie scene, which is also a multi-modal retrieval problem. By considering these two tasks as multi-modal retrieval problems, we propose a unified multi-modal retrieval framework, which can make full use of the models pre-trained on large scale multi-modal datasets, experiments show that it is critical for the tasks which have only hundreds of training examples. To further improve the performance on movie video understanding task, we also collect a large scale video-text dataset, which contains 427,603 movie-shot and text pairs. Experimental results validate the effectiveness of this dataset.
Chen-Wei Xie, Siyang Sun, Jianmin Wu, Dangwei Li
ACM Multimedia4
2022 Token Embeddings Alignment for Cross-Modal Retrieval
abstract
Cross-modal retrieval has achieved significant progress in recent years with the help of token embeddings interaction methods. Most existing methods first extract embedding for each token of input image and text, then feed the token-level embeddings into a multi-modal transformer to learn a joint representation, this joint representation can be used to predict matching score between input image and text. However, these methods don't explicitly supervise the alignment between visual and textual tokens. In this paper, we propose a novel Token Embeddings AlignMent (TEAM) block, it first explicitly aligns visual tokens and textual tokens, then produces token-level matching scores to measure fine-grained similarity between input image and text. TEAM achieves new state-of-the-art performance on commonly used cross-modal retrieval benchmarks. Moreover, TEAM is interpretable and we provide visualization experiments to show how it works. At last, we construct a new billion-scale vision-language pre-training dataset in Chinese, which is the largest Chinese vision-language pre-training dataset so far. After pre-training on this dataset, our framework also achieves state-of-the-art performance on Chinese cross-modal retrieval benchmarks.
Chen-Wei Xie, Jianmin Wu, Xian-Sheng Hua 0001
ACM Multimedia2
2021 Exploring Visual-Audio Composition Alignment Network for Quality Fashion Retrieval in Video
abstract
Fashion retrieval in video suffers from the issues of imperfect visual representation and low quality of search results under the E-commercial circumstance. Previous works generally focus on searching the identical images from visual perspective only, but lack of leveraging multi-modal information for high quality commodities. As a cross-domain problem, instructional or exhibiting audio reveals rich semantic information to facilite the video-to-shop task. In this paper, we present a novel Visual-Audio Composition Alignment Network (VACANet) to deal with quality fashion retrieval in video. Firstly, we introduce the visual-audio composition module in VACANet aiming to distinguish attentive and residual entities by learning semantic embedding from both visual and audio streams. Secondly, a quality alignment training scheme is then designed by quality-aware triplet mining and domain alignment constraint for video-to-image adaptation. Finally, extensive experiments conducted on challenging video datasets demonstrate the scalable effectiveness of our model in alleviating quality fashion retrieval.
Yanhao Zhang 0002, Jianmin Wu, Dangwei Li, Chenwei Xie
ICASSP2
2020 Research on the Evaluation Model of Rural Information Demand Based on Big Data
abstract
In recent years, the imbalance of rural information supply and demand has seriously hindered the process of rural informatization. Rural information demand is a decisive factor in the relationship between rural information supply and demand. Therefore, research on the influencing factors of rural information demand has attracted much attention. The traditional rural information demand factor analysis does not consider the correlation between factors. The factors themselves carry a lot of repeated information, which seriously interferes with the objectivity of the analysis results. Proceeding from the complexity and diversity of influencing factors of rural information demand, based on the selected subjective and objective factors, based on the forward partial correlation analysis and post-ROC test, a probit discriminant model of influencing factors of rural information demand was constructed, and the relationship with Lingshou was determined. There are 24 factors that are significantly related to county rural information needs. The research results show that this method not only eliminates the factors that carry highly repetitive information and the correlation is not significant but also makes the results more reliable. At the same time, it also found that rural information supply is related to farmers’ information cognition ability, acceptance awareness, and acceptance ability. This study provides new methods and new ideas for solving related problems.
Yanfeng Jin, Gang Li 0027, Jianmin Wu
Wirel. Commun. Mob. Comput.3
2019 Virtual ID Discovery from E-commerce Media at Alibaba: Exploiting Richness of User Click Behavior for Visual Search Relevance
abstract
Visual search plays an essential role for E-commerce. To meet the search demands of users and promote shopping experience at Alibaba, visual search relevance of real-shot images is becoming the bottleneck. Traditional visual search paradigm is usually based upon supervised learning with labeled data. However, large-scale categorical labels are required with expensive human annotations, which limits its applicability and also usually fails in distinguishing the real-shot images.
Yanhao Zhang 0002, Jianmin Wu, Rong Jin 0001
CIKM5
2014 Challenges in Quantifying Narcotic Use from Drug Dispensing Records
Jianmin Wu, Jon D. Duke, John T. Finnell
AMIA1
2014 WikiReviz: An Edit History Visualization for Wiki Systems
Jianmin Wu, Mizuho Iwaihara
APWeb1
2014 Tracking Topics on Revision Graphs of Wikipedia Edit History
Bonan Li, Jianmin Wu, Mizuho Iwaihara
WAIM2
2013 Evaluating Congruence Between Laboratory LOINC Value Sets for Quality Measures, Public Health Reporting, and Mapping Common Tests
Jianmin Wu, John T. Finnell, Daniel J. Vreeman
AMIA1
2013 A Practical Method for Predicting Frequent Use of Emergency Department Care Using Routinely Available Electronic Registration Data
Jianmin Wu, Huiping Xu, John T. Finnell, Shaun J. Grannis
AMIA1
2013 Revision graph extraction in Wikipedia based on supergram decomposition
abstract
As one of the popular social media that many people turn to in recent years, collaborative encyclopedia Wikipedia provides information in a more "Neutral Point of View" way than others. Towards this core principle, plenty of efforts have been put into collaborative contribution and editing. The trajectories of how such collaboration appears by revisions are valuable for group dynamics and social media research, which suggest that we should extract the underlying derivation relationships among revisions from chronologically-sorted revision history in a precise way. In this paper, we propose a revision graph extraction method based on supergram decomposition in the document collection of near-duplicates. The plain text of revisions would be measured by its frequency distribution of supergram, which is the variable-length token sequence that keeps the same through revisions. We show that this method can effectively perform the task than existing methods.
Jianmin Wu, Mizuho Iwaihara
OpenSym1
2013 Reorder user's tweets
abstract
Twitter displays the tweets a user received in a reversed chronological order, which is not always the best choice. As Twitter is full of messages of very different qualities, many informative or relevant tweets might be flooded or displayed at the bottom while some nonsense buzzes might be ranked higher. In this work, we present a supervised learning method for personalized tweets reordering based on user interests. User activities on Twitter, in terms of tweeting, retweeting, and replying, are leveraged to obtain the training data for reordering models. Through exploring a rich set of social and personalized features, we model the relevance of tweets by minimizing the pairwise loss of relevant and irrelevant tweets. The tweets are then reordered according to the predicted relevance scores. Experimental results with real twitter user activities demonstrated the effectiveness of our method. The new method achieved above 30% accuracy gain compared with the default ordering in twitter based on time.
Keyi Shen, Jianmin Wu, Ya Zhang 0002, Yiping Han, Xiaokang Yang 0001, Li Song 0001, Xiao Gu 0001
ACM Trans. Intell. Syst. Technol.2
2012 A multi-cue mean-shift target tracking approach based on fuzzified region dynamic image fusion
Xiao Yun, Jianmin Wu
Sci. China Inf. Sci.3
2012 Training the max-margin sequence model with the relaxed slack variables
Lingfeng Niu, Jianmin Wu, Yong Shi 0001
Neural Networks2
2009 Exploiting term relationship to boost text classification
abstract
Document classification provides an effective way to handle the explosive online textual data. However, in practical classification settings, we face the so-called feature sparsity problem caused by a lack of training documents or the shortness of text to be classified. In this paper, we solve the sparsity problem by exploiting term relationships along with Naive Bayes classifiers. The first method is to estimate term relationships based on the co-occurrence information of two terms in a certain context. The second method estimates the term relationships based on the distribution of terms over different hierarchical categories in a publicly available document taxonomy. Thereafter, term relationship is used to augment Naive Bayes classifiers. We test our methods on two open-domain data sets to demonstrate its advantages. The experimental results show that our method can significantly improve the classification performance, especially when we do not have enough training data or the texts are Web search queries.
Dou Shen, Jianmin Wu, Bin Cao 0001, Jian-Tao Sun, Qiang Yang 0001, Zheng Chen 0001, Ying Li 0040
CIKM2
2008 Learning Bidirectional Similarity for Collaborative Filtering
Bin Cao 0001, Jian-Tao Sun, Jianmin Wu, Qiang Yang 0001, Zheng Chen 0001
ECML/PKDD (1)3