Zhongwei Xie

dblp:164/7444 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
abstract
Novel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of reconstructing photorealistic 3D scenes from sparse or single-view inputs. We introduce SpatialCrafter, a framework that leverages the rich knowledge in video diffusion models to generate plausible additional observations, thereby alleviating reconstruction ambiguity. Through a trainable camera encoder and an epipolar attention mechanism for explicit geometric constraints, we achieve precise camera control and 3D consistency, further reinforced by a unified scale estimation strategy to handle scale discrepancies across datasets. Furthermore, by integrating monocular depth priors with semantic features in the video latent space, our framework directly regresses 3D Gaussian primitives and efficiently processes long-sequence features using a hybrid network structure. Extensive experiments show our method enhances sparse view reconstruction and restores the realistic appearance of 3D scenes.
Songchun Zhang, Huiyao Xu, Sitong Guo, Zhongwei Xie, Hujun Bao, Weiwei Xu 0003, Changqing Zou
ICCV4
2024 Video Frame-wise Explanation Driven Contrastive Learning for Procedural Text Generation
Zhongwei Xie, Chuanbo Liu
Comput. Vis. Image Underst.3
2024 Boosting Healthiness Exposure in Category-Constrained Meal Recommendation Using Nutritional Standards
abstract
Food computing, a newly emerging topic, is closely linked to human life through computational methodologies. Meal recommendation, a food-related study about human health, aims to provide users a meal with courses constrained from specific categories (e.g., appetizers, main dishes) that can be enjoyed as a service. Historical interaction data, important user information, is often used by existing models to learn user preferences. However, if a user’s preferences favor less healthy meals, the model will follow that preference and make similar recommendations, potentially negatively impacting the user’s long-term health. This emphasizes the necessity for health-oriented and responsible meal recommendation systems. In this article, we propose a healthiness-aware and category-wise meal recommendation model called CateRec, which boosts healthiness exposure by using nutritional standards as knowledge to guide the model training. Two fundamental questions are raised and answered: (1) How can the healthiness of meals be evaluated? Two well-known nutritional standards from the World Health Organization and the United Kingdom Food Standards Agency are used to calculate the healthiness score of the meal. (2) How can the model training be guided in a health-oriented manner? We construct category-wise personalization partial rankings and category-wise healthiness partial rankings, and theoretically analyze that they meet the necessary properties and assumptions required to be trained by the maximum posterior estimator under Bayesian probability. The data analysis confirms the existence of user preferences leaning towards less healthy meals in two public datasets. A comprehensive experiment demonstrates that our CateRec effectively boosts healthiness exposure in terms of mean healthiness score and ranking exposure while being comparable to the state-of-the-art model in terms of recommendation accuracy.
Ming Li 0072, Lin Li 0001, Xiaohui Tao 0001, Zhongwei Xie, Qing Xie 0002, Jingling Yuan
ACM Trans. Intell. Syst. Technol.4
2023 Transductive Cross-Lingual Scene-Text Visual Question Answering
Lin Li 0001, Haohan Zhang 0002, Zeqin Fang, Zhongwei Xie, Jianquan Liu
ICONIP (6)4
2022 Cross-Modal Attention Networks with Modality Disentanglement for Scene-Text VQA
abstract
Understanding the texts in visual scenes is essential for reasoning tasks when texts carry key information. Therefore, it is a common way to learn multi-modal representation from visual scenes to support Scene-Text Visual Question&Answering. Different modalities, such as text and image, are embedded into a joint semantic space where attention mechanism is widely applied. This paper regards Scene-Text VQA as a kind of cross-modal task where the joint embedding between a scene image and its text answer is a semantic bridge to share the strong semantic clues with the two modalities. To this end, our proposed framework firstly divides the scene recognition features into two types: visual and textual features. And then external cross-modal pre-trained features are introduced to simultaneously guide the representation of visual and text features through cross-modal attention net-works. Finally, Transformer is applied as a decoder to output answers. Conducted on a public Scene-Text VQA dataset, experimental results show that our proposed framework out-performs existing approaches in terms of accuracy.
Zeqin Fang, Lin Li 0001, Zhongwei Xie, Jingling Yuan
ICME3
2022 Cross-Modal Retrieval between Event-Dense Text and Image
abstract
This paper presents a novel approach to the problem of event-dense text and image cross-modal retrieval where the text contains the descriptions of numerous events. It is known that modality alignment is crucial for retrieval performance. However, due to the lack of event sequence information in the image, it is challenging to perform the fine-grain alignment of the event-dense text with the image. Our proposed approach incorporates the event-oriented features to enhance the cross-modal alignment, and applies the event-dense text-image retrieval to the food domain for empirical validation. Specifically, we capture the significance of each event by Transformer, and combine it with the identified key event elements, to enhance the discriminative ability of the learned text embedding that summarizes all the events. Next, we produce the image embedding by combining the event tag jointly shared by the text and image with the visual embedding of the event-related image regions, which describes the eventual consequence of all the events and facilitates the event-based cross-modal alignment. Finally, we integrate text embedding and image embedding with the loss optimization empowered with the event tag by iteratively regulating the joint embedding learning for cross-modal retrieval. Extensive experiments demonstrate that our proposed event-oriented modality alignment approach significantly outperforms the state-of-the-art approach with a 23.3% improvement on top-1 Recall for image-to-recipe retrieval on Recipe1M 10k test set.
Zhongwei Xie, Lin Li 0001, Luo Zhong, Jianquan Liu, Ling Liu 0001
ICMR1
2022 Learning Text-image Joint Embedding for Efficient Cross-modal Retrieval with Deep Feature Engineering
abstract
This article introduces a two-phase deep feature engineering framework for efficient learning of semantics enhanced joint embedding, which clearly separates the deep feature engineering in data preprocessing from training the text-image joint embedding model. We use the Recipe1M dataset for the technical description and empirical validation. In preprocessing, we perform deep feature engineering by combining deep feature engineering with semantic context features derived from raw text-image input data. We leverage LSTM to identify key terms, deep NLP models from the BERT family, TextRank, or TF-IDF to produce ranking scores for key terms before generating the vector representation for each key term by using Word2vec. We leverage Wide ResNet50 and Word2vec to extract and encode the image category semantics of food images to help semantic alignment of the learned recipe and image embeddings in the joint latent space. In joint embedding learning, we perform deep feature engineering by optimizing the batch-hard triplet loss function with soft-margin and double negative sampling, taking into account also the category-based alignment loss and discriminator-based alignment loss. Extensive experiments demonstrate that our SEJE approach with deep feature engineering significantly outperforms the state-of-the-art approaches.
Zhongwei Xie, Ling Liu 0001, Yanzhao Wu 0001, Luo Zhong, Lin Li 0001
ACM Trans. Inf. Syst.1
2022 Learning TFIDF Enhanced Joint Embedding for Recipe-Image Cross-Modal Retrieval Service
abstract
It is widely acknowledged that learning joint embeddings of recipes with images is challenging due to the diverse composition and deformation of ingredients in cooking procedures. We present a Multi-modal Semantics enhanced Joint Embedding approach (MSJE) for learning a common feature space between the two modalities (text and image), with the ultimate goal of providing high-performance cross-modal retrieval services. Our MSJE approach has three unique features. First, we extract the TFIDF feature from the title, ingredients and cooking instructions of recipes. By determining the significance of word sequences through combining LSTM learned features with their TFIDF features, we encode a recipe into a TFIDF weighted vector for capturing significant key terms and how such key terms are used in the corresponding cooking instructions. Second, we combine the recipe TFIDF feature with the recipe sequence feature extracted through two-stage LSTM networks, which is effective in capturing the unique relationship between a recipe and its associated image(s). Third, we further incorporate TFIDF enhanced category semantics to improve the mapping of image modality and to regulate the similarity loss function during the iterative learning of cross-modal joint embedding. Experiments on the benchmark dataset Recipe1M show the proposed approach outperforms the state-of-the-art approaches.
Zhongwei Xie, Ling Liu 0001, Yanzhao Wu 0001, Lin Li 0001, Luo Zhong
IEEE Trans. Serv. Comput.1
2021 Learning Joint Embedding with Modality Alignments for Cross-Modal Retrieval of Recipes and Food Images
abstract
This paper presents a three-tier modality alignment approach to learning text-image joint embedding, coined as JEMA, for cross-modal retrieval of cooking recipes and food images. The first tier improves recipe text embedding by optimizing the LSTM networks with term extraction and ranking enhanced sequence patterns, and optimizes the image embedding by combining the ResNeXt-101 image encoder with the category embedding using wideResNet-50 with word2vec. The second tier modality alignment optimizes the textual-visual joint embedding loss function using a double batch-hard triplet loss with soft-margin optimization. The third modality alignment incorporates two types of cross-modality alignments as the auxiliary loss regularizations to further reduce the alignment errors in the joint learning of the two modality-specific embedding functions. The category-based cross-modal alignment aims to align the image category with the recipe category as a loss regularization to the joint embedding. The cross-modal discriminator-based alignment aims to add the visual-textual embedding distribution alignment to further regularize the joint embedding loss. Extensive experiments with the one-million recipes benchmark dataset Recipe1M demonstrate that the proposed JEMA approach outperforms the state-of-the-art cross-modal embedding methods for both image-to-recipe and recipe-to-image retrievals.
Zhongwei Xie, Ling Liu 0001, Lin Li 0001, Luo Zhong
CIKM1
2021 Boosting Ensemble Accuracy by Revisiting Ensemble Diversity Metrics
abstract
Neural network ensembles are gaining popularity by harnessing the complementary wisdom of multiple base models. Ensemble teams with high diversity promote high failure independence, which is effective for boosting the overall ensemble accuracy. This paper provides an in-depth study on how to design and compute ensemble diversity, which can capture the complementary decision capacity of ensemble member models. We make three original contri-butions. First, we revisit the ensemble diversity metrics in the literature and analyze the inherent problems of poor correlation between ensemble diversity and ensemble ac-curacy, which leads to the low quality ensemble selection using such diversity metrics. Second, instead of computing diversity scores for ensemble teams of different sizes using the same criteria, we introduce focal model based ensemble diversity metrics, coined as FQ-diversity metrics. Our new metrics significantly improve the intrinsic correlation between high ensemble diversity and high ensemble accuracy. Third, we introduce a diversity fusion method, coined as the EQ-diversity metric, by integrating the top three most representative FQ-diversity metrics. Comprehensive experiments on two benchmark datasets (CIFAR-10 and ImageNet) show that our FQ and EQ diversity metrics are effective for selecting high diversity ensemble teams to boost overall ensemble accuracy.
Yanzhao Wu 0001, Ling Liu 0001, Zhongwei Xie, Ka-Ho Chow 0001, Wenqi Wei 0001
CVPR3
2021 Efficient Deep Feature Calibration for Cross-Modal Joint Embedding Learning
abstract
This paper introduces a two-phase deep feature calibration framework for efficient learning of semantics enhanced text-image cross-modal joint embedding, which clearly separates the deep feature calibration in data preprocessing from training the joint embedding model. We use the Recipe1M dataset for the technical description and empirical validation. In preprocessing, we perform deep feature calibration by combining deep feature engineering with semantic context features derived from raw text-image input data. We leverage LSTM to identify key terms, NLP methods to produce ranking scores for key terms before generating the key term feature. We leverage wideResNet50 to extract and encode the image category semantics to help semantic alignment of the learned recipe and image embeddings in the joint latent space. In joint embedding learning, we perform deep feature calibration by optimizing the batch-hard triplet loss function with soft-margin and double negative sampling, also utilizing the category-based alignment loss and discriminator-based alignment loss. Extensive experiments demonstrate that our SEJE approach with the deep feature calibration significantly outperforms the state-of-the-art approaches.
Zhongwei Xie, Ling Liu 0001, Lin Li 0001, Luo Zhong
ICMI1
2020 primirTSS: an R package for identifying cell-specific microRNA transcription start sites
abstract
SUMMARY: The R/Bioconductor package primirTSS is a fast and convenient tool that allows implementation of the analytical method to identify transcription start sites of microRNAs by integrating ChIP-seq data of H3K4me3 and Pol II. It further ensures the precision by employing the conservation score and sequence features. The tool showed a good performance when using H3K4me3 or Pol II Chip-seq data alone as input, which brings convenience to applications where multiple datasets are hard to acquire. This flexible package is provided with both R-programming interfaces as well as graphical web interfaces. AVAILABILITY AND IMPLEMENTATION: primirTSS is available at: http://bioconductor.org/packages/primirTSS. The documentation of the package including an accompanying tutorial was deposited at: https://bioconductor.org/packages/release/bioc/vignettes/primirTSS/inst/doc/primirTSS.html. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pumin Li, Xu Hua, Zhongwei Xie
Bioinform.4
2020 Image-to-video person re-identification with cross-modal embeddings
Zhongwei Xie, Lin Li 0001, Xian Zhong, Luo Zhong, Jianwen Xiang
Pattern Recognit. Lett.1
2019 Enhancing multimodal deep representation learning by fixed model reuse
Zhongwei Xie, Lin Li 0001, Xian Zhong, Yang He 0003, Luo Zhong
Multim. Tools Appl.1
2018 A Hybrid Model Reuse Training Approach for Multilingual OCR
Zhongwei Xie, Lin Li 0001, Xian Zhong, Luo Zhong, Qing Xie 0002, Jianwen Xiang
WISE (1)1
2017 Co-training an Improved Recurrent Neural Network with Probability Statistic Models for Named Entity Recognition
Yueqing Sun, Lin Li 0001, Zhongwei Xie, Qing Xie 0002, Xin Li 0064, Guandong Xu
DASFAA (2)3
2015 WeChecker: efficient and precise detection of privilege escalation vulnerabilities in Android apps
abstract
Due to the rapid increase of Android apps and their wide usage to handle personal data, a precise and large-scaling checker is in need to validate the apps' permission flow before they are listed on the market. Several tools have been proposed to detect sensitive data leaks in Android apps. But these tools are not applicable to large-scale analysis since they fail to deal with the arbitrary execution orders of different event handlers smartly. Event handlers are invoked by the framework based on the system state, therefore we cannot pre-determine their order of execution. Besides, since all exported components can be invoked by an external app, the execution orders of these components are also arbitrary. A naive way to simulate these two types of arbitrary execution orders yields a permutation of all event handlers in an app. The time complexity is O(n!) where n is the number of event handlers in an app. This leads to a high analysis overhead when n is big. To give an illustration, CHEX [10] found 50.73 entry points of 44 unique class types in an app on average. In this paper we propose an improved static taint analysis to deal with the challenge brought by the arbitrary execution orders without sacrificing the high precision. Our analysis does not need to make permutations and achieves a polynomial time complexity. We also propose to unify the array and map access with object reference by propagating access paths to reduce the number of false positives due to field-insensitivity and over approximation of array access and map access.
Xingmin Cui, Lucas C. K. Hui, Zhongwei Xie, Tian Zeng, Siu-Ming Yiu
WISEC4