EDBT 2026 Demo / reviewers in the wild / expert
Luo Zhong
dblp:62/6237
· DBLP profile ↗
28ranked-venue papers
2as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Computer networks · 3Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fragrant: frequency-auxiliary guided relational attention network for low-light action recognition
Wenxuan Liu 0008, Xuemei Jia, Yihao Ju, Yakun Ju, Kui Jiang, Shifeng Wu, Luo Zhong, Xian Zhong |
Vis. Comput. | 7 |
| 2022 | Visual-Aware Attention Dual-Stream Decoder for Video CaptioningabstractVideo captioning is a challenging task that captures different visual parts and describes them in sentences, for it requires visual and linguistic coherence. The attention mechanism in the current video captioning method learns to assign weight to each frame, promoting the decoder dynamically. This may not explicitly model the correlation and the temporal coherence of the visual features extracted in the sequence frames. To generate semantically coherent sentences, we propose a new Visual-aware Attention (VA) model, which concatenates dynamic changes of temporal sequence frames with the words at the previous moment, as the input of attention mechanism to extract sequence features. In addition, the prevalent approaches widely use the Teacher-forcing (TF) learning during training, where the next token is generated conditioned on the previous ground-truth tokens. The semantic information in the previously generated tokens is lost. Therefore, we design a Self-forcing (SF) stream that takes the semantic information in the probability distribution of the previous token as input to enhance the current token. The Dual-stream Decoder (DD) architecture unifies the TF and SF streams, generating sen-tences to promote the annotated captioning for both streams. Meanwhile, with the Dual-stream Decoder utilized, the ex-posure bias problem is alleviated, caused by the discrepancy between the training and testing in the TF learning. The effectiveness of the proposed Visual-aware Attention Dual-stream Decoder (VADD) is demonstrated through the result of ex-perimental studies on Microsoft video description (MSVD) corpus and MSR-Video to text (MSR-VTT) datasets. Zhixin Sun, Shuqin Chen, Luo Zhong |
ICME | 3 |
| 2022 | Cross-Modal Retrieval between Event-Dense Text and ImageabstractThis paper presents a novel approach to the problem of event-dense text and image cross-modal retrieval where the text contains the descriptions of numerous events. It is known that modality alignment is crucial for retrieval performance. However, due to the lack of event sequence information in the image, it is challenging to perform the fine-grain alignment of the event-dense text with the image. Our proposed approach incorporates the event-oriented features to enhance the cross-modal alignment, and applies the event-dense text-image retrieval to the food domain for empirical validation. Specifically, we capture the significance of each event by Transformer, and combine it with the identified key event elements, to enhance the discriminative ability of the learned text embedding that summarizes all the events. Next, we produce the image embedding by combining the event tag jointly shared by the text and image with the visual embedding of the event-related image regions, which describes the eventual consequence of all the events and facilitates the event-based cross-modal alignment. Finally, we integrate text embedding and image embedding with the loss optimization empowered with the event tag by iteratively regulating the joint embedding learning for cross-modal retrieval. Extensive experiments demonstrate that our proposed event-oriented modality alignment approach significantly outperforms the state-of-the-art approach with a 23.3% improvement on top-1 Recall for image-to-recipe retrieval on Recipe1M 10k test set. Zhongwei Xie, Lin Li 0001, Luo Zhong, Jianquan Liu, Ling Liu 0001 |
ICMR | 3 |
| 2022 | Joint Re-Detection and Re-Identification for Multi-Object Tracking
Xian Zhong, Jingling Yuan, Luo Zhong |
MMM (1) | 6 |
| 2022 | Learning Text-image Joint Embedding for Efficient Cross-modal Retrieval with Deep Feature EngineeringabstractThis article introduces a two-phase deep feature engineering framework for efficient learning of semantics enhanced joint embedding, which clearly separates the deep feature engineering in data preprocessing from training the text-image joint embedding model. We use the Recipe1M dataset for the technical description and empirical validation. In preprocessing, we perform deep feature engineering by combining deep feature engineering with semantic context features derived from raw text-image input data. We leverage LSTM to identify key terms, deep NLP models from the BERT family, TextRank, or TF-IDF to produce ranking scores for key terms before generating the vector representation for each key term by using Word2vec. We leverage Wide ResNet50 and Word2vec to extract and encode the image category semantics of food images to help semantic alignment of the learned recipe and image embeddings in the joint latent space. In joint embedding learning, we perform deep feature engineering by optimizing the batch-hard triplet loss function with soft-margin and double negative sampling, taking into account also the category-based alignment loss and discriminator-based alignment loss. Extensive experiments demonstrate that our SEJE approach with deep feature engineering significantly outperforms the state-of-the-art approaches. Zhongwei Xie, Ling Liu 0001, Yanzhao Wu 0001, Luo Zhong, Lin Li 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2022 | Learning TFIDF Enhanced Joint Embedding for Recipe-Image Cross-Modal Retrieval ServiceabstractIt is widely acknowledged that learning joint embeddings of recipes with images is challenging due to the diverse composition and deformation of ingredients in cooking procedures. We present a Multi-modal Semantics enhanced Joint Embedding approach (MSJE) for learning a common feature space between the two modalities (text and image), with the ultimate goal of providing high-performance cross-modal retrieval services. Our MSJE approach has three unique features. First, we extract the TFIDF feature from the title, ingredients and cooking instructions of recipes. By determining the significance of word sequences through combining LSTM learned features with their TFIDF features, we encode a recipe into a TFIDF weighted vector for capturing significant key terms and how such key terms are used in the corresponding cooking instructions. Second, we combine the recipe TFIDF feature with the recipe sequence feature extracted through two-stage LSTM networks, which is effective in capturing the unique relationship between a recipe and its associated image(s). Third, we further incorporate TFIDF enhanced category semantics to improve the mapping of image modality and to regulate the similarity loss function during the iterative learning of cross-modal joint embedding. Experiments on the benchmark dataset Recipe1M show the proposed approach outperforms the state-of-the-art approaches. Zhongwei Xie, Ling Liu 0001, Yanzhao Wu 0001, Lin Li 0001, Luo Zhong |
IEEE Trans. Serv. Comput. | 5 |
| 2021 | Learning Joint Embedding with Modality Alignments for Cross-Modal Retrieval of Recipes and Food ImagesabstractThis paper presents a three-tier modality alignment approach to learning text-image joint embedding, coined as JEMA, for cross-modal retrieval of cooking recipes and food images. The first tier improves recipe text embedding by optimizing the LSTM networks with term extraction and ranking enhanced sequence patterns, and optimizes the image embedding by combining the ResNeXt-101 image encoder with the category embedding using wideResNet-50 with word2vec. The second tier modality alignment optimizes the textual-visual joint embedding loss function using a double batch-hard triplet loss with soft-margin optimization. The third modality alignment incorporates two types of cross-modality alignments as the auxiliary loss regularizations to further reduce the alignment errors in the joint learning of the two modality-specific embedding functions. The category-based cross-modal alignment aims to align the image category with the recipe category as a loss regularization to the joint embedding. The cross-modal discriminator-based alignment aims to add the visual-textual embedding distribution alignment to further regularize the joint embedding loss. Extensive experiments with the one-million recipes benchmark dataset Recipe1M demonstrate that the proposed JEMA approach outperforms the state-of-the-art cross-modal embedding methods for both image-to-recipe and recipe-to-image retrievals. Zhongwei Xie, Ling Liu 0001, Lin Li 0001, Luo Zhong |
CIKM | 4 |
| 2021 | Efficient Deep Feature Calibration for Cross-Modal Joint Embedding LearningabstractThis paper introduces a two-phase deep feature calibration framework for efficient learning of semantics enhanced text-image cross-modal joint embedding, which clearly separates the deep feature calibration in data preprocessing from training the joint embedding model. We use the Recipe1M dataset for the technical description and empirical validation. In preprocessing, we perform deep feature calibration by combining deep feature engineering with semantic context features derived from raw text-image input data. We leverage LSTM to identify key terms, NLP methods to produce ranking scores for key terms before generating the key term feature. We leverage wideResNet50 to extract and encode the image category semantics to help semantic alignment of the learned recipe and image embeddings in the joint latent space. In joint embedding learning, we perform deep feature calibration by optimizing the batch-hard triplet loss function with soft-margin and double negative sampling, also utilizing the category-based alignment loss and discriminator-based alignment loss. Extensive experiments demonstrate that our SEJE approach with the deep feature calibration significantly outperforms the state-of-the-art approaches. Zhongwei Xie, Ling Liu 0001, Lin Li 0001, Luo Zhong |
ICMI | 4 |
| 2021 | Local-enhanced Multi-resolution Representation Learning for Vehicle Re-identificationabstractIn real traffic scenarios, the changes of vehicle resolution that the camera captures tend to be relatively obvious considering the distances to the vehicle, different directions, and height of the camera. When the resolution difference exists between the probe and the gallery vehicle, the resolution mismatch will occur, which will seriously influence the performance of the vehicle re-identification (Re-ID). This problem is also known as multi-resolution vehicle Re-ID. An effective strategy is equivalent to utilize image super-resolution to handle the resolution gap. However, existing methods conduct super-resolution on global images instead of local representation of each image, leading to much more noisy information generated from the background and illumination variations. In our work, a local-enhanced multi-resolution representation learning (LMRL) is therefore proposed to address these problems by combining the training of local-enhanced super-resolution (LSR) module and local-guided contrastive learning (LCL) module. Specifically, we use a parsing network to parse a vehicle into four different parts to extract local-enhanced vehicle representation. And then, the LSR module, which consists of two auto-encoders that share parameters, transforms low-resolution images into high-resolution in both global and local branches. LCL module can learn discriminative vehicle representation by contrasting local representation between the high-resolution reconstructed image and the ground truth. We evaluate our approach on two public datasets that contain vehicle images at a wide range of resolutions, in which our approach shows significant superiority to the existing solution. Xian Zhong, Jingling Yuan, Rongbo Zhang, Duxiu Feng, Luo Zhong |
MMAsia | 7 |
| 2021 | Subspace Enhancement and Colorization Network for Infrared Video Action Recognition
Xian Zhong, Wenxuan Liu 0008, Zhengwei Yang 0001, Luo Zhong |
PRICAI (3) | 6 |
| 2020 | A staged adaptive firefly algorithm for UAV charging planning in wireless sensor networks
Linhui Cheng, Luo Zhong, Xiao Zhang 0006, Jiaxu Xing |
Comput. Commun. | 2 |
| 2020 | An emotion classification algorithm based on SPT-CapsNet
Xian Zhong, Jinhang Liu, Lin Li 0001, Shuqin Chen, Yuyu Dong, Bingqing Wu, Luo Zhong |
Neural Comput. Appl. | 8 |
| 2020 | Adaptively Converting Auxiliary Attributes and Textual Embedding for Video Captioning Based on BiLSTM
Shuqin Chen, Xian Zhong, Lin Li 0001, Luo Zhong |
Neural Process. Lett. | 6 |
| 2020 | Image-to-video person re-identification with cross-modal embeddings
Zhongwei Xie, Lin Li 0001, Xian Zhong, Luo Zhong, Jianwen Xiang |
Pattern Recognit. Lett. | 4 |
| 2020 | Image inspired Chinese couplet generationabstractChinese couplets, as one of the traditional Chinese culture, is the treasure of Chinese civilization and the inheritance of Chinese history. Given a sentence (namely an antecedent clause), people reply with another sentence (namely a subsequent clause) equal in length. Because of the complexity of the semantic and grammatical rules of couplet, it is not easy to create a suitable couplet that meets the requirements of sentence pattern, context, and flatness. With the development of neural models and natural language processing, automatic generation of Chinese couplets has drawn significant attention due to its artistic and cultural value, most of these works mainly focus on generating couplet by given text information, while visual inspirations for couplet generation have been rarely explored. In this paper, we design a Chinese couplet generation model based on NIC (Neural Image Caption), which can compose a piece of couplet suitable to the artistic conception in an image. At first, we use the improved VGG16 model to predict the input image. The content of the image can be automatically recognized and the corresponding description are generated and translated into Chinese keywords. Then, the encoder-decoder framework is used repeatedly to process these keywords, and finally the couplet can be generated. Moreover, to satisfy special characteristics of couplets, we incorporate the attention mechanism into the encoding-decoding process, which greatly improves the accuracy of couplets generated automatically. Shengqiong Yuan, Luo Zhong, Lin Li 0001 |
Web Intell. | 2 |
| 2020 | An Encoder-Decoder Network Based FCN Architecture for Semantic SegmentationabstractIn recent years, the convolutional neural network (CNN) has made remarkable achievements in semantic segmentation. The method of semantic segmentation has a desirable application prospect. Nowadays, the methods mostly use an encoder-decoder architecture as a way of generating pixel by pixel segmentation prediction. The encoder is for extracting feature maps and decoder for recovering feature map resolution. An improved semantic segmentation method on the basis of the encoder-decoder architecture is proposed. We can get better segmentation accuracy on several hard classes and reduce the computational complexity significantly. This is possible by modifying the backbone and some refining techniques. Finally, after some processing, the framework has achieved good performance in many datasets. In comparison with the traditional architecture, our architecture does not need additional decoding layer and further reuses the encoder weight, thus reducing the complete quantity of parameters needed for processing. In this paper, a modified focal loss function is also put forward, as a replacement for the cross-entropy function to achieve a better treatment of the imbalance problem of the training data. In addition, more context information is added to the decode module as a way of improving the segmentation results. Experiments prove that the presented method can get better segmentation results. As an integral part of a smart city, multimedia information plays an important role. Semantic segmentation is an important basic technology for building a smart city. Yongfeng Xing, Luo Zhong, Xian Zhong |
Wirel. Commun. Mob. Comput. | 2 |
| 2019 | Enhancing multimodal deep representation learning by fixed model reuse
Zhongwei Xie, Lin Li 0001, Xian Zhong, Yang He 0003, Luo Zhong |
Multim. Tools Appl. | 5 |
| 2018 | A Hybrid Model Reuse Training Approach for Multilingual OCR
Zhongwei Xie, Lin Li 0001, Xian Zhong, Luo Zhong, Qing Xie 0002, Jianwen Xiang |
WISE (1) | 4 |
| 2018 | Image Restoration via Group l2, 1 Norm-Based Structural Sparse RepresentationabstractSparse representation has recently been extensively studied in the field of image restoration. Many sparsity-based approaches enforce sparse coding on patches with certain constraints. However, extracting structural information is a challenging task in the field image restoration. Motivated by the fact that structured sparse representation (SSR) method can capture the inner characteristics of image structures, which helps in finding sparse representations of nonlinear features or patterns, we propose the SSR approach for image restoration. Specifically, a generalized model is developed using structured restraint, namely, the group [Formula: see text]-norm of the coefficient matrix is introduced in the traditional sparse representation with respect to minimizing the differences within classes and maximizing the differences between classes for sparse representation, and its applications with image restoration are also explored. The sparse coefficients of SSR are obtained through iterative optimization approach. Experimental results have shown that the proposed SSR technique can significantly deliver the reconstructed images with high quality, which manifest the effectiveness of our approach in both peak signal-to-noise ratio performance and visual perception. Kaisong Zhang, Luo Zhong, Xuan Ya Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2018 | Unsupervised Keyword Extraction from Microblog Posts via Hashtags
Lin Li 0001, Jinghang Liu, Yueqing Sun, Guandong Xu, Jingling Yuan, Luo Zhong |
J. Web Eng. | 6 |
| 2017 | Understanding Trajectory Data Based on Heterogeneous Information Network Using Visual Analytics
Rui Zhang 0066, Luo Zhong, Hongbo Jiang 0001 |
MSN | 3 |
| 2013 | QUBiC: An adaptive approach to query-based recommendation
Lin Li 0001, Luo Zhong, Zhenglu Yang, Masaru Kitsuregawa |
J. Intell. Inf. Syst. | 2 |
| 2012 | Activity-centered design for temporal data managementabstractWith our interviews, we found that professors need to organize and use fine-grained information from diverse sources and formats in their everyday tasks, such as taking a meeting, hosting a project and etc. Although there is a bunch of computer applications and devices which can help them easily manage all these cross-media information, professors are still often confused about reviewing their old information as time goes on. In this paper, we present activity-centered learning design, a meaningful structure for task information, to assist professors with their present and future teaching tasks. We also illustrate a prototype of the Activity construct from the end-user's perspective. Shengqiong Yuan, Luo Zhong |
CSCWD | 2 |
| 2012 | A feature-free search query classification approach using semantic distance
Lin Li 0001, Luo Zhong, Guandong Xu, Masaru Kitsuregawa |
Expert Syst. Appl. | 2 |
| 2009 | Grey Neural Network Based Predictive Model for Multi-core Architecture 2D Spatial Characteristics
Jingling Yuan, Luo Zhong |
ISNN (1) | 4 |
| 2009 | An Improved Greedy Based Global Optimized Placement Algorithm
Luo Zhong, Kejing Wang, Jingling Yuan |
ISNN (4) | 1 |
| 2008 | Multiple Sources Data Fusion Strategies Based on Multi-class Support Vector Machine
Luo Zhong, Zichun Ding, Cuicui Guo, Huazhu Song |
ISNN (1) | 1 |
| 2005 | Structural Damage Detection by Integrating Independent Component Analysis and Support Vector Machine
Huazhu Song, Luo Zhong |
ADMA | 2 |