VLDB 2026 Research / reviewers in the wild / expert
Ge Luo 0003
dblp:321/6589
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0009-0005-2757-1879ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Smooth-Guided Implicit Data Augmentation for Domain GeneralizationabstractThe training process of a domain generalization (DG) model involves utilizing one or more interrelated source domains to attain optimal performance on an unseen target domain. Existing DG methods often use auxiliary networks or require high computational costs to improve the model's generalization ability by incorporating a diverse set of source domains. In contrast, this work proposes a method called Smooth-Guided Implicit Data Augmentation (SGIDA) that operates in the feature space to capture the diversity of source domains. To amplify the model's generalization capacity, a distance metric learning (DML) loss function is incorporated. Additionally, rather than depending on deep features, the suggested approach employs logits produced from cross entropy (CE) losses with infinite augmentations. A theoretical analysis shows that logits are effective in estimating distances defined on original features, and the proposed approach is thoroughly analyzed to provide a better understanding of why logits are beneficial for DG. Moreover, to increase the diversity of the source domain, a sampling-based method called smooth is introduced to obtain semantic directions from interclass relations. The effectiveness of the proposed approach is demonstrated through extensive experiments on widely used DG, object detection, and remote sensing datasets, where it achieves significant improvements over existing state-of-the-art methods across various backbone networks. Mengzhu Wang, Junze Liu, Ge Luo 0003, Shanshan Wang 0008, Wei Wang 0335, Long Lan, Ye Wang 0023, Feiping Nie 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Engaging Live Video Comments GenerationabstractMultimodal Dialogue agents are often required to respond to conversation history using both textual and visual content. Even though current dialogue studies predominantly strive to generate natural texts or images, they fall short in considering the relevance of multimodal responses within a dialogue context, consequently confining agents from making prudent choices based on multiple alternatives and their associated relevance scores for decision-making. In this paper, we present a bidirectional multimodal dialogue framework that skillfully combines the forward generation of multiple text and image response candidates with reverse selection guided by relevance scores evaluated on dialogue context, facilitating agents in selecting the most suitable multimodal responses. Specifically, the forward generation aspect of our framework leverages a stage-wise approach, first producing textual replies and composite visual descriptions from the dialogue context, followed by the generation of visual responses aligned with the descriptions. In the reverse selection process, visual responses are translated into tangible descriptive texts that, in conjunction with textual responses, are inversely tied back to the dialogue context for relevance assessment, assigning a reference score to each multimodal response candidate to assist the intelligent agent in making informed decisions. Experimental outcomes demonstrate that our proposed bidirectional dialogue response framework markedly elevates performance in both automatic and human evaluations, yielding a range of contextually fitting multimodal responses for selection. Ge Luo 0003, Junqiang Huang, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001 |
ACM Multimedia | 1 |
| 2024 | Disentangled Style Domain for Implicit z-Watermark Towards Copyright ProtectionabstractText-to-image models have shown surprising performance in high-quality image generation, while also raising intensified concerns about the unauthorized usage of personal dataset in training and personalized fine-tuning. Recent approaches, embedding watermarks, introducing perturbations, and inserting backdoors into datasets, rely on adding minor information vulnerable to adversarial training, limiting their ability to detect unauthorized data usage. In this paper, we introduce a novel implicit Zero-Watermarking scheme that first utilizes the disentangled style domain to detect unauthorized dataset usage in text-to-image models. Specifically, our approach generates the watermark from the disentangled style domain, enabling self-generalization and mutual exclusivity within the style domain anchored by protected units. The domain achieves the maximum concealed offset of probability distribution through both the injection of identifier $z$ and dynamic contrastive learning, facilitating the structured delineation of dataset copyright boundaries for multiple sources of styles and contents. Additionally, we introduce the concept of watermark distribution to establish a verification mechanism for copyright ownership of hybrid or partial infringements, addressing deficiencies in the traditional mechanism of dataset copyright ownership for AI mimicry. Notably, our method achieves one-sample verification for copyright ownership in AI mimic generations. The code is available at: [https://github.com/Hlufies/ZWatermarking](https://github.com/Hlufies/ZWatermarking) Junqiang Huang, Zhaojun Guo, Ge Luo 0003, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001 |
NeurIPS | 3 |
| 2023 | Forward Creation, Reverse Selection: Achieving Highly Pertinent Multimodal Responses in Dialogue ContextsabstractMultimodal Dialogue agents are often required to respond to conversation history using both textual and visual content. Even though current dialogue studies predominantly strive to generate natural texts or images, they fall short in considering the relevance of multimodal responses within a dialogue context, consequently confining agents from making prudent choices based on multiple alternatives and their associated relevance scores for decision-making. In this paper, we present a bidirectional multimodal dialogue framework that skillfully combines the forward generation of multiple text and image response candidates with reverse selection guided by relevance scores evaluated on dialogue context, facilitating agents in selecting the most suitable multimodal responses. Specifically, the forward generation aspect of our framework leverages a stage-wise approach, first producing textual replies and composite visual descriptions from the dialogue context, followed by the generation of visual responses aligned with the descriptions. In the reverse selection process, visual responses are translated into tangible descriptive texts that, in conjunction with textual responses, are inversely tied back to the dialogue context for relevance assessment, assigning a reference score to each multimodal response candidate to assist the intelligent agent in making informed decisions. Experimental outcomes demonstrate that our proposed bidirectional dialogue response framework markedly elevates performance in both automatic and human evaluations, yielding a range of contextually fitting multimodal responses for selection. Ge Luo 0003, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001 |
CIKM | 1 |
| 2023 | VCMaster: Generating Diverse and Fluent Live Video Comments Based on Multimodal ContextsabstractLive video commenting, or "bullet screen," is a popular social style on video platforms. Automatic live commenting has been explored as a promising approach to enhance the appeal of videos. However, existing methods neglect the diversity of generated sentences, limiting the potential to obtain human-like comments. In this paper, we introduce a novel framework called "VCMaster" for multimodal live video comments generation, which balances the diversity and quality of generated comments to create human-like sentences. We involve images, subtitles, and contextual comments as inputs to better understand complex video contexts. Then, we propose an effective Hierarchical Cross-Fusion Decoder to integrate high-quality trimodal feature representations by cross-fusing critical information from previous layers. Additionally, we develop a Sentence-Level Contrastive Loss to enlarge the distance between generated and contextual comments by contrastive learning. It helps the model to avoid the pitfall of simply imitating provided contextual comments and losing creativity, encouraging the model to achieve more diverse comments while maintaining high quality. We also construct a large-scale multimodal live video comments dataset with 292,507 comments and three sub-datasets that cover nine general categories. Extensive experiments demonstrate that our model achieves a level of human-like language expression and remarkably fluent, diverse, and engaging generated comments compared to baselines. Ge Luo 0003, Sheng Li 0006, Zhenxing Qian, Xinpeng Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | Image Steganalysis with Convolutional Vision TransformerabstractRecent research has shown that deep learning based methods offer more accurate detection for image steganalysis than the traditional detection paradigm based on rich media models. Existing network architectures based on deep learning, however, stack more and more convolutional layers to increase local receptive fields for image stegananlysis. Limited by hardware, the detector with several convolutional layers may not extract features of steganography images from a global perspective effectively. In this paper, we propose a Convolutional Vision Transformer for image stegananlysis, which can capture both local and global dependencies among noise features. In image processing phase, our network preserves CNN frame for its capacity of producing image noise residuals. Different from previous methods, we utilize the attention mechanism of vision transformer for feature extraction and classification. The proposed network is validated on two public image datasets (BOSSbase 1.01 and ALASKA #2). Experimental results demonstrate that our network performs well over fixed-size dataset and arbitrary-size dataset. Ge Luo 0003, Ping Wei 0004, Shuwen Zhu, Xinpeng Zhang 0001, Zhenxing Qian, Sheng Li 0006 |
ICASSP | 1 |
| 2022 | Joint Learning for Addressee Selection and Response Generation in Multi-Party ConversationabstractA large number of multi-party conversation scenarios exist in social networks, which have been seldom studied in the field of human-machine conversation. In this paper, we study a novel task of joint learning for addressee selection and response generation in multi-party conversations. Systems are expected to select whom they address and generate the corresponding response. To solve it, we propose an end-to-end addressee selection and response generation (ASRG) model, containing an addressee selection module and a response generation module. In the selection module, we develop an addressee prediction attention scheme to obtain a unique context vector for each candidate, thereby calculating the probability of the candidate more accurately. In the generation module, we propose a Focus Transformer to generate responses. These two modules are jointly learnt to fully explore the correlations between addressee and response. Experimental results show ASRG remarkably outperforms baselines and generates relevant content for different addressees. Sheng Li 0006, Ping Wei 0004, Ge Luo 0003, Xinpeng Zhang 0001, Zhenxing Qian |
ICASSP | 4 |
| 2022 | Generative Steganographic FlowabstractGenerative steganography (GS) is a new data hiding manner, featuring direct generation of stego media from secret data. Existing GS methods are generally criticized for their poor performances. In this paper, we propose a novel flow based GS approach - Generative Steganographic Flow (GSF), which provides direct generation of stego images without cover image. We take the stego image generation and secret data recovery process as an invertible transformation, and build a reversible bijective mapping between input secret data and generated stego images. In the forward mapping, secret data is hidden in the input latent of Glow model to generate stego images. By reversing the mapping, hidden data can be extracted exactly from generated stego images. Furthermore, we propose a novel latent optimization strategy to improve the fidelity of stego images. Experimental results show our proposed GSF has far better performances than SOTA works. Ping Wei 0004, Ge Luo 0003, Xinpeng Zhang 0001, Zhenxing Qian, Sheng Li 0006 |
ICME | 2 |
| 2022 | Generative Steganography NetworkabstractSteganography usually modifies cover media to embed secret data. A new steganographic approach called generative steganography (GS) has emerged recently, in which stego images (images containing secret data) are generated from secret data directly without cover media. However, existing GS schemes are often criticized for their poor performances. In this paper, we propose an advanced generative steganography network (GSN) that can generate realistic stego images without using cover images. We firstly introduce the mutual information mechanism in GS, which helps to achieve high secret extraction accuracy. Our model contains four sub-networks, i.e., an image generator (G), a discriminator (D), a steganalyzer (S), and a data extractor (E). D and S act as two adversarial discriminators to ensure the visual quality and security of generated stego images. E is to extract the hidden secret from generated stego images. The generator G is flexibly constructed to synthesize either cover or stego images with different inputs. It facilitates covert communication by concealing the function of generating stego images in a normal generator. A module named secret block is designed to hide secret data in the feature maps during image generation, with which high hiding capacity and image fidelity are achieved. In addition, a novel hierarchical gradient decay (HGD) skill is developed to resist steganalysis detection. Experiments demonstrate the superiority of our work over existing methods. Ping Wei 0004, Sheng Li 0006, Xinpeng Zhang 0001, Ge Luo 0003, Zhenxing Qian |
ACM Multimedia | 4 |
| 2022 | Breaking Robust Data Hiding in Online Social NetworksabstractSome robust data hiding approaches have been proposed to transmit secret data through online social networks (OSNs). Traditional steganalysis tools are inefficient in detecting these tailored steganographic methods. Although some algorithms have been developed to remove the hidden data of images, they are criticized for their low secret removal rate and poor image quality. Moreover, most of them are nongeneric methods, and multiple models must be trained to fit different algorithms. In this letter, we propose a general end-to-end data hiding break approach for OSNs, called the secret data remover (SDR). It is universal for algorithms of both robust steganography and robust watermarking, with which stego images are directly input and clean ones with the same appearances will then be generated. Moreover, we develop two novel techniques, namely, image fusion and latent renewal, to enhance the image quality and improve the overall performance. Experiments show that our proposed method achieves superior performance compared to state-of-the-art works. Hidden secret data are cleared while image quality is maintained or even slightly improved. At the same time, our work can be easily deployed in OSNs. Ping Wei 0004, Zhiying Zhu 0001, Ge Luo 0003, Zhenxing Qian, Xinpeng Zhang 0001 |
IEEE Signal Process. Lett. | 3 |