VLDB 2026 Research / reviewers in the wild / expert
Longwen Gao
dblp:149/1242
· DBLP profile ↗
11ranked-venue papers
5as first author
5since 2021 · last 2026
0009-0000-0621-2950ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text SynthesisabstractAbstract Diffusion‐based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large‐scale annotated data to support multilingual generation. In this work, we revisit the necessity of complex auxiliary modules and further explore an approach that simultaneously ensures glyph accuracy and achieves high‐fidelity scene integration, by leveraging diffusion models' inherent capabilities for contextual reasoning. To this end, we introduce TextFlux, a DiT‐based framework that enables multilingual scene text synthesis. The advantages of TextFlux can be summarized as follows: (1) OCR‐free model architecture. TextFlux eliminates the need for OCR encoders that are specifically used to extract visual text‐related features. (2) Strong multilingual scalability. TextFlux is effective in low‐resource multilingual settings, and achieves strong performance in newly added languages with fewer than 1,000 samples. (3) Streamlined training setup. TextFlux is trained with only 1% of the training data required by competing methods. (4) Controllable multi‐line text generation. TextFlux offers flexible multi‐line synthesis with precise line‐level control, outperforming methods restricted to single‐line or rigid layouts. Extensive experiments and visualizations demonstrate that TextFlux outperforms previous methods in both qualitative and quantitative evaluations. Our code is available at https://github.com/yyyyyxie/textflux . Jielei Zhang, Weihang Wang 0011, Longwen Gao, Zhouhui Lian |
Comput. Graph. Forum | 5 |
| 2025 | MX-Font++: Mixture of Heterogeneous Aggregation Experts for Few-shot Font GenerationabstractFew-shot Font Generation (FFG) aims to create new font libraries using limited reference glyphs, with crucial applications in digital accessibility and equity for low-resource languages, especially in multilingual artificial intelligence systems. Although existing methods have shown promising performance, transitioning to unseen characters in low-resource languages remains a significant challenge, especially when font glyphs vary considerably across training sets. MX-Font considers the content of a character from the perspective of a local component, employing a Mixture of Experts (MoE) approach to adaptively extract the component for better transition. However, the lack of a robust feature extractor prevents them from adequately decoupling content and style, leading to sub-optimal generation results. To alleviate these problems, we propose Heterogeneous Aggregation Experts (HAE), a powerful feature extraction expert that helps decouple content and style downstream from being able to aggregate information in channel and spatial dimensions. Additionally, we propose a novel content-style homogeneity loss to enhance the untangling. Extensive experiments on several datasets demonstrate that our MX-Font++ yields superior visual results in FFG and effectively outperforms state-of-the-art methods. Weihang Wang 0011, Duolin Sun, Jielei Zhang, Longwen Gao |
ICASSP | 4 |
| 2023 | Mining and Applying Composition Knowledge of Dance Moves for Style-Concentrated Dance GenerationabstractChoreography refers to creation of dance motions according to both music and dance knowledge, where the created dances should be style-specific and consistent. However, most of the existing methods generate dances using the given music as the only reference, lacking the stylized dancing knowledge, namely, the flag motion patterns contained in different styles. Without the stylized prior knowledge, these approaches are not promising to generate controllable style or diverse moves for each dance style, nor new dances complying with stylized knowledge. To address this issue, we propose a novel music-to-dance generation framework guided by style embedding, considering both input music and stylized dancing knowledge. These style embeddings are learnt representations of style-consistent kinematic abstraction of reference dance videos, which can act as controllable factors to impose style constraints on dance generation in a latent manner. Hence, we can make the style embedding fit into any given style while allowing the flexibility to generate new compatible dance moves by modifying the style embedding according to the learnt representations of a certain style. We are the first to achieve knowledge-driven style control in dance generation tasks. To support this study, we build a large multi-style music-to-dance dataset referred to as I-Dance. The qualitative and quantitative evaluations demonstrate the advantage of the proposed framework, as well as the ability to synthesize diverse moves under a dance style directed by style embedding. Xinjian Zhang, Su Yang 0001, Yi Xu 0003, Weishan Zhang, Longwen Gao |
AAAI | 5 |
| 2023 | Video Compression Artifact Reduction by Fusing Motion Compensation and Global Context in a Swin-CNN Based Parallel ArchitectureabstractVideo Compression Artifact Reduction aims to reduce the artifacts caused by video compression algorithms and improve the quality of compressed video frames. The critical challenge in this task is to make use of the redundant high-quality information in compressed frames for compensation as much as possible. Two important possible compensations: Motion compensation and global context, are not comprehensively considered in previous works, leading to inferior results. The key idea of this paper is to fuse the motion compensation and global context together to gain more compensation information to improve the quality of compressed videos. Here, we propose a novel Spatio-Temporal Compensation Fusion (STCF) framework with the Parallel Swin-CNN Fusion (PSCF) block, which can simultaneously learn and merge the motion compensation and global context to reduce the video compression artifacts. Specifically, a temporal self-attention strategy based on shifted windows is developed to capture the global context in an efficient way, for which we use the Swin transformer layer in the PSCF block. Moreover, an additional Ada-CNN layer is applied in the PSCF block to extract the motion compensation. Experimental results demonstrate that our proposed STCF framework outperforms the state-of-the-art methods up to 0.23dB (27% improvement) on the MFQEv2 dataset. Xinjian Zhang, Su Yang 0001, Wuyang Luo, Longwen Gao, Weishan Zhang |
AAAI | 4 |
| 2021 | GIF Thumbnails: Attract More Clicks to Your VideosabstractWith the rapid increase of mobile devices and online media, more and more people prefer posting/viewing videos online. Generally, these videos are presented on video streaming sites with image thumbnails and text titles. While facing huge amounts of videos, a viewer clicks through a certain video with high probability because of its eye-catching thumbnail. However, current video thumbnails are created manually, which is time-consuming and quality-unguaranteed. And static image thumbnails contain very limited information of the corresponding videos, which prevents users from successfully clicking what they really want to view. In this paper, we address a novel problem, namely GIF thumbnail generation, which aims to automatically generate GIF thumbnails for videos and consequently boost their Click-Through-Rate (CTR). Here, a GIF thumbnail is an animated GIF file consisting of multiple segments from the video, containing more information of the target video than a static image thumbnail. To support this study, we build the first GIF thumbnails benchmark dataset that consists of 1070 videos covering a total duration of 69.1 hours, and 5394 corresponding manually-annotated GIFs. To solve this problem, we propose a learning-based automatic GIF thumbnail generation model, which is called Generative Variational Dual-Encoder (GEVADEN). As not relying on any user interaction information (e.g. time-sync comments and real-time view counts), this model is applicable to newly-uploaded/rarely-viewed videos. Experiments on our built dataset show that GEVADEN significantly outperforms several baselines, including video-summarization and highlight-detection based ones. Furthermore, we develop a pilot application of the proposed model on an online video platform with 9814 videos covering 1231 hours, which shows that our model achieves a 37.5% CTR improvement over traditional image thumbnails. This further validates the effectiveness of the proposed model and the promising application prospect of GIF thumbnails. Yi Xu 0003, Fan Bai 0001, Yingxuan Shi, Qiuyu Chen, Longwen Gao, Kai Tian 0001, Shuigeng Zhou, Huyang Sun |
AAAI | 5 |
| 2019 | Non-Local ConvLSTM for Video Compression Artifact ReductionabstractVideo compression artifact reduction aims to recover high-quality videos from low-quality compressed videos. Most existing approaches use a single neighboring frame or a pair of neighboring frames (preceding and/or following the target frame) for this task. Furthermore, as frames of high quality overall may contain low-quality patches, and high-quality patches may exist in frames of low quality overall, current methods focusing on nearby peak-quality frames (PQFs) may miss high-quality details in low-quality frames. To remedy these shortcomings, in this paper we propose a novel end-to-end deep neural network called non-local ConvLSTM (NL-ConvLSTM in short) that exploits multiple consecutive frames. An approximate non-local strategy is introduced in NL-ConvLSTM to capture global motion patterns and trace the spatiotemporal dependency in a video sequence. This approximate strategy makes the non-local module work in a fast and low space-cost way. Our method uses the preceding and following frames of the target frame to generate a residual, from which a higher quality frame is reconstructed. Experiments on two datasets show that NL-ConvLSTM outperforms the existing methods. Yi Xu 0003, Longwen Gao, Kai Tian 0001, Shuigeng Zhou, Huyang Sun |
ICCV | 2 |
| 2016 | Group and Graph Joint Sparsity for Linked Data ClassificationabstractVarious sparse regularizers have been applied to machine learning problems, among which structured sparsity has been proposed for a better adaption to structured data. In this paper, motivated by effectively classifying linked data (e.g. Web pages, tweets, articles with references, and biological network data) where a group structure exists over the whole dataset and links exist between specific samples, we propose a joint sparse representation model that combines group sparsity and graph sparsity, to select a small number of connected components from the graph of linked samples, meanwhile promoting the sparsity of edges that link samples from different groups in each connected component. Consequently, linked samples are selected from a few sparsely-connected groups. Both theoretical analysis and experimental results on four benchmark datasets show that the joint sparsity model outperforms traditional group sparsity model and graph sparsity model, as well as the latest group-graph sparsity model. Longwen Gao, Shuigeng Zhou |
AAAI | 1 |
| 2016 | Semi-Supervised Group Sparse Representation: Model, Algorithm and ApplicationsabstractGroup sparse representation (GSR) exploits group structure in data and works well on many problems. However, the group structure must be manually given in advance. In many practical scenarios such as classification, samples are grouped according to their labels. Constructing a consistent group structure in such cases is not easy. The reasons are: 1) samples may be incorrectly labeled; and 2) label assigning in big data is time-consuming and expensive. In this paper, we propose and formulate a new problem, semi-supervised group sparse representation (SS-GSR) to support group sparse representation among both labeled and unlabeled data, while learning a more robust group structure, which can be further exploited to more effectively represent other unlabeled data. We develop a model to tackle the SS-GSR problem, based on the manifold assumption in subspace segmentation that samples in the same group lie close in feature space and span the same subspace. We also propose an alternating algorithm to solve the model. Finally, we validate the model via extensive experiments. Longwen Gao, Yeqing Li, Junzhou Huang, Shuigeng Zhou |
ECAI | 1 |
| 2015 | Learning Sparse Representations from Datasets with Uncertain Group Structures: Model, Algorithm and ApplicationsabstractGroup sparsity has drawn much attention in machine learning. However, existing work can handle only datasets with certain group structures, where each sample has a certain membership with one or more groups. This paper investigates the learning of sparse representations from datasets with uncertain group structures, where each sample has an uncertain member-ship with all groups in terms of a probability distribution. We call this problem uncertain group sparse representation (UGSR in short), which is a generalization of the standard group sparse representation (GSR). We formulate the UGSR model and propose an efficient algorithm to solve this problem. We apply UGSR to text emotion classification and aging face recognition. Experiments show that UGSR outperforms standard sparse representation (SR) and standard GSR as well as fuzzy kNN classification. Longwen Gao, Shuigeng Zhou |
AAAI | 1 |
| 2015 | Effectively classifying short texts by structured sparse representation with dictionary filtering
Longwen Gao, Shuigeng Zhou, Jihong Guan |
Inf. Sci. | 1 |
| 2014 | Towards Topological-Transformation Robust Shape Comparison: A Sparse Representation Based Manifold Embedding ApproachabstractNon-rigid shape comparison based on manifold embeddingusing Generalized Multidimensional Scaling(GMDS) has attracted much attention for its highaccuracy. However, this method requires that shape surfaceis not elastic. In other words, it is sensitive totopological transformations such as stretching and compressing.To tackle this problem, we propose a new approachthat constructs a high-dimensional space to embedthe manifolds of shapes based on sparse representation,which is able to completely withstand rigid transformationsand considerably tolerate topological transformations.Experiments on TOSCA shapes validate theproposed approach. Longwen Gao, Shuigeng Zhou |
AAAI | 1 |