EDBT 2026 Demo / reviewers in the wild / expert
Quan Chen 0006
dblp:40/3858-6
· DBLP profile ↗
15ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0002-4865-2396ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-free subject-enhanced attention guidance for compositional text-to-image generation
Bo Wang 0011, Te Yang, Quan Chen 0006, Di Dong |
Pattern Recognit. | 5 |
| 2026 | ASR-Enhanced Multimodal Representation Learning for Cross-Domain Product RetrievalabstractE-commerce is increasinglymultimedia-enriched, with products exhibited in a broad-domain manner as images, short videos, or live stream promotions. A unified and vectorized cross-domain production representation is essential. Due to large intra-product variance and high inter-product similarity in the broad-domain scenario, a visual-only representation is inadequate. While Automatic Speech Recognition (ASR) text derived from the short or live-stream videos is readily accessible, how to de-noise the excessively noisy text for multimodal representation learning is mostly untouched. We proposeASR-enhancedMultimodalProduct Representation Learning (AMPere). In order to extract product-specific information from the raw ASR text,AMPereuses an easy-to-implement LLM-based ASR text summarizer. The LLM-summarized text, together with visual data, is then fed into a multi-branch network to generate compact multimodal embeddings. Extensive experiments on a large-scale tri-domain dataset verify the effectiveness ofAMPerein obtaining a unified multimodal product representation that clearly improves cross-domain product retrieval. Ruixiang Zhao, Jian Jia, Yan Li 0043, Xuehan Bai, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Xirong Li 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Industrial ApplicationabstractContemporary recommendation systems predominantly rely on ID embedding to capture latent associations among users and items. However, this approach overlooks the wealth of semantic information embedded within textual descriptions of items, leading to suboptimal performance and poor generalizations. Leveraging the capability of large language models to comprehend and reason about textual content presents a promising avenue for advancing recommendation systems. To achieve this, we propose an Llm-driven knowlEdge Adaptive RecommeNdation (LEARN) framework that synergizes open-world knowledge with collaborative knowledge. We address computational complexity concerns by utilizing pretrained LLMs as item encoders and freezing LLM parameters to avoid catastrophic forgetting and preserve open-world knowledge. To bridge the gap between the open-world and collaborative domains, we design a twin-tower structure supervised by the recommendation task and tailored for practical industrial application. Through experiments on the real large-scale industrial dataset and online A/B tests, we demonstrate the efficacy of our approach in industry application. We also achieve state-of-the-art performance on six Amazon Review datasets to verify the superiority of our method. Jian Jia, Yan Li 0043, Honggang Chen, Xuehan Bai, Zhaocheng Liu, Jian Liang 0001, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Kun Gai |
AAAI | 8 |
| 2025 | D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX MatchingabstractVideos showcasing specific products are increasingly important for E-commerce. Key moments naturally exist as the first appearance of a specific product, presentation of its distinctive features, the presence of a buying link, etc. Adding proper sound effects (SFX) to such moments, or video decoration with SFX (VDSFX), is crucial for enhancing user engaging experience. Previous work adds SFX to videos by video-to-SFX matching at a holistic level, lacking the ability of adding SFX to a specific moment. Meanwhile, previous studies on video highlight detection or video moment retrieval consider only moment localization, leaving moment to SFX matching untouched. By contrast, we propose in this paper D&M, a unified method that accomplishes key moment detection and moment-to-SFX matching simultaneously. Moreover, for the new VDSFX task we build a large-scale dataset SFX-Moment from an E-commerce video creation platform. For a fair comparison, we build competitive baselines by extending a number of current video moment detection methods to the new task. Extensive experiments on SFX-Moment show the superior performance of the proposed method over the baselines. Minquan Wang, Bo Wang 0011, Aozhu Chen, Quan Chen 0006, Peng Jiang 0002, Xirong Li 0001 |
AAAI | 6 |
| 2025 | SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
Zhentao Tan, Ben Xue, Jian Jia, Wencai Ye, Shaoyun Shi, Mingjie Sun, Wenjin Wu, Quan Chen 0006, Peng Jiang 0002 |
ICCV | 9 |
| 2025 | Music Grounding by Short Video
Zijie Xin, Minquan Wang, Quan Chen 0006, Peng Jiang 0002, Xirong Li 0001 |
ICCV | 4 |
| 2025 | MOTION: Multi-object Video Editing with Training-Free Attention Guidance
Qitong Yan, Jian Jia, Bo Wang 0071, Quan Chen 0006, Peng Jiang 0002, Minfeng Zhu 0001, Linchao Zhu, Wei Chen 0001 |
ICIC (18) | 6 |
| 2025 | Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific HeadsabstractWe introduce Orthus, a unified multimodal model that excels in generating interleaved images and text from mixed-modality inputs by simultaneously handling discrete text tokens and continuous image features under the AR modeling principle. The continuous treatment of visual signals minimizes the information loss while the fully AR formulation renders the characterization of the correlation between modalities straightforward. Orthus leverages these advantages through its modality-specific heads—one regular language modeling (LM) head predicts discrete text tokens and one diffusion head generates continuous image features. We devise an efficient strategy for building Orthus—by substituting the Vector Quantization (VQ) operation in the existing unified AR model with a soft alternative, introducing a diffusion head, and tuning the added modules to reconstruct images, we can create an Orthus-base model effortlessly (e.g., within 72 A100 GPU hours). Orthus-base can further embrace post-training to craft lengthy interleaved image-text, reflecting the potential for handling intricate real-world tasks. For visual understanding and generation, Orthus achieves a GenEval score of 0.58 and an MME-P score of 1265.8 using 7B parameters, outperforming competing baselines including Show-o and Chameleon. Siqi Kou, Jiachun Jin, Jian Jia, Quan Chen 0006, Peng Jiang 0002, Zhijie Deng |
ICML | 7 |
| 2024 | Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation LearningabstractIn recent years, text-to-video retrieval methods based on CLIP have experienced rapid development. The primary direction of evolution is to exploit the much wider gamut of visual and textual cues to achieve alignment. Concretely, those methods with impressive performance often design a heavy fusion block for sentence (words)-video (frames) interaction, regardless of the prohibitive computation complexity. Nevertheless, these approaches are not optimal in terms of feature utilization and retrieval efficiency. To address this issue, we adopt multi-granularity visual feature learning, ensuring the model's comprehensiveness in capturing visual content features spanning from abstract to detailed levels during the training phase. To better leverage the multi-granularity features, we devise a two-stage retrieval architecture in the retrieval phase. This solution ingeniously balances the coarse and fine granularity of retrieval content. Moreover, it also strikes a harmonious equilibrium between retrieval effectiveness and efficiency. Specifically, in training phase, we design a parameter-free text-gated interaction block (TIB) for fine-grained video representation learning and embed an extra Pearson Constraint to optimize cross-modal representation learning. In retrieval phase, we use coarse-grained video representations for fast recall of top-k candidates, which are then reranked by fine-grained video representations. Extensive experiments on four benchmarks demonstrate the efficiency and effectiveness. Notably, our method achieves comparable performance with the current state-of-the-art methods while being nearly 50 times faster. Kaibin Tian, Yanhua Cheng, Xinglin Hou, Quan Chen 0006, Han Li 0005 |
AAAI | 5 |
| 2024 | Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product Retrieval
Xiaowan Hu, Yan Li 0043, Minquan Wang, Haoqian Wang, Quan Chen 0006, Han Li 0005, Peng Jiang 0002 |
ACM Multimedia | 6 |
| 2024 | Spatiotemporal Fine-grained Video Description for Short VideosabstractIn the mobile internet era, short videos are inundating people's lives. However, research on visual language models specifically designed for short videos has not yet received sufficient attention. Short videos are not just videos of limited duration. The prominent visual details and high information density of short videos differentiate them to long videos. In this paper, we propose the SpatioTemporal Fine-grained Description (STFVD) emphasizing on the uniqueness of short videos, which entails capturing the intricate details of the main subject and fine-grained movements. To this end, we create a comprehensive Short Video Advertisements Description (SVAD) dataset, comprising 34,930 clips from 5,046 videos. The dataset covers a range of topics, including 191 sub-industries, 649 popular products, and 470 trending games. Various efforts have been made in the data annotation process to ensure the inclusion of fine-grained spatiotemporal information, resulting in 34,930 high-quality annotations. Compared to existing datasets, samples in SVAD exhibit a superior text information density, suggesting that SVAD is more appropriate for the analysis of short videos. Based on the SVAD dataset, we develop a visual language model (SVAD-VLM) to generate spatiotemporal fine-grained description for short videos. We use a prompt-guided keyword generation task to efficiently learn key visual information. Moreover, we also utilize dual visual alignment to exploit the advantage of mixed-datasets training. Experiments on SVAD dataset demonstrate the challenge of STFVD and the competitive performance of proposed method compared to previous ones. Te Yang, Jian Jia, Bo Wang 0071, Yanhua Cheng, Yan Li 0043, Dongze Hao, Xipeng Cao, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Xiangyu Zhu 0001, Zhen Lei 0001 |
ACM Multimedia | 8 |
| 2023 | Cross-Domain Product Representation Learning for Rich-Content E-CommerceabstractThe proliferation of short video and live-streaming platforms has revolutionized how consumers engage in online shopping. Instead of browsing product pages, consumers are now turning to rich-content e-commerce, where they can purchase products through dynamic and interactive media like short videos and live streams. This emerging form of online shopping has introduced technical challenges, as products may be presented differently across various media domains. Therefore, a unified product representation is essential for achieving cross-domain product recognition to ensure an optimal user search experience and effective product recommendations. Despite the urgent industrial need for a unified cross-domain product representation, previous studies have predominantly focused only on product pages without taking into account short videos and live streams. To fill the gap in the rich-content e-commerce area, in this paper, we introduce a large-scale cRoss-dOmain Product rEcognition dataset, called ROPE. ROPE covers a wide range of product categories and contains over 180,000 products, corresponding to millions of short videos and live streams. It is the first dataset to cover product pages, short videos, and live streams simultaneously, providing the basis for establishing a unified product representation across different media domains. Furthermore, we propose a Cross-dOmain Product rEpresentation framework, namely COPE, which unifies product representations in different domains through multimodal learning including text and vision. Extensive experiments on downstream tasks demonstrate the effectiveness of COPE in learning a joint feature space for all product domains. Xuehan Bai, Yan Li 0043, Yanhua Cheng, Quan Chen 0006, Han Li 0005 |
ICCV | 5 |
| 2023 | Cross-view Semantic Alignment for Livestreaming Product RecognitionabstractLive commerce is the act of selling products online through live streaming. The customer’s diverse demands for online products introduce more challenges to Livestreaming Product Recognition. Previous works have primarily focused on fashion clothing data or utilize single-modal input, which does not reflect the real-world scenario where multimodal data from various categories are present. In this paper, we present LPR4M, a large-scale multimodal dataset that covers 34 categories, comprises 3 modalities (image, video, and text), and is 50× larger than the largest publicly available dataset. LPR4M contains diverse videos and noise modality pairs while exhibiting a long-tailed distribution, resembling real-world problems. Moreover, a cRoss-vIew semantiC alignmEnt (RICE) model is proposed to learn discriminative instance features from the image and video views of the products. This is achieved through instance-level contrastive learning and cross-view patch-level feature propagation. A novel Patch Feature Reconstruction loss is proposed to penalize the semantic misalignment between cross-view patches. Extensive experiments demonstrate the effectiveness of RICE and provide insights into the importance of dataset diversity and expressivity. The dataset and code are available at https://github.com/adxcreative/RICE. Yan Li 0043, Yanhua Cheng, Quan Chen 0006, Han Li 0005 |
ICCV | 6 |
| 2020 | Progressive Feature Polishing Network for Salient Object DetectionabstractFeature matters for salient object detection. Existing methods mainly focus on designing a sophisticated structure to incorporate multi-level features and filter out cluttered features. We present Progressive Feature Polishing Network (PFPN), a simple yet effective framework to progressively polish the multi-level features to be more accurate and representative. By employing multiple Feature Polishing Modules (FPMs) in a recurrent manner, our approach is able to detect salient objects with fine details without any post-processing. A FPM parallelly updates the features of each level by directly incorporating all higher level context information. Moreover, it can keep the dimensions and hierarchical structures of the feature maps, which makes it flexible to be integrated with any CNN-based models. Empirical experiments show that our results are monotonically getting better with increasing number of FPMs. Without bells and whistles, PFPN outperforms the state-of-the-art methods significantly on five benchmark datasets under various evaluation metrics. Our code is available at: https://github.com/chenquan-cq/PFPN. Bo Wang 0071, Quan Chen 0006, Zhiqiang Zhang 0011, Xiaogang Jin 0001, Kun Gai |
AAAI | 2 |
| 2018 | Semantic Human MattingabstractHuman matting, high quality extraction of humans from natural images, is crucial for a wide variety of applications. Since the matting problem is severely under-constrained, most previous methods require user interactions to take user designated trimaps or scribbles as constraints. This user-in-the-loop nature makes them difficult to be applied to large scale data or time-sensitive scenarios. In this paper, instead of using explicit user input constraints, we employ implicit semantic constraints learned from data and propose an automatic human matting algorithm Semantic Human Matting(SHM). SHM is the first algorithm that learns to jointly fit both semantic information and high quality details with deep networks. In practice, simultaneously learning both coarse semantics and fine details is challenging. We propose a novel fusion strategy which naturally gives a probabilistic estimation of the alpha matte. We also construct a very large dataset with high quality annotations consisting of 35,513 unique foregrounds to facilitate the learning and evaluation of human matting. Extensive experiments on this dataset and plenty of real images show that SHM achieves comparable results with state-of-the-art interactive matting methods. Quan Chen 0006, Tiezheng Ge, Yanyu Xu 0001, Zhiqiang Zhang 0011, Kun Gai |
ACM Multimedia | 1 |