EDBT 2026 Demo / reviewers in the wild / expert
Ziyun Zeng
dblp:282/8373
· DBLP profile ↗
19ranked-venue papers
4as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Detection and Quantitative Assessment of Dental Plaque in Intraoral ImagesabstractDental plaques are biofilms formed by microorganisms and extracellular matrix in the oral cavity. The most common oral diseases, including dental caries and periodontal diseases, are initiated by dental plaque accumulating on tooth surfaces. We develop an AI-based algorithm (AI-Plaque) to automatically segment dental plaque areas in intraoral images, enabling automatic plaque detection and quantitative severity assessment. To mitigate the limited segmentation capabilities of existing models and dataset limitations, we introduce a data preprocessing procedure, a specialized fine-tuning approach, and a novel self-training pipeline that leverages synthesized plaque images. Our approach demonstrates significant improvements in plaque segmentation and high effectiveness in plaque assessment, over several baseline methods, as evaluated by both image segmentation metrics and a custom-designed quantitative Visual Plaque Index. The AI-Plaque algorithm could empower individuals to receive timely, personalized feedback on their oral hygiene practices for dental plaque control, such as brushing and flossing, and ultimately reduce the risk of developing oral diseases. Ziyun Zeng, Junyu Chen 0005, Noha Rashwan, Nisreen Al Jallad, Jiebo Luo 0001 |
ACM Trans. Comput. Heal. | 1 |
| 2025 | LiveCC: Learning Video LLM with Streaming Speech Transcription at ScaleabstractRecent video large language models (Video LLMs) often depend on costly human annotations or proprietary APIs (e.g., GPT-4o) to produce training data, which limits their training at scale. In this paper, we explore large-scale training for Video LLM with cheap automatic speech recognition (ASR) transcripts. Specifically, we propose a novel streaming training approach that densely interleaves the ASR words and video frames according to their timestamps. Compared to previous studies in vision-language representation with ASR, our method naturally fits the streaming characteristics of ASR, thus enabling the model to learn temporally-aligned, fine-grained vision-language modeling. To support the training algorithm, we introduce a data pipeline for YouTube videos and their closed captions (CC), resulting in Live-CC-10M pre-training set and Live-WhisperX-408K high-quality supervised fine-tuning (SFT) set. Remarkably, even without SFT, the pre-trained model LiveCC-7B demonstrates significant improvements in general video QA and exhibits a new capability in real-time video commentary. To evaluate this, we carefully design a new benchmark LiveSports-3K, using LLM-as-a-judge to measure the free-form commentary. Experiments show our final model LiveCC-7B can surpass LLaVA-Video-72B in commentary quality even working in a real-time mode. Meanwhile, it achieves state-of-the-art results at the 7B scale on popular benchmarks such as VideoMME, demonstrating its broad generalizability. All resources of this paper have been released at showlab.github.io/livecc. Joya Chen, Ziyun Zeng, Zejun Ma 0001, Zheng Shou 0001 |
CVPR | 2 |
| 2025 | Omnipaint: Mastering Object-Oriented Editing Via Disentangled Insertion-Removal InpaintingabstractDiffusion-based generative models have revolutionized object-oriented image editing, yet their deployment in realistic object removal and insertion remains hampered by challenges such as the intricate interplay of physical effects and insufficient paired training data. In this work, we introduce OmniPaint, a unified framework that re-conceptualizes object removal and insertion as interdependent processes rather than isolated tasks. Leveraging a pre-trained diffusion prior along with a progressive training pipeline comprising initial paired sample optimization and subsequent large-scale unpaired refinement via CycleFlow, OmniPaint achieves precise foreground elimination and seamless object insertion while faithfully preserving scene geometry and intrinsic properties. Furthermore, our novel CFD metric offers a robust, reference-free evaluation of context consistency and object hallucination, establishing a new benchmark for high-fidelity image editing. Project page: https://yeates.github.io/OmniPaint-Page/ Ziyun Zeng, Haitian Zheng, Jiebo Luo 0001 |
ICCV | 2 |
| 2025 | MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation ModelsabstractRecent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits: text-to-image (T2I) benchmarks that lacks multi-modal conditioning, and customized image generation benchmarks that overlook compositional semantics and common knowledge. We propose MMIG-Bench, a comprehensive Multi-Modal Image Generation Benchmark that unifies these tasks by pairing 4,850 richly annotated text prompts with 1,750 multi-view reference images across 380 subjects, spanning humans, animals, objects, and artistic styles. MMIG-Bench is equipped with a three-level evaluation framework: (1) low-level metrics for visual artifacts and identity preservation of objects; (2) novel Aspect Matching Score (AMS): a VQA-based mid-level metric that delivers fine-grained prompt-image alignment and shows strong correlation with human judgments; and (3) high-level metrics for aesthetics and human preference. Using MMIG-Bench, we benchmark 17 state-of-the-art models, including Gemini 2.5 Pro, FLUX, DreamBooth, and IP-Adapter, and validate our metrics with 32k human ratings, yielding in-depth insights into architecture and data design. Hang Hua, Ziyun Zeng, Yunlong Tang 0002, Daniel G. Aliaga, Wei Xiong 0008, Jiebo Luo 0001 |
NeurIPS | 2 |
| 2024 | GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video RetrievalabstractGiven a text query, partially relevant video retrieval (PRVR) seeks to find untrimmed videos containing pertinent moments in a database. For PRVR, clip modeling is essential to capture the partial relationship between texts and videos. Current PRVR methods adopt scanning-based clip construction to achieve explicit clip modeling, which is information-redundant and requires a large storage overhead. To solve the efficiency problem of PRVR methods, this paper proposes GMMFormer, a Gaussian-Mixture-Model based Transformer which models clip representations implicitly. During frame interactions, we incorporate Gaussian-Mixture-Model constraints to focus each frame on its adjacent frames instead of the whole video. Then generated representations will contain multi-scale clip information, achieving implicit clip modeling. In addition, PRVR methods ignore semantic differences between text queries relevant to the same video, leading to a sparse embedding space. We propose a query diverse loss to distinguish these text queries, making the embedding space more intensive and contain more semantic information. Extensive experiments on three large-scale video datasets (i.e., TVR, ActivityNet Captions, and Charades-STA) demonstrate the superiority and efficiency of GMMFormer. Jinpeng Wang 0002, Bin Chen 0011, Ziyun Zeng, Shutao Xia |
AAAI | 4 |
| 2024 | VideoCutLER: Surprisingly Simple Unsupervised Video Instance SegmentationabstractExisting approaches to unsupervised video instance segmentation typically rely on motion estimates and experi-ence difficulties tracking small or divergent motions. We present VideoCutLER, a simple method for unsupervised multi-instance video segmentation without using motion-based learning signals like optical flow or training on natural videos. Our key insight is that using high-quality pseudo masks and a simple video synthesis method for model training is surprisingly sufficient to enable the resulting video model to effectively segment and track multiple instances across video frames. We show the first competitive unsupervised learning results on the challenging YouTube Vis-2019 benchmark, achieving 50.7%$AP_{50}^{video}$, surpassing the previous state-of-the-art by a large margin. VideoCutLER can also serve as a strong pretrained model for supervised video instance segmentation tasks, exceeding DINO by 15.9% on YouTubeVIS-2019 in terms of$AP^{video}$. Xudong Wang 0007, Ishan Misra, Ziyun Zeng, Rohit Girdhar, Trevor Darrell |
CVPR | 3 |
| 2024 | Making LLaMA SEE and Draw with SEED TokenizerabstractThe great success of Large Language Models (LLMs) has expanded the potential of multimodality, contributing to the gradual evolution of General Artificial Intelligence (AGI). A true AGI agent should not only possess the capability to perform predefined multi-tasks but also exhibit emergent abilities in an open-world context. However, despite the considerable advancements made by recent multimodal LLMs, they still fall short in effectively unifying comprehension and generation tasks, let alone open-world emergent abilities. We contend that the key to overcoming the present impasse lies in enabling text and images to be represented and processed interchangeably within a unified autoregressive Transformer. To this end, we introduce $\textbf{SEED}$, an elaborate image tokenizer that empowers LLMs with the ability to $\textbf{SEE}$ and $\textbf{D}$raw at the same time. We identify two crucial design principles: (1) Image tokens should be independent of 2D physical patch positions and instead be produced with a $\textit{1D causal dependency}$, exhibiting intrinsic interdependence that aligns with the left-to-right autoregressive prediction mechanism in LLMs. (2) Image tokens should capture $\textit{high-level semantics}$ consistent with the degree of semantic abstraction in words, and be optimized for both discriminativeness and reconstruction during the tokenizer training phase. With SEED tokens, LLM is able to perform scalable multimodal autoregression under its original training recipe, i.e., next-word prediction. SEED-LLaMA is therefore produced by large-scale pretraining and instruction tuning on the interleaved textual and visual data, demonstrating impressive performance on a broad range of multimodal comprehension and generation tasks. More importantly, SEED-LLaMA has exhibited compositional emergent abilities such as multi-turn in-context multimodal generation, acting like your AI assistant. The code (training and inference) and models are released in https://github.com/AILab-CVC/SEED. Yuying Ge, Sijie Zhao, Ziyun Zeng, Yixiao Ge, Chen Li 0046, Xintao Wang 0002, Ying Shan |
ICLR | 3 |
| 2024 | PromptFix: You Prompt and We Fix the PhotoabstractDiffusion models equipped with language models demonstrate excellent controllability in image generation tasks, allowing image processing to adhere to human instructions. However, the lack of diverse instruction-following data hampers the development of models that effectively recognize and execute user-customized instructions, particularly in low-level tasks. Moreover, the stochastic nature of the diffusion process leads to deficiencies in image generation or editing tasks that require the detailed preservation of the generated images. To address these limitations, we propose PromptFix, a comprehensive framework that enables diffusion models to follow human instructions to perform a wide variety of image-processing tasks. First, we construct a large-scale instruction-following dataset that covers comprehensive image-processing tasks, including low-level tasks, image editing, and object creation. Next, we propose a high-frequency guidance sampling method to explicitly control the denoising process and preserve high-frequency details in unprocessed areas. Finally, we design an auxiliary prompting adapter, utilizing Vision-Language Models (VLMs) to enhance text prompts and improve the model's task generalization. Experimental results show that PromptFix outperforms previous methods in various image-processing tasks. Our proposed model also achieves comparable inference efficiency with these baseline models and exhibits superior zero-shot capabilities in blind restoration and combination tasks. Ziyun Zeng, Hang Hua, Jianlong Fu, Jiebo Luo 0001 |
NeurIPS | 2 |
| 2024 | Hugs Bring Double Benefits: Unsupervised Cross-Modal Hashing with Multi-granularity Aligned Transformers
Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Dongliang Liao, Gongfu Li, Shutao Xia |
Int. J. Comput. Vis. | 2 |
| 2024 | Pyramid hybrid pooling quantization for efficient fine-grained image retrieval
Ziyun Zeng, Jinpeng Wang 0002, Bin Chen 0011, Tao Dai 0001, Shutao Xia, Zhi Wang 0001 |
Pattern Recognit. Lett. | 1 |
| 2023 | Contrastive Masked Autoencoders for Self-Supervised Video HashingabstractSelf-Supervised Video Hashing (SSVH) models learn to generate short binary representations for videos without ground-truth supervision, facilitating large-scale video retrieval efficiency and attracting increasing research attention. The success of SSVH lies in the understanding of video content and the ability to capture the semantic relation among unlabeled videos. Typically, state-of-the-art SSVH methods consider these two points in a two-stage training pipeline, where they firstly train an auxiliary network by instance-wise mask-and-predict tasks and secondly train a hashing model to preserve the pseudo-neighborhood structure transferred from the auxiliary network. This consecutive training strategy is inflexible and also unnecessary. In this paper, we propose a simple yet effective one-stage SSVH method called ConMH, which incorporates video semantic information and video similarity relationship understanding in a single stage. To capture video semantic information for better hashing learning, we adopt an encoder-decoder structure to reconstruct the video from its temporal-masked frames. Particularly, we find that a higher masking ratio helps video understanding. Besides, we fully exploit the similarity relationship between videos by maximizing agreement between two augmented views of a video, which contributes to more discriminative and robust hash codes. Extensive experiments on three large-scale video datasets (i.e., FCVID, ActivityNet and YFCC) indicate that ConMH achieves state-of-the-art results. Code is available at https://github.com/huangmozhi9527/ConMH. Jinpeng Wang 0002, Bin Chen 0011, Ziyun Zeng, Shutao Xia |
AAAI | 4 |
| 2023 | Learning Transferable Spatiotemporal Representations from Natural Script KnowledgeabstractPre-training on large-scale video data has become a common recipe for learning transferable spatiotemporal representations in recent years. Despite some progress, existing methods are mostly limited to highly curated datasets (e.g., K400) and exhibit unsatisfactory out-of-the-box representations. We argue that it is due to the fact that they only capture pixel-level knowledge rather than spatiotemporal semantics, which hinders further progress in video understanding. Inspired by the great success of image-text pre-training (e.g., CLIP), we take the first step to exploit language semantics to boost transferable spatiotemporal representation learning. We introduce a new pre-text task, Turning to Video for Transcript Sorting (TVTS), which sorts shuffled ASR scripts by attending to learned video representations. We do not rely on descriptive captions and learn purely from video, i.e., leveraging the natural transcribed speech knowledge to provide noisy but useful semantics over time. Our method enforces the vision model to contextualize what is happening over time so that it can re-organize the narrative transcripts, and can seamlessly apply to large-scale uncurated video data in the real world. Our method demonstrates strong out-of-the-box spatiotemporal representations on diverse benchmarks, e.g., +13.6% gains over VideoMAE on SSV2 via linear probing. The code is available at https://github.com/TencentARC/TVTS. Ziyun Zeng, Yuying Ge, Xihui Liu, Bin Chen 0011, Ping Luo 0002, Shutao Xia, Yixiao Ge |
CVPR | 1 |
| 2023 | ConCAP: Contrastive Context-Aware Prompt for Resource-hungry Action RecognitionabstractExisting large-scale image-language pre-trained models, e.g., CLIP [1], have revealed strong spatial recognition capability on various vision tasks. However, they achieve inferior performance in action recognition due to lack of temporal reasoning ability. Moreover, fully tuning large models require expensive computational infrastructures, and state-of-the-art video models yield slow inference speed due to the high frame sampling rate. The above drawbacks make existing video action recognition works impractical to be applied in resource-hungry scenarios, which is common in the real world. In this work, we propose Contrastive Context-Aware Prompt (ConCAP) for resource-hungry action recognition. Specifically, we develop a lightweight PromptFormer to learn the spatio-temporal representations stacking on top of frozen frame-wise visual backbones, where learnable prompt tokens are plugged between frame tokens during self-attention. These prompt tokens are expected to auto-complete the contextual spatiotemporal information between frames and therefore enhance the model’s representation capability. To achieve this goal, we align the prompt-enhanced representation with both category-level textual representations and video representations from densely sampled frames. Extensive experiments on four video benchmarks show that we achieve state-of-the-art or competitive performance compared to existing methods with far fewer trainable parameters and faster inference speed with limited frames, demonstrating the superiority of ConCAP in resource-hungry scenarios. Ziyun Zeng, Qijun Zhao, Zhen Zhai |
ICME | 2 |
| 2023 | MISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for RecommendationabstractThe goal of sequential recommendation (SR) is to predict a user's potential interested items based on her/his historical interaction sequences. Most existing sequential recommenders are developed based on ID features, which, despite their widespread use, often underperform with sparse IDs and struggle with the cold-start problem. Besides, inconsistent ID mappings hinder the model's transferability, isolating similar recommendation domains that could have been co-optimized. This paper aims to address these issues by exploring the potential of multi-modal information in learning robust and generalizable sequence representations. We propose MISSRec, a multi-modal pre-training and transfer learning framework for SR. On the user side, we design a Transformer-based encoder-decoder model, where the contextual encoder learns to capture the sequence-level multi-modal synergy while a novel interest-aware decoder is developed to grasp item-modality-interest relations for better sequence representation. On the candidate item side, we adopt a dynamic fusion module to produce user-adaptive item representation, providing more precise matching between users and items. We pre-train the model with contrastive learning objectives and fine-tune it in an efficient manner. Extensive experiments demonstrate the effectiveness and flexibility of MISSRec, promising an practical solution for real-world recommendation scenarios. Jinpeng Wang 0002, Ziyun Zeng, Jun Yuan 0008, Rui Zhang 0003, Hai-Tao Zheng 0002, Shutao Xia |
ACM Multimedia | 2 |
| 2022 | Contrastive Quantization with Code Memory for Unsupervised Image RetrievalabstractThe high efficiency in computation and storage makes hashing (including binary hashing and quantization) a common strategy in large-scale retrieval systems. To alleviate the reliance on expensive annotations, unsupervised deep hashing becomes an important research problem. This paper provides a novel solution to unsupervised deep quantization, namely Contrastive Quantization with Code Memory (MeCoQ). Different from existing reconstruction-based strategies, we learn unsupervised binary descriptors by contrastive learning, which can better capture discriminative visual semantics. Besides, we uncover that codeword diversity regularization is critical to prevent contrastive learning-based quantization from model degeneration. Moreover, we introduce a novel quantization code memory module that boosts contrastive learning with lower feature drift than conventional feature memories. Extensive experiments on benchmark datasets show that MeCoQ outperforms state-of-the-art methods. Code and configurations are publicly released. Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Tao Dai 0001, Shutao Xia |
AAAI | 2 |
| 2022 | Hugs Are Better Than Handshakes: Unsupervised Cross-Modal Transformer Hashing with Multi-granularity Alignment
Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Dongliang Liao, Gongfu Li, Shutao Xia |
BMVC | 2 |
| 2022 | Motion-Aware Graph Reasoning Hashing for Self-supervised Video Retrieval
Ziyun Zeng, Jinpeng Wang 0002, Bin Chen 0011, Shutao Xia |
BMVC | 1 |
| 2022 | Hybrid Contrastive Quantization for Efficient Cross-View Video RetrievalabstractWith the recent boom of video-based social platforms (e.g., YouTube and TikTok), video retrieval using sentence queries has become an important demand and attracts increasing research attention. Despite the decent performance, existing text-video retrieval models in vision and language communities are impractical for large-scale Web search because they adopt brute-force search based on high-dimensional embeddings. To improve efficiency, Web search engines widely apply vector compression libraries (e.g., FAISS [26]) to post-process the learned embeddings. Unfortunately, separate compression from feature encoding degrades the robustness of representations and incurs performance decay. To pursue a better balance between performance and efficiency, we propose the first quantized representation learning method for cross-view video retrieval, namely Hybrid Contrastive Quantization (HCQ). Specifically, HCQ learns both coarse-grained and fine-grained quantizations with transformers, which provide complementary understandings for texts and videos and preserve comprehensive semantic information. By performing Asymmetric-Quantized Contrastive Learning (AQ-CL) across views, HCQ aligns texts and videos at coarse-grained and multiple fine-grained levels. This hybrid-grained learning strategy serves as strong supervision on the cross-view video quantization model, where contrastive learning at different levels can be mutually promoted. Extensive experiments on three Web video benchmark datasets demonstrate that HCQ achieves competitive performance with state-of-the-art non-compressed retrieval methods while showing high efficiency in storage and computation. Code and configurations are available at https://github.com/gimpong/WWW22-HCQ. Jinpeng Wang 0002, Bin Chen 0011, Dongliang Liao, Ziyun Zeng, Gongfu Li, Shutao Xia, Jin Xu 0014 |
WWW | 4 |
| 2021 | SwinFGHash: Fine-grained Image Retrieval via Transformer-based Hashing Network
Jinpeng Wang 0002, Ziyun Zeng, Bin Chen 0011, Shudeng Wu, Shutao Xia |
BMVC | 3 |