EDBT 2026 Demo / reviewers in the wild / expert
Jun Peng 0007
dblp:87/6982-7
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0003-0655-1594ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMsabstractTraining-free video understanding methods leverage the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating videos as a sequences of static frames, thus obviating the need for costly video-specific training. However, this paradigm often suffers from severe visual redundancy and high computational overhead, especially when processing long videos. Crucially, existing keyframe selection strategies, especially those based on CLIP similarity, are prone to biases and may inadvertently overlook critical frames, resulting in suboptimal video comprehension. To address these significant challenges, we propose KTV, a novel two-stage framework for efficient and effective training-free video understanding. In the first stage, KTV performs question-agnostic keyframe selection by clustering frame-level visual features, yielding a compact, diverse, and representative subset of frames that mitigates temporal redundancy. In the second stage, KTV applies key visual token selection, pruning redundant or less informative tokens from each selected keyframe based on token importance and redundancy, which significantly reduces the number of tokens fed into the LLM. Extensive experiments on the Multiple-Choice VideoQA task demonstrate that KTV outperforms state-of-the-art training-free baselines while using significantly fewer visual tokens, e.g., only 504 tokens for a 60 min video with 10800 frames, achieving 44.8% accuracy on the MLVU-Test benchmark. In particular, KTV also exceeds several training-based approaches on certain benchmarks. Baiyang Song, Jun Peng 0007, Yuxin Zhang 0002, Feidiao Yang, Jianyuan Guo |
AAAI | 2 |
| 2026 | Graph-empowered Text-to-SQL generation on Electronic Medical Records
Jun Peng 0007, Baiyang Song, Yiyi Zhou, Rongrong Ji |
Pattern Recognit. | 2 |
| 2025 | TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt TuningabstractDespite the efficiency of prompt learning in transferring vision-language models (VLMs) to downstream tasks, existing methods mainly learn the prompts in a coarse-grained manner where the learned prompt vectors are shared across all categories. Consequently, the tailored prompts often fail to discern class-specific visual concepts, thereby hindering the transferred performance for classes that share similar or complex visual attributes. Recent advances mitigate this challenge by leveraging external knowledge from Large Language Models (LLMs) to furnish class descriptions, yet incurring notable inference costs. In this paper, we introduce TextRefiner, a plug-and-play method to refine the text prompts of existing methods by leveraging the internal knowledge of VLMs. Particularly, TextRefiner builds a novel local cache module to encapsulate fine-grained visual concepts derived from local tokens within the image branch. By aggregating and aligning the cached visual descriptions with the original output of the text branch, TextRefiner can efficiently refine and enrich the learned prompts from existing methods without relying on any external expertise. For example, it improves the performance of CoOp from 71.66% to 76.96% on 11 benchmarks, surpassing CoCoOp which introduced instance-wise feature for text prompts. Equipped with TextRefiner, PromptKD achieves state-of-the-art performance while keep inference efficient. Jingjing Xie, Yuxin Zhang 0002, Jun Peng 0007, Zhaohong Huang, Liujuan Cao |
AAAI | 3 |
| 2025 | From Objects to Events: Unlocking Complex Visual Understanding in Object Detectors Via LLM-guided Symbolic Reasoning
Yuhui Zeng, Haoxiang Wu, Wenjie Nie, Xiawu Zheng, Yunhang Shen, Jun Peng 0007, Yonghong Tian 0001, Rongrong Ji |
ICCV | 7 |
| 2024 | Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric RegularizationabstractText-to-3D-aware face (T3D Face) generation and manipulation is an emerging research hot spot in machine learning, which still suffers from low efficiency and poor quality. In this paper, we propose an ***E**nd-to-End **E**fficient and **E**ffective* network for fast and accurate T3D face generation and manipulation, termed $E^3$-FaceNet. Different from existing complex generation paradigms, $E^3$-FaceNet resorts to a direct mapping from text instructions to 3D-aware visual space. We introduce a novel *Style Code Enhancer* to enhance cross-modal semantic alignment, alongside an innovative *Geometric Regularization* objective to maintain consistency across multi-view generations. Extensive experiments on three benchmark datasets demonstrate that $E^3$-FaceNet can not only achieve picture-like 3D face generation and manipulation, but also improve inference speed by orders of magnitudes. For instance, compared with Latent3D, $E^3$-FaceNet speeds up the five-view generations by almost 470 times, while still exceeding in generation quality. Our code is released at <https://github.com/Aria-Zhangjl/E3-FaceNet>. Jinlu Zhang 0002, Yiyi Zhou, Qiancheng Zheng, Xiaoxiong Du, Gen Luo, Jun Peng 0007, Xiaoshuai Sun, Rongrong Ji |
ICML | 6 |
| 2024 | Efficient Event Stream Super-Resolution with Recursive Multi-Branch Fusion
Quanmin Liang, Zhilin Huang, Xiawu Zheng, Feidiao Yang, Jun Peng 0007, Kai Huang 0001, Yonghong Tian 0001 |
IJCAI | 5 |
| 2023 | PixelFace+: Towards Controllable Face Generation and Manipulation with Text Descriptions and Segmentation MasksabstractSynthesizing vivid human portraits is a research hot spot in image generation with a wide scope of applications. In addition to fidelity, generation controllability is another key factor that has long plagued its development. To address this issue, existing solutions usually adopt either textual or visual conditions for the target face synthesis, e.g., descriptions or segmentation masks, which still cannot fully control the generation due to the intrinsic shortages of each condition. In this paper, we propose to make use of both types of prior information to facilitate controllable face generation. In particular, we hope to produce coarse-grained information about faces based on the segmentation masks, such as face shapes and poses, and the text description is used to render detailed face attributes, e.g., face color, makeup and gender. More importantly, we hope that the generation can be easily controlled via interactively editing both types of information, making face generation more applicable to real-world applications. To accomplish this target, we propose a novel face generation model termed PixelFace+. In PixelFace+, both the text and mask are encoded as pixel-wise priors, based on which the pixel synthesis process is conducted to produce the expected portraits. Meanwhile, the loss objectives are also carefully designed to make sure that the generated faces are semantically aligned with both text and mask inputs. To validate the proposed PixelFace+, we conducted a comprehensive set of experiments on the widely recognized benchmark called MMCelebA. We not only quantitatively compare PixelFace+ with a bunch of newly proposed Text-to-Face(T2F) generation methods, but also give plenty of qualitative analyses. The experimental results demonstrate that PixelFace+ not only outperforms existing generation methods in both image quality and conditional matching but also shows a much superior controllability of face generation. More importantly, PixelFace+ presents a convenient and interactive way of face generation and manipulation via editing the text and mask inputs. Our SOURCE CODE and DEMO are given in our supplementary materials. Xiaoxiong Du, Jun Peng 0007, Yiyi Zhou, Jinlu Zhang 0002, Siting Chen, Guannan Jiang, Xiaoshuai Sun, Rongrong Ji |
ACM Multimedia | 2 |
| 2022 | PixelFolder: An Efficient Progressive Pixel Synthesis Network for Image Generation
Yiyi Zhou, Qi Zhang 0066, Jun Peng 0007, Yunhang Shen, Xiaoshuai Sun, Chao Chen 0026, Rongrong Ji |
ECCV (14) | 4 |
| 2022 | Learning Dynamic Prior Knowledge for Text-to-Face Pixel SynthesisabstractText-to-face (T2F) generation is an emerging research hot spot in multimedia, which aims to synthesize vivid portraits based on the given descriptions. Its main challenge lies in the accurate alignments from texts to image pixels, which pose a high demand in generation fidelity. We define T2F as a pixel synthesis problem conditioned on the texts and propose a novel dynamic pixel synthesis network, PixelFace, for end-to-end T2F generation in this paper. To fully exploit the prior knowledge for T2F synthesis, we propose a novel dynamic parameter generation module, which transforms text features into dynamic knowledge embeddings for end-to-end pixel regression. These knowledge embeddings are example-dependent and spatially related to image pixels, based on which PixelFace can exploit the text priors for high-quality text-guided face generation. To validate the proposed PixelFace, we conduct extensive experiments on the MMCelebA, and compare PixelFace with a set of state-of-the-art methods in T2F and T2I generations, e.g., StyleCLIP and TediGAN. The experimental results not only show the greater performance of PixelFace than the compared methods but also validates its merits over existing T2F methods in both text-image matching and inference speed. Codes will be released at: \textcolormagenta \urlhttps://github.com/pengjunn/PixelFace . Jun Peng 0007, Xiaoxiong Du, Yiyi Zhou, Yunhang Shen, Xiaoshuai Sun, Rongrong Ji |
ACM Multimedia | 1 |
| 2022 | Towards Open-Ended Text-to-Face Generation, Combination and ManipulationabstractText-to-face (T2F) generation is an emerging research hot spot in multimedia, and its main challenge lies in the high fidelity requirement of generated portraits. Many existing works resort to exploring the latent space in a pre-trained generator, e.g., StyleGAN, which has obvious shortcomings in efficiency and generalization ability. In this paper, we propose a generative network for open-ended text-to-face generation, which is termed OpenFaceGAN. Differing from existing StyleGAN-based methods, OpenFaceGAN constructs an effective multi-modal latent space that directly converts the natural language description into a face. This mapping paradigm can fit the real data distribution well and make the model capable of open-ended and even zero-shot T2F generation. Our method improves the inference speed by an order of magnitude, e.g., 294 times than TediGAN. Based on OpenFaceGAN, we further explore text-guided face manipulation (editing). In particular, we propose a parameterized module, OpenEditor, to automatically disentangle the target latent code and update the original style information. OpenEditor also makes OpenFaceGAN directly applicable for most manipulation instructions without example-dependent searches or optimizations, greatly improving the efficiency of face manipulation. We conduct extensive experiments on two benchmark datasets namely Multi-Modal CelebA-HQ and Face2Text-v1.0. The experimental results not only show the superior performance of OpenFaceGAN to the existing T2F methods in both image quality and image-text matching degree but also greatly confirm its outstanding ability in the zero-shot generation. Codes will be released at: \textcolormagenta \urlhttps://github.com/pengjunn/OpenFace Jun Peng 0007, Han Pan, Yiyi Zhou, Xiaoshuai Sun, Yan Wang 0059, Yongjian Wu 0001, Rongrong Ji |
ACM Multimedia | 1 |
| 2022 | Knowledge-Driven Generative Adversarial Network for Text-to-Image SynthesisabstractText-to-Image (T2I) synthesis is a challenging task that aims to convert natural language descriptions to real images. It remains an open problem mainly due to the diversity of text descriptions, which poses a huge obstacle in generating vivid and relevant images. Moreover, the existing evaluation metrics in T2I synthesis are mainly used to evaluate the visual quality of the generated images, while the semantic consistency between the two modalities is often ignored. To address these issues, we present a novelKnowledge-Driven Generative Adversarial Network, termed KD-GAN, and a new evaluation system, namedPseudo Turing Test(PTT for short). Concretely, KD-GAN takes a further step in imitating the behavior of human painting,i.e., drawing an image according to reference knowledge. The introduction of reference knowledge in KD-GAN not only improves the quality of the generated images but also enhances the semantic consistency between them and the input texts. In addition, KD-GAN can also greatly avoid some flaws against common sense during image generation,e.g., skiing in the blue sky. The proposed PTT is an important supplement to the existing evaluation system of T2I synthesis. It includes a set of pseudo-experts of different multimedia tasks to evaluate the semantic consistency between the given texts and the generated images. To validate the proposed KD-GAN, we conducted extensive experiments on two benchmark datasets,i.e., Caltech-UCSD Birds (CUB), and MS-COCO (COCO). The experimental results demonstrate that KD-GAN outperforms state-of-the-art methods on IS, FID, and the proposed PTT metrics.11The codes of KD-GAN are at [Online]. Available:https://github.com/pengjunn/KD-GANand the codes and models of PTT are at [Online]. Available:https://github.com/pengjunn/PTT. Jun Peng 0007, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Yongjian Wu 0001, Feiyue Huang, Rongrong Ji |
IEEE Trans. Multim. | 1 |
| 2019 | Towards Cross-modality Topic Modelling via Deep Topical Correlation AnalysisabstractThe cross-modality topic detection in social media retains as an open problem mainly due to the difficulty of dealing with modality independence and modality missing. In this paper, we present a novel Deep Topical Correlation Analysis (DTCA) approach, which achieves robust and accurate topic detecting for micro-blogs and handles aforementioned challenges simultaneously. In particular, bidirectional recurrent neural networks and convolutional neural networks are used to learn deep textual and visual features, respectively. Then a Canonical Correlation Analysis based fusion scheme is proposed, which has two innovations to deal with both two problems mentioned above. We further release a large-scale cross-modal twitter dataset for topic detection. Extensive and quantitative evaluations are conducted with comparisons to several state-of-the-arts and alternative approaches on this dataset. Significant performance gains are reported to demonstrate the merits of proposed approach. Jun Peng 0007, Yiyi Zhou, Liujuan Cao, Xiaoshuai Sun, Jinsong Su, Rongrong Ji |
ICASSP | 1 |