EDBT 2026 Demo / reviewers in the wild / expert
Jun-Kun Chen
dblp:333/0859
· DBLP profile ↗
22ranked-venue papers
8as first author
19since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HDEdit: Editing Videos and 3D Scenes with Video Diffusion Through Hierarchical Task DecompositionabstractWe introduce HDEdit, a training-free framework for instruction-guided video and 3D scene editing that resolves the fundamental tension between instruction fulfillment and original content preservation through Hierarchical task Decomposition. Our key insight is to progressively decompose complex edits into simpler subtasks. This hierarchical strategy aligns with dual objectives: an LLMguided planner structures high-level subgoals for reliable instruction fulfillment, while embedding-space interpolation further refines each subgoal to preserve unedited content. Two tailored control mechanisms - word-level attention map propagation and parallel denoising synchronization - ensure temporally consistent, hyperparameter tuning-free execution. Beyond video, we extend HDEdit to 3D editing via a simple yet effective render-edit-reconstruct process that maintains strong geometric consistency. Extensive experiments demonstrate our state-of-the-art results across diverse and challenging edits, including longduration videos, fast camera motion, and significant 3D geometric changes. Jun-Kun Chen, Jipeng Lyu, Yu-Xiong Wang |
3DV | 2 |
| 2024 | ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene EditingabstractThis paper proposes ConsistDreamer - a novel framework that lifts 2D diffusion models with 3D awareness and 3D consistency, thus enabling high-fidelity instruction-guided scene editing. To overcome the fundamental limitation of missing 3D consistency in 2D diffusion models, our key insight is to introduce three synergistic strategies that augment the input of the 2D diffusion model to become 3D-aware and to explicitly enforce 3D consistency during the training process. Specifically, we design sur-rounding views as context-rich input for the 2D diffusion model, and generate 3D-consistent structured noise instead of image-independent noise. Moreover, we introduce self-supervised consistency-enforcing training within the per-scene editing procedure. Extensive evaluation shows that our ConsistDreamer achieves state-of-the-art performance for instruction-guided scene editing across various scenes and editing instructions, particularly in complicated large-scale indoor scenes from ScanNet++, with significantly improved sharpness and fine-grained textures. Notably, ConsistDreamer stands as the first work capable of success-fully editing complex (e.g., plaid/checkered) patterns. Our project page is at immortalco.github.io/ConsistDreamer. Jun-Kun Chen, Samuel Rota Bulò, Norman Müller, Lorenzo Porzi, Peter Kontschieder, Yu-Xiong Wang |
CVPR | 1 |
| 2024 | Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D DiffusionabstractThis paper proposes Instruct 4D-to-4D that achieves 4D awareness and spatial-temporal consistency for 2D diffusion models to generate high-quality instruction-guided dynamic scene editing results. Traditional applications of 2D diffusion models in dynamic scene editing often result in inconsistency, primarily due to their inherent frame-by-frame editing methodology. Addressing the complexities of extending instruction-guided editing to 4D, our key insight is to treat a 4D scene as a pseudo-3D scene, decoupled into two sub-problems: achieving temporal consistency in video editing and applying these edits to the pseudo-3D scene. Following this, we first enhance the Instruct-Pix2Pix (IP2P) model with an anchor-aware attention module for batch processing and consistent editing. Additionally, we integrate optical flow-guided appearance propagation in a sliding window fashion for more precise frame-to-frame editing and incorporate depth-based projection to manage the extensive data of pseudo-3D scenes, followed by iterative editing to achieve convergence. We extensively evaluate our approach in various scenes and editing instructions, and demonstrate that it achieves spatially and temporally consistent editing results, with significantly enhanced detail and sharpness over the prior art. Notably, Instruct 4D-to-4D is general and applicable to both monocular and challenging multi-camera scenes. Code and more results are available at immortalco.github.io/Instruct-4D-to-4D. Linzhan Mou, Jun-Kun Chen, Yu-Xiong Wang |
CVPR | 2 |
| 2024 | Leveraging Timestamp Information for Serialized Joint Streaming Recognition and TranslationabstractThe growing need for instant spoken language transcription and translation is driven by increased global communication and cross-lingual interactions. This has made offering translations in multiple languages essential for user applications. Traditional approaches to automatic speech recognition (ASR) and speech translation (ST) have often relied on separate systems, leading to inefficiencies in computational resources, and increased synchronization complexity in real time. In this paper, we propose a streaming Transformer-Transducer (T-T) model able to jointly produce many-to-one and one-to-many transcription and translation using a single decoder. We introduce a novel method for joint token-level serialized output training based on timestamp information to effectively produce ASR and ST outputs in the streaming setting. Experiments on {it,es,de}↔en prove the effectiveness of our approach, enabling the generation of one-to-many joint outputs with a single decoder for the first time. Sara Papi, Jun-Kun Chen, Naoyuki Kanda, Jinyu Li 0001, Yashesh Gaur |
ICASSP | 3 |
| 2024 | Diarist: Streaming Speech Translation with Speaker DiarizationabstractEnd-to-end speech translation (ST) for conversation recordings involves several under-explored challenges such as speaker diarization (SD) without accurate word time stamps and handling of overlapping speech in a streaming fashion. In this work, we propose DiariST, the first streaming ST and SD solution. It is built upon a neural transducer-based streaming ST system and integrates tokenlevel serialized output training and t-vector, which were originally developed for multi-talker speech recognition. Due to the absence of evaluation benchmarks in this area, we develop a new evaluation dataset, DiariST-AliMeeting, by translating the reference Chinese transcriptions of the AliMeeting corpus into English. We also propose new metrics, called speaker-agnostic BLEU and speaker-attributed BLEU, to measure the ST quality while taking SD accuracy into account. Our system achieves a strong ST and SD capability compared to offline systems based on Whisper, while performing streaming inference for overlapping speech. To facilitate the research in this new direction, we release the evaluation data, the offline baseline systems, and the evaluation code. Mu Yang, Naoyuki Kanda, Xiaofei Wang 0009, Jun-Kun Chen, Jinyu Li 0001, Takuya Yoshioka |
ICASSP | 4 |
| 2024 | Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
Jinyu Li 0001, Jun-Kun Chen, Aswin Shanmugam Subramanian |
INTERSPEECH | 4 |
| 2024 | ProEdit: Simple Progression is All You Need for High-Quality 3D Scene EditingabstractThis paper proposes ProEdit - a simple yet effective framework for high-quality 3D scene editing guided by diffusion distillation in a novel progressive manner. Inspired by the crucial observation that multi-view inconsistency in scene editing is rooted in the diffusion model’s large feasible output space (FOS), our framework controls the size of FOS and reduces inconsistency by decomposing the overall editing task into several subtasks, which are then executed progressively on the scene. Within this framework, we design a difficulty-aware subtask decomposition scheduler and an adaptive 3D Gaussian splatting (3DGS) training strategy, ensuring high efficiency in performing each subtask. Extensive evaluation shows that our ProEdit achieves state-of-the-art results in various scenes and challenging editing tasks, all through a simple framework without any expensive or sophisticated add-ons like distillation losses, components, or training procedures. Notably, ProEdit also provides a new way to preview, control, and select the aggressivity of editing operation during the editing process. Jun-Kun Chen, Yu-Xiong Wang |
NeurIPS | 1 |
| 2024 | SceneCraft: Layout-Guided 3D Scene GenerationabstractThe creation of complex 3D scenes tailored to user specifications has been a tedious and challenging task with traditional 3D modeling tools. Although some pioneering methods have achieved automatic text-to-3D generation, they are generally limited to small-scale scenes with restricted control over the shape and texture. We introduce SceneCraft, a novel method for generating detailed indoor scenes that adhere to textual descriptions and spatial layout preferences provided by users. Central to our method is a rendering-based technique, which converts 3D semantic layouts into multi-view 2D proxy maps. Furthermore, we design a semantic and depth conditioned diffusion model to generate multi-view images, which are used to learn a neural radiance field (NeRF) as the final scene representation. Without the constraints of panorama image generation, we surpass previous methods in supporting complicated indoor space generation beyond a single room, even as complicated as a whole multi-bedroom apartment with irregular shapes and layouts. Through experimental analysis, we demonstrate that our method significantly outperforms existing approaches in complex indoor scene generation with diverse textures, consistent geometry, and realistic visual quality. Xiuyu Yang, Yunze Man, Jun-Kun Chen, Yu-Xiong Wang |
NeurIPS | 3 |
| 2024 | Investigating Neural Audio Codecs For Speech Language Model-Based Speech GenerationabstractNeural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the codec system affects the speech generation performance of the SLM. In this work, we examine codec tokens within SLM framework for speech generation to provide insights for effective codec design. We retrain existing high-performing neural codec models on the same data set and loss functions to compare their performance in a uniform setting. We integrate codec tokens into two SLM systems: masked-based parallel speech generation system and an auto-regressive (AR) plus non-auto-regressive (NAR) model-based system. Our findings indicate that better speech reconstruction in codec systems does not guarantee improved speech generation in SLM. A high-quality codec decoder is crucial for natural speech production in SLM, while speech intelligibility depends more on quantization mechanism. Jiaqi Li 0030, Dongmei Wang, Xiaofei Wang 0009, Yao Qian, Shujie Liu 0001, Midia Yousefi, Canrun Li, Chung-Hsien Tsai, Jun-Kun Chen, Sheng Zhao 0002, Jinyu Li 0001, Zhizheng Wu 0001, Michael Zeng 0001 |
SLT | 12 |
| 2023 | Improving Stability in Simultaneous Speech Translation: A Revision-Controllable Decoding ApproachabstractSimultaneous Speech-to-Text translation serves a critical role in real-time crosslingual communication. Despite the advancements in recent years, challenges remain in achieving stability in the translation process, a concern primarily manifested in the flickering of partial results. In this paper, we propose a novel revision-controllable method designed to address this issue. Our method introduces an allowed revision window within the beam search pruning process to screen out candidate translations likely to cause extensive revisions, leading to a substantial reduction in flickering and, crucially, providing the capability to completely eliminate flickering. The experiments demonstrate the proposed method can significantly improve the decoding stability without compromising substantially on the translation quality. Jun-Kun Chen, Jinyu Li 0001 |
ASRU | 1 |
| 2023 | Token-Level Serialized Output Training for Joint Streaming ASR and ST Leveraging Textual AlignmentsabstractIn real-world applications, users often require both translations and transcriptions of speech to enhance their comprehension, particularly in streaming scenarios where incremental generation is necessary. This paper introduces a streaming Transformer-Transducer that jointly generates automatic speech recognition (ASR) and speech translation (ST) outputs using a single decoder. To produce ASR and ST content effectively with minimal latency, we propose a joint token-level serialized output training method that interleaves source and target words by leveraging an off-the-shelf textual aligner. Experiments in monolingual (it-en) and multilingual ({de,es,it}-en) settings demonstrate that our approach achieves the best quality-latency balance. With an average ASR latency of 1s and ST latency of $1.3 \mathrm{~s}$, our model shows no degradation or even improves output quality compared to separate ASR and ST models, yielding an average improvement of 1.1 WER and 0.4 BLEU in the multilingual case. Sara Papi, Jun-Kun Chen, Jinyu Li 0001, Yashesh Gaur |
ASRU | 3 |
| 2023 | NeuralEditor: Editing Neural Radiance Fields via Manipulating Point CloudsabstractThis paper proposes NeuralEditor that enables neural radiance fields (NeRFs) natively editable for general shape editing tasks. Despite their impressive results on novel-view synthesis, it remains a fundamental challenge for NeRFs to edit the shape of the scene. Our key insight is to exploit the explicit point cloud representation as the underlying structure to construct NeRFs, inspired by the intuitive interpretation of NeRF rendering as a process that projects or “plots” the associated 3D point cloud to a 2D image plane. To this end, NeuralEditor introduces a novel rendering scheme based on deterministic integration within K-D tree-guided density-adaptive voxels, which produces both high-quality rendering results and precise point clouds through optimization. NeuralEditor then performs shape editing via mapping associated points between point clouds. Extensive evaluation shows that NeuralEditor achieves state-of-the-art performance in both shape deformation and scene morphing tasks. Notably, NeuralEditor supports both zeroshot inference and further fine-tuning over the edited scene. Our code, benchmark, and demo video are available at immortalco.github.io/NeuralEditor. Jun-Kun Chen, Jipeng Lyu, Yu-Xiong Wang |
CVPR | 1 |
| 2023 | Contrastive Learning Relies More on Spatial Inductive Bias Than Supervised Learning: An Empirical StudyabstractThough self-supervised contrastive learning (CL) has shown its potential to achieve state-of-the-art accuracy without any supervision, its behavior still remains under-investigated. Different from most previous work that understands CL from learning objectives, we focus on an unexplored yet natural aspect: the spatial inductive bias which seems to be implicitly exploited via data augmentations in CL. We design an experiment to study the reliance of CL on such spatial inductive bias, by destroying the global or local spatial structures of an image with global or local patch shuffling, and comparing the performance drop between experiments on original and corrupted dataset to quantify the reliance on certain inductive bias. We also use the uniformity of feature space to further research how CL-pre-trained models behave with the corrupted dataset. Our results and analysis show that CL has a much higher reliance on spatial inductive bias than SL, regardless of specific CL algorithm or backbones, opening a new direction for studying the behavior of CL. Yuanyi Zhong, Jun-Kun Chen, Yu-Xiong Wang |
ICCV | 3 |
| 2022 | PointTree: Transformation-Robust Point Cloud Encoder with Relaxed K-D Trees
Jun-Kun Chen, Yu-Xiong Wang |
ECCV (3) | 1 |
| 2022 | A3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and EditingabstractRecently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of speech understanding, but for the inverse direction, speech synthesis, the potential of representation learning is yet to be realized, due to the challenging nature of generating high-quality speech. To address this problem, we propose our framework, Alignment-Aware Acoustic-Text Pretraining (A$^3$T), which reconstructs masked acoustic signals with text input and acoustic-text alignment during training. In this way, the pretrained model can generate high quality reconstructed spectrogram, which can be applied to the speech editing and unseen speaker TTS directly. Experiments show A$^3$T outperforms SOTA models on speech editing, and improves multi-speaker speech synthesis without the external speaker verification model. He Bai 0002, Renjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang Huang 0001 |
ICML | 3 |
| 2021 | Improving Simultaneous Translation by Incorporating Pseudo-References with Fewer ReorderingsabstractSimultaneous translation is vastly different from full-sentence translation, in the sense that it starts translation before the source sentence ends, with only a few words delay.However, due to the lack of large-scale, high-quality simultaneous translation datasets, most such systems are still trained on conventional fullsentence bitexts.This is far from ideal for the simultaneous scenario due to the abundance of unnecessary long-distance reorderings in those bitexts.We propose a novel method that rewrites the target side of existing fullsentence corpora into simultaneous-style translation.Experiments on Zh!En and Ja!En simultaneous translation show substantial improvements (up to +2.7 BLEU) with the addition of these generated pseudo-references. Jun-Kun Chen, Renjie Zheng, Atsuhito Kita, Mingbo Ma, Liang Huang 0001 |
EMNLP (1) | 1 |
| 2021 | RNNLogic: Learning Logic Rules for Reasoning on Knowledge Graphs
Meng Qu, Jun-Kun Chen, Louis-Pascal A. C. Xhonneux, Yoshua Bengio, Jian Tang 0005 |
ICLR | 2 |
| 2021 | Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech TranslationabstractRecently, representation learning for text and speech has successfully improved many language related tasks. However, all existing methods suffer from two limitations: (a) they only learn from one input modality, while a unified representation for both speech and text is needed by tasks such as end-to-end speech translation, and as a result, (b) they can not exploit various large-scale text and speech data and their performance is limited by the scarcity of parallel speech translation data. To address these problems, we propose a Fused Acoustic and Text Masked Language Model (FAT-MLM) which jointly learns a unified representation for both acoustic and text input from various types of corpora including parallel data for speech recognition and machine translation, and even pure speech and text data. Within this cross-modal representation learning framework, we further present an end-to-end model for Fused Acoustic and Text Speech Translation (FAT-ST). Experiments on three translation directions show that by fine-tuning from FAT-MLM, our proposed speech translation models substantially improve translation quality by up to +5.9 BLEU. Renjie Zheng, Jun-Kun Chen, Mingbo Ma, Liang Huang 0001 |
ICML | 2 |
| 2021 | SpecRec: An Alternative Solution for Improving End-to-End Speech-to-Text Translation via Spectrogram Reconstruction
Jun-Kun Chen, Mingbo Ma, Renjie Zheng, Liang Huang 0001 |
Interspeech | 1 |
| 2019 | VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchabstractWe present a new large-scale multilingual video description dataset, VATEX1, which contains over 41,250 videos and 825, 000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation pairs. Compared to the widely-used MSRVTT dataset [64], VATEX is multilingual, larger, linguistically complex, and more diverse in terms of both video and natural language descriptions. We also introduce two tasks for video-and-language research based on VATEX: (1) Multilingual Video Captioning, aimed at describing a video in various languages with a compact unified captioning model, and (2) Video-guided Machine Translation, to translate a source language description into the target language using the video information as additional spatiotemporal context. Extensive experiments on the VATEX dataset show that, first, the unified multilingual model can not only produce both English and Chinese descriptions for a video more efficiently, but also offer improved performance over the monolingual models. Furthermore, we demonstrate that the spatiotemporal video context can be effectively utilized to align source and target languages and thus assist machine translation. In the end, we discuss the potentials of using VATEXfor other video-and-language research. Xin Wang 0061, Jiawei Wu 0003, Jun-Kun Chen, Lei Li 0005, Yuan-Fang Wang, William Yang Wang |
ICCV | 3 |
| 2018 | Meta Multi-Task Learning for Sequence ModelingabstractSemantic composition functions have been playing a pivotal role in neural representation learning of text sequences. In spite of their success, most existing models suffer from the underfitting problem: they use the same shared compositional function on all the positions in the sequence, thereby lacking expressive power due to incapacity to capture the richness of compositionality. Besides, the composition functions of different tasks are independent and learned from scratch. In this paper, we propose a new sharing scheme of composition function across multiple tasks. Specifically, we use a shared meta-network to capture the meta-knowledge of semantic composition and generate the parameters of the task-specific semantic composition models. We conduct extensive experiments on two types of tasks, text classification and sequence tagging, which demonstrate the benefits of our approach. Besides, we show that the shared meta-knowledge learned by our proposed model can be regarded as off-the-shelf knowledge and easily transferred to new tasks. Jun-Kun Chen, Xipeng Qiu, Pengfei Liu 0003, Xuanjing Huang 0001 |
AAAI | 1 |
| 2018 | Same Representation, Different Attentions: Shareable Sentence Representation Learning from Multiple TasksabstractDistributed representation plays an important role in deep learning based natural language processing. However, the representation of a sentence often varies in different tasks, which is usually learned from scratch and suffers from the limited amounts of training data. In this paper, we claim that a good sentence representation should be invariant and can benefit the various subsequent tasks. To achieve this purpose, we propose a new scheme of information sharing for multi-task learning. More specifically, all tasks share the same sentence representation and each task can select the task-specific information from the shared sentence representation with attention mechanisms. The query vector of each task's attention could be either static parameters or generated dynamically. We conduct extensive experiments on 16 different text classification tasks, which demonstrate the benefits of our architecture. Source codes of this paper are available on Github. Renjie Zheng, Jun-Kun Chen, Xipeng Qiu |
IJCAI | 2 |