EDBT 2026 Demo / reviewers in the wild / expert
Zilong Chen
dblp:87/8238
· DBLP profile ↗
22ranked-venue papers
4as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationabstractIn this paper, we introduce MeshGen, an advanced image-to-3D pipeline that generates high-quality 3D meshes with detailed geometry and physically based rendering (PBR) textures. Addressing the challenges faced by existing 3D native diffusion models, such as suboptimal auto-encoder performance, limited controllability, poor generalization, and inconsistent image-based PBR texturing, Mesh-Gen employs several key innovations to overcome these limitations. We pioneer a render-enhanced point-to-shape auto-encoder that compresses meshes into a compact latent space by designing perceptual optimization with ray-based regularization. This ensures that the 3D shapes are accurately represented and reconstructed to preserve geometric details within the latent space. To address data scarcity and image-shape misalignment, we further propose geometric augmentation and generative rendering augmentation techniques, which enhance the model’s controllability and gen-eralization ability, allowing it to perform well even with limited public datasets. For the texture generation, Mesh-Gen employs a reference attention-based multi-view Con-trolNet for consistent appearance synthesis. This is further complemented by our multi-view PBR decomposer that estimates PBR components and a UV inpainter that fills invisible areas, ensuring a seamless and consistent texture across the 3D mesh. Our extensive experiments demonstrate that MeshGen largely outperforms previous methods in both shape and texture generation, setting a new standard for the quality of 3D meshes generated with PBR textures. Zilong Chen, Yikai Wang 0001, Wenqiang Sun, Feng Wang 0034, Huaping Liu 0001 |
CVPR | 1 |
| 2025 | MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving PerceptionabstractAdvanced driver assistance systems require a comprehensive understanding of the driver’s mental/physical state and traffic context but existing works often neglect the potential benefits of joint learning between these tasks. This paper proposes MMTL-UniAD, a unified multi-modal multitask learning framework that simultaneously recognizes driver behavior (e.g., looking around, talking), driver emotion (e.g., anxiety, happiness), vehicle behavior (e.g., parking, turning), and traffic context (e.g., traffic jam, traffic smooth). A key challenge is avoiding negative transfer between tasks, which can impair learning performance. To address this, we introduce two key components into the framework: one is the multi-axis region attention network to extract global context-sensitive features, and the other is the dual-branch multimodal embedding to learn multi-modal embeddings from both task-shared and task-specific features. The former uses a multi-attention mechanism to extract task-relevant features, mitigating negative transfer caused by task-unrelated features. The latter employs a dual-branch structure to adaptively adjust task-shared and task-specific parameters, enhancing cross-task knowledge transfer while reducing task conflicts. We assess MMTL-UniAD on the AIDE dataset, using a series of ablation studies, and show that it outperforms state-of-the-art methods across all four tasks. The code is available on https://github.com/Wenzhuo-Liu/MMTL-UniAD. Wenzhuo Liu, Wenshuo Wang 0001, Yicheng Qiao, Qiannan Guo, Jiayin Zhu, Zilong Chen, Huiming Yang, Zhiwei Li 0011, Tiao Tan, Huaping Liu 0001 |
CVPR | 7 |
| 2025 | MeshAnything V2: Artist-Created Mesh Generation with Adjacent Mesh TokenizationabstractMeshes are the de facto 3D representation in the industry but are labor-intensive to produce. Recently, a line of research has focused on autoregressively generating meshes. This approach processes meshes into a sequence composed of vertices and then generates them vertex by vertex, similar to how a language model generates text. These methods have achieved some success but still struggle to generate complex meshes. One primary reason for this limitation is their inefficient tokenization methods. To address this issue, we introduce MeshAnything V2, an advanced mesh generation model designed to create Artist-Created Meshes that align precisely with specified shapes. A key innovation behind MeshAnything V2 is our novel Adjacent Mesh Tokenization (AMT) method. Unlike traditional approaches that represent each face using three vertices, AMT optimizes this by employing a single vertex wherever feasible, effectively reducing the token sequence length by about half on average. This not only streamlines the tokenization process but also results in more compact and well-structured sequences, enhancing the efficiency of mesh generation. With these improvements, MeshAnything V2 effectively doubles the face limit compared to previous models, delivering superior performance without increasing computational costs. We will make our code and models publicly available. Project Page: https://buaacyw.github.io/meshanything-v2/ Yikai Wang 0001, Yihao Luo, Zilong Chen, Jun Zhu 0001, Chi Zhang 0007, Guosheng Lin |
ICCV | 5 |
| 2025 | Dimensionx: Create Any 3D and 4D Scenes From a Single Image With Decoupled Video Diffusion
Wenqiang Sun, Fangfu Liu, Zilong Chen, Yueqi Duan, Jun Zhu 0001, Jun Zhang 0004, Yikai Wang 0001 |
ICCV | 4 |
| 2025 | TEM3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive DrivingabstractMulti-task learning (MTL) can advance assistive driving by exploring inter-task correlations through shared representations. However, existing methods face two critical limitations: single-modality constraints limiting comprehensive scene understanding and inefficient architectures impeding real-time deployment. This paper proposes TEM3-Learning (Time-Efficient Multimodal Multi-task Learning), a novel framework that jointly optimizes driver emotion recognition, driver behavior recognition, traffic context recognition, and vehicle behavior recognition through a two-stage architecture. The first component, the mamba-based multi-view temporal-spatial feature extraction subnetwork (MTS-Mamba), introduces a forward-backward temporal scanning mechanism and global-local spatial attention to efficiently extract low-cost temporal-spatial features from multi-view sequential images. The second component, the MTL-based gated multimodal feature integrator (MGMI), employs task-specific multi-gating modules to adaptively highlight the most relevant modality features for each task, effectively alleviating the negative transfer problem in MTL. Evaluation on the AIDE dataset, our proposed model achieves state-of-the-art accuracy across all four tasks, maintaining a lightweight architecture with fewer than 6 million parameters and delivering an impressive 142.32 FPS inference speed. Rigorous ablation studies further validate the effectiveness of the proposed framework and the independent contributions of each module. The code is available on https://github.com/Wenzhuo-Liu/TEM3-Learning. Wenzhuo Liu, Yicheng Qiao, Qiannan Guo, Zilong Chen, Meihua Zhou, Zhiwei Li 0011, Huaping Liu 0001, Wenshuo Wang 0001 |
IROS | 5 |
| 2025 | EOOD: End-to-end oriented object detection
Caiguang Zhang, Zilong Chen, Boli Xiong, Kefeng Ji, Gangyao Kuang |
Neurocomputing | 2 |
| 2025 | High-Frequency Centimeter-Accuracy Water Level Estimation in the Yangtze River Using Multi-GNSS Interferometric ReflectometryabstractWater level measurement is essential for managing and conserving water resources. The Yangtze River has a significant impact on agriculture, transportation, and local ecosystems, while traditional water level measurement still has some limitations, e.g. high cost and low coverage. Recently, Global Navigation Satellite System-Interferometric Reflectometry (GNSS-IR) has become a new means for water level monitoring. This paper integrates the reflected signals from GPS, GLONASS, Galileo and BDS constellations for real time high-frequency water level estimation at Yangtze River stations (BADO and DATO). By integrating a sliding window with Variational Mode Decomposition (VMD), Iteratively Reweighted Least Squares (IRLS), and Savitzky-Golay (S-G) filtering, water level measurements at a 5-minute temporal resolution are obtained and evaluated. Our individual and combined results show that the VMD method efficiently provides more accurate and stable inversion results using signal-to-noise ratio (SNR) data in each window. The combined and filtered results increased the accuracy by over 45.88% when compared to single-system measurements. The root mean square error (RMSE) was 4.86 cm at BADO and 2.28 cm at DATO with coefficient of determination (R2) 0.99 at both stations. Our results demonstrated that the combined method enabled precise and continuous water level monitoring, which provides new pathways for hydrological monitoring and real-time water resource management. Shuanggen Jin, Zilong Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Text-to-3D using Gaussian SplattingabstractAutomatic text-to-3D generation that combines Score Distillation Sampling (SDS) with the optimization of volume rendering has achieved remarkable progress in synthesizing realistic 3D objects. Yet most existing text-to-3D methods by SDS and volume rendering suffer from inaccurate geometry, e.g., the Janus issue, since it is hard to explicitly integrate 3D priors into implicit 3D representations. Besides, it is usually time-consuming for them to generate elaborate 3D models with rich colors. In response, this paper proposes GSGEN, a novel method that adopts Gaussian Splatting, a recent state-of-the-art representation, to text-to-3D generation. GSGEN aims at generating high-quality 3D objects and addressing existing shortcomings by exploiting the explicit nature of Gaussian Splatting that enables the incorporation of 3D prior. Specifically, our method adopts a progressive optimization strategy, which includes a geometry optimization stage and an appearance refinement stage. In geometry optimization, a coarse representation is established under 3D point cloud diffusion prior along with the ordinary 2D SDS optimization, ensuring a sensible and 3D-consistent rough shape. Subsequently, the obtained Gaussians undergo an iterative appearance refinement to enrich texture details. In this stage, we increase the number of Gaussians by compactness-based densification to enhance continuity and improve fidelity. With these designs, our approach can generate 3D assets with delicate details and accurate geometry. Extensive evaluations demonstrate the effectiveness of our method, especially for capturing high-frequency components. Our code is available at https://github.com/gsgen3d/gsgen. Zilong Chen, Feng Wang 0034, Yikai Wang 0001, Huaping Liu 0001 |
CVPR | 1 |
| 2024 | GaussianEditor: Swift and Controllable 3D Editing with Gaussian Splattingabstract3D editing plays a crucial role in many areas such as gaming and virtual reality. Traditional 3D editing methods, which rely on representations like meshes and point clouds, often fall short in realistically depicting complex scenes. On the other hand, methods based on implicit 3D representations, like Neural Radiance Field (NeRF), render complex scenes effectively but suffer from slow processing speeds and limited control over specific scene areas. In response to these challenges, our paper presents GaussianEditor, the first 3D editing algorithm based on Gaussian Splatting (GS), a novel 3D representation. GaussianEditor enhances precision and control in editing through our proposed Gaussian semantic tracing, which traces the editing target throughout the training process. Additionally, we propose Hierarchical Gaussian splatting (HGS) to achieve stabilized and fine results under stochastic generative guidance from 2D diffusion models. We also develop editing strategies for efficient object removal and integration, a challenging task for existing methods. Our comprehensive experiments demonstrate GaussianEditor's superior control, effective, and efficient performance, marking a significant advancement in 3D editing. Zilong Chen, Chi Zhang 0007, Feng Wang 0034, Yikai Wang 0001, Zhongang Cai, Lei Yang 0059, Huaping Liu 0001, Guosheng Lin |
CVPR | 2 |
| 2024 | Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian SurfelsabstractVideo generative models are receiving particular attention given their ability to generate realistic and imaginative frames. Besides, these models are also observed to exhibit strong 3D consistency, significantly enhancing their potential to act as world simulators. In this work, we present Vidu4D, a novel reconstruction model that excels in accurately reconstructing 4D (i.e., sequential 3D) representations from single generated videos, addressing challenges associated with non-rigidity and frame distortion. This capability is pivotal for creating high-fidelity virtual contents that maintain both spatial and temporal coherence. At the core of Vidu4D is our proposed Dynamic Gaussian Surfels (DGS) technique. DGS optimizes time-varying warping functions to transform Gaussian surfels (surface elements) from a static state to a dynamically warped state. This transformation enables a precise depiction of motion and deformation over time. To preserve the structural integrity of surface-aligned Gaussian surfels, we design the warped-state geometric regularization based on continuous warping fields for estimating normals. Additionally, we learn refinements on rotation and scaling parameters of Gaussian surfels, which greatly alleviates texture flickering during the warping process and enhances the capture of fine-grained appearance details. Vidu4D also contains a novel initialization state that provides a proper start for the warping fields in DGS. Equipping Vidu4D with an existing video generative model, the overall framework demonstrates high-fidelity text-to-4D generation in both appearance and geometry. Yikai Wang 0001, Zilong Chen, Fuchun Sun 0001, Jun Zhu 0001 |
NeurIPS | 3 |
| 2024 | Efficient image generation with Contour Wavelet Diffusion
Dimeng Zhang, Zilong Chen, Yuntao Zou |
Comput. Graph. | 3 |
| 2023 | BIC: Twitter Bot Detection with Text-Graph Interaction and Semantic ConsistencyabstractZhenyu Lei, Herun Wan, Wenqian Zhang, Shangbin Feng, Zilong Chen, Jundong Li, Qinghua Zheng, Minnan Luo. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhenyu Lei 0004, Herun Wan, Shangbin Feng, Zilong Chen, Jundong Li, Minnan Luo |
ACL (1) | 5 |
| 2023 | Masked Space-Time Hash Encoding for Efficient Dynamic Scene ReconstructionabstractIn this paper, we propose the Masked Space-Time Hash encoding (MSTH), a novel method for efficiently reconstructing dynamic 3D scenes from multi-view or monocular videos. Based on the observation that dynamic scenes often contain substantial static areas that result in redundancy in storage and computations, MSTH represents a dynamic scene as a weighted combination of a 3D hash encoding and a 4D hash encoding. The weights for the two components are represented by a learnable mask which is guided by an uncertainty-based objective to reflect the spatial and temporal importance of each 3D position. With this design, our method can reduce the hash collision rate by avoiding redundant queries and modifications on static areas, making it feasible to represent a large number of space-time voxels by hash tables with small size.Besides, without the requirements to fit the large numbers of temporally redundant features independently, our method is easier to optimize and converge rapidly with only twenty minutes of training for a 300-frame dynamic scene. We evaluate our method on extensive dynamic scenes. As a result, MSTH obtains consistently better results than previous state-of-the-art methods with only 20 minutes of training time and 130 MB of memory storage. Feng Wang 0034, Zilong Chen, Guokang Wang, Huaping Liu 0001 |
NeurIPS | 2 |
| 2023 | KRACL: Contrastive Learning with Graph Context Modeling for Sparse Knowledge Graph CompletionabstractKnowledge Graph Embeddings (KGE) aim to map entities and relations to low dimensional spaces and have become the de-facto standard for knowledge graph completion. Most existing KGE methods suffer from the sparsity challenge, where it is harder to predict entities that appear less frequently in knowledge graphs. In this work, we propose a novel framework KRACL1 to alleviate the widespread sparsity in KGs with graph context and contrastive learning. Firstly, we propose the Knowledge Relational Attention Network (KRAT) to leverage the graph context by simultaneously projecting neighboring triples to different latent spaces and jointly aggregating messages with the attention mechanism. KRAT is capable of capturing the subtle semantic information and importance of different context triples as well as leveraging multi-hop information in knowledge graphs. Secondly, we propose the knowledge contrastive loss by combining the contrastive loss with cross entropy loss, which introduces more negative samples and thus enriches the feedback to sparse entities. Our experiments demonstrate that KRACL achieves superior results across various standard knowledge graph benchmarks, especially on WN18RR and NELL-995 which have large numbers of low in-degree entities. Extensive experiments also bear out KRACL’s effectiveness in handling sparse knowledge graphs and robustness against noisy triples. Zhaoxuan Tan, Zilong Chen, Shangbin Feng, Qingyue Zhang 0003, Jundong Li, Minnan Luo |
WWW | 2 |
| 2023 | A trajectory prediction method based on bayonet importance encoding and bidirectional LSTM
Lechen Guan, Jingtian Shi, Dongle Wang, Hu Shao, Zilong Chen, Danhua Chu |
Expert Syst. Appl. | 5 |
| 2022 | PAR: Political Actor Representation Learning with Social Context and Expert KnowledgeabstractShangbin Feng, Zhaoxuan Tan, Zilong Chen, Ningnan Wang, Peisheng Yu, Qinghua Zheng, Xiaojun Chang, Minnan Luo. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Shangbin Feng, Zhaoxuan Tan, Zilong Chen, Ningnan Wang, Peisheng Yu, Xiaojun Chang, Minnan Luo |
EMNLP | 3 |
| 2022 | Attention-based Broad Self-guided Network for Low-light Image EnhancementabstractLow-light image enhancement is widely used in many fields, such as target detection, face recognition, and image segmentation. In recent years, Deep Learning methods have achieved impressive breakthroughs in low-light image enhancement. However, most of them mine high-dimensional features of images by stacking network structures and deepening the depth of the network, which causes more runtime costs for single image enhancement. To reduce inference time while fully extracting local features and global features of low-light images, we propose an Attention-based Broad Self-guided Network for real-world low-light image Enhancement. Compared to U-net structure [1], such a self-guided structure requires a smaller number of parameters and enables us to achieve better effectiveness. To extract local information more efficiently, we designed a Multilevel Guided Densely Connected Attention block, which can be considered as a novel extension of the densely connected blocks in feature space. In addition, we also offer a more efficient module to extract global information from images, which is called the Global Spatial Attention module. The proposed network is validated by lots of mainstream benchmarks. Many experimental results show that the proposed network outperforms most state-of-the-art low-light image Enhancement methods. Our code is available at https://github.com/paullenwyue/ABSGNet. Zilong Chen, Yaling Liang, Minghui Du |
ICPR | 1 |
| 2022 | KCD: Knowledge Walks and Textual Cues Enhanced Political Perspective Detection in News MediaabstractWenqian Zhang, Shangbin Feng, Zilong Chen, Zhenyu Lei, Jundong Li, Minnan Luo. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Shangbin Feng, Zilong Chen, Zhenyu Lei 0004, Jundong Li, Minnan Luo |
NAACL-HLT | 3 |
| 2022 | TwiBot-22: Towards Graph-Based Twitter Bot DetectionabstractTwitter bot detection has become an increasingly important task to combat misinformation, facilitate social media moderation, and preserve the integrity of the online discourse. State-of-the-art bot detection methods generally leverage the graph structure of the Twitter network, and they exhibit promising performance when confronting novel Twitter bots that traditional methods fail to detect. However, very few of the existing Twitter bot detection datasets are graph-based, and even these few graph-based datasets suffer from limited dataset scale, incomplete graph structure, as well as low annotation quality. In fact, the lack of a large-scale graph-based Twitter bot detection benchmark that addresses these issues has seriously hindered the development and evaluation of novel graph-based bot detection approaches. In this paper, we propose TwiBot-22, a comprehensive graph-based Twitter bot detection benchmark that presents the largest dataset to date, provides diversified entities and relations on the Twitter network, and has considerably better annotation quality than existing datasets. In addition, we re-implement 35 representative Twitter bot detection baselines and evaluate them on 9 datasets, including TwiBot-22, to promote a fair comparison of model performance and a holistic understanding of research progress. To facilitate further research, we consolidate all implemented codes and datasets into the TwiBot-22 evaluation framework, where researchers could consistently evaluate new models and datasets. The TwiBot-22 Twitter bot detection benchmark and evaluation framework are publicly available at \url{https://twibot22.github.io/}. Shangbin Feng, Zhaoxuan Tan, Herun Wan, Ningnan Wang, Zilong Chen, Binchi Zhang, Zhenyu Lei 0004, Xinshun Feng, Qingyue Zhang 0003, Hongrui Wang 0004, Yuhan Liu 0028, Yuyang Bai, Heng Wang 0008, Zijian Cai, Lijing Zheng, Zihan Ma 0001, Jundong Li, Minnan Luo |
NeurIPS | 5 |
| 2018 | Fingerprint Image Enhancement Method based on Adaptive Median FilterabstractThe traditional median filtering method uses a fixed filter window size method to remove the impulse noise in a fingerprint image. If the filtering window size is small, the traditional median filtering method will not filter out the impulse noise completely. If the filtering window size is large, the fingerprint image may become blurred. To solve the problem, a method based on adaptive median filter is proposed for fingerprint image enhancement processing and impulse noise removal in the paper. The use of adaptive median filtering to remove the impulse noise of the fingerprint image mainly involves three steps. First, the size of the adaptive median filter window is initialized, and it is judged whether the center pixel of the filter window in the fingerprint image is impulse noise. Second, the size of the filter window is determined based on the median value, the maximum value, and the minimum value within the filter window. Finally, median filtering is performed on the fingerprint image under the filter window size obtained in the previous steps, and the filter output value is used instead of the window center pixel value. The method is tested on rolled fingerprint images contaminated by impulse noise and fingerprint images contaminated by impulse noise from a crime scene. Experimental results show that the method based on adaptive median filter for fingerprint image enhancement outperforms the traditional median filtering method in filtering impulse noise performance. Zizheng Wang, Zilong Chen |
APCC | 3 |
| 2013 | Delay and backlog distribution analysis of Amplify-and-Forward cooperative channels: A stochastic network calculus perspectiveabstractAn accurate and comprehensive delay and backlog performance evaluation is crucial for the quality of service (QoS) guarantee in cooperative communication. However, the classical deterministic network calculus and queueing theory is not suitable for the wireless networks due to their inherent random behaviors. For this aim, we employ stochastic network calculus to analyze the Amplify-and-Forward cooperative communication where the inherent time-varying service characteristic and traffic self-similarity are considered. In particular, a stochastic service curve for Amplify-and-Forward cooperative channel is derived. We use a stochastic arrival curve to describe the self-similar traffic which is characterized by Fractional Brownian Motion (FBM). Furthermore, the closed-form expressions of delay and backlog distribution are derived. Simulation results validate the accuracy of the theoretical analysis. More importantly, the proposed analytical framework can be used to evaluate the performance distribution under various self-similar traffic patterns. Lie Jiang, Zilong Chen |
WCNC | 4 |
| 2010 | Using Text Classification Method in Relevance Feedback
Zilong Chen |
ACIIDS (2) | 1 |