EDBT 2026 Demo / reviewers in the wild / expert
Feng Chen 0047
dblp:21/3047-47
· DBLP profile ↗
20ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0003-1800-8441ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMsabstractExisting sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge the training–inference gap and lack the capacity for fine-grained token selection across multiple dimensions—such as queries, key-values (KV), and heads—leading to suboptimal performance and acceleration gains. In this paper, we introduce OmniSparse, a training-aware fine-grained sparse attention of long-video MLLMs, which is applied in both training and inference with dynamic token budget allocation. Specifically, OmniSparse contains three adaptive and complementary mechanisms: (1) query selection as lazy-active classification, aiming to retain active queries that capture broader semantic similarity, while discarding most of lazy ones that focus on limited local context and exhibit high functional redundancy with their neighbors, (2) KV selection with head-level dynamic budget allocation, where a shared budget is determined based on the flattest head and applied uniformly across all heads to ensure attention recall after selection, and (3) KV cache slimming to alleviate head-level redundancy, which selectively fetches visual KV cache according to the head-level decoding query pattern. Experimental results demonstrate that OmniSparse can achieve comparable performance with full attention, achieving 2.7x speedup during prefill and 2.4x memory reduction for decoding. Feng Chen 0047, Yefei He, Shaoxuan He, Yuanyu He, Jing Liu 0048, Lequan Lin, Akide Liu, Zhenbang Sun, Bohan Zhuang, Qi Wu 0001 |
AAAI | 1 |
| 2025 | Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image SynthesisabstractConditional image synthesis is a crucial task with broad applications, such as artistic creation and virtual reality. However, current generative methods are often task-oriented with a narrow scope, handling a restricted condition with constrained applicability. In this paper, we propose a novel approach that treats conditional image synthesis as the modular combination of diverse fundamental condition units. Specifically, we divide conditions into three primary units: text, layout, and drag. To enable effective control over these conditions, we design a dedicated alignment module for each. For the text condition, we introduce a Dense Concept Alignment (DCA) module, which achieves dense visual-text alignment by drawing on diverse textual concepts. For the layout condition, we propose a Dense Geometry Alignment (DGA) module to enforce comprehensive geometric constraints that preserve the spatial configuration. For the drag condition, we introduce a Dense Motion Alignment (DMA) module to apply multi-level motion regularization, ensuring that each pixel follows its desired trajectory without visual artifacts. By flexibly inserting and combining these alignment modules, our framework enhances the model’s adaptability to diverse conditional generation tasks and greatly expands its application range. Extensive experiments demonstrate the superior performance of our framework across a variety of conditions, including textual description, segmentation mask (bounding box), drag manipulation, and their combinations. Code is available at https://github.com/ZixuanWang0525/DADG Duo Peng, Feng Chen 0047, Yinjie Lei |
CVPR | 3 |
| 2025 | ZipVL: Accelerating Vision-Language Models Through Dynamic Token Sparsity
Yefei He, Feng Chen 0047, Wenqi Shao, Kaipeng Zhang, Bohan Zhuang |
ICCV | 2 |
| 2025 | Neighboring Autoregressive Modeling for Efficient Visual GenerationabstractVisual autoregressive models typically adhere to a raster-order ``next-token prediction" paradigm, which overlooks the spatial and temporal locality inherent in visual content. Specifically, visual tokens exhibit significantly stronger correlations with their spatially or temporally adjacent tokens compared to those that are distant. In this paper, we propose Neighboring Autoregressive Modeling (NAR), a novel paradigm that formulates autoregressive visual generation as a progressive outpainting procedure, following a near-to-far ``next-neighbor prediction" mechanism. Starting from an initial token, the remaining tokens are decoded in ascending order of their Manhattan distance from the initial token in the spatial-temporal space, progressively expanding the boundary of the decoded region. To enable parallel prediction of multiple adjacent tokens in the spatial-temporal space, we introduce a set of dimension-oriented decoding heads, each predicting the next token along a mutually orthogonal dimension. During inference, all tokens adjacent to the decoded tokens are processed in parallel, substantially reducing the model forward steps for generation. Experiments on ImageNet$256\times 256$ and UCF101 demonstrate that NAR achieves 2.4$\times$ and 8.6$\times$ higher throughput respectively, while obtaining superior FID/FVD scores for both image and video generation tasks compared to the PAR-4X approach. When evaluating on text-to-image generation benchmark GenEval, NAR with 0.8B parameters outperforms Chameleon-7B while using merely 0.4 of the training data. Code is available at https://github.com/ThisisBillhe/NAR. Yefei He, Yuanyu He, Shaoxuan He, Feng Chen 0047, Kaipeng Zhang, Bohan Zhuang |
ICCV | 4 |
| 2025 | ZipAR: Parallel Autoregressive Image Generation through Spatial LocalityabstractIn this paper, we propose ZipAR, a training-free, plug-and-play parallel decoding framework for accelerating autoregressive (AR) visual generation. The motivation stems from the observation that images exhibit local structures, and spatially distant regions tend to have minimal interdependence. Given a partially decoded set of visual tokens, in addition to the original next-token prediction scheme in the row dimension, the tokens corresponding to spatially adjacent regions in the column dimension can be decoded in parallel. To ensure alignment with the contextual requirements of each token, we employ an adaptive local window assignment scheme with rejection sampling analogous to speculative decoding. By decoding multiple tokens in a single forward pass, the number of forward passes required to generate an image is significantly reduced, resulting in a substantial improvement in generation efficiency. Experiments demonstrate that ZipAR can reduce the number of model forward passes by up to 91% on the Emu3-Gen model without requiring any additional retraining. Yefei He, Feng Chen 0047, Yuanyu He, Shaoxuan He, Kaipeng Zhang, Bohan Zhuang |
ICML | 2 |
| 2025 | Learning multi-granularity representation with transformer for visible-infrared person re-identification
Yujian Feng, Feng Chen 0047, Guozi Sun, Fei Wu 0004, Yimu Ji 0001, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001 |
Pattern Recognit. | 2 |
| 2025 | Homogeneous and heterogeneous relational graph for visible-infrared person re-identification
Yujian Feng, Feng Chen 0047, Jian Yu 0007, Yimu Ji 0001, Fei Wu 0004, Shangdong Liu, Xiaoyuan Jing |
Pattern Recognit. | 2 |
| 2024 | Local aggressive and physically realizable adversarial attacks on 3D point cloud
Zhiyu Chen 0004, Feng Chen 0047, Mingjie Wang 0001, Shangdong Liu, Yimu Ji 0001 |
Comput. Secur. | 2 |
| 2024 | Cross-Modality Spatial-Temporal Transformer for Video-Based Visible-Infrared Person Re-IdentificationabstractVideo-based visible-infrared person re-identification (VVI-ReID) aims to match the identity of a person captured in video sequences from both visible and infrared cameras. The VVI-ReID task requires considering both the spatial relationship between body parts within each frame and the temporal change of appearance between successive frames. Existing VVI Re-ID methods employ Convolutional Neural Networks to extract local spatial features and Long Short-Term Memory to form temporal associations. However, these methods can not effectively capture the global spatial feature and the long-range temporal dependencies in ultra-long sequences. In this paper, we propose a Cross-modality Spatial-temporal Transformer (CST) including a Cross-frame Tube Transformer Module (CTTM) and a Multi-frame Transformer Fusion Module (MTFM) to address these challenges. Firstly, CTTM tokenizes a video clip into multiple 3D tubes, each encapsulating local spatial-temporal information of pedestrians, and then obtains global spatial-temporal representations by establishing the relationship between tubes. Secondly, we design MTFM to exchange information between multiple frames using message tokens, thus modeling the long-range temporal dependencies of features of pedestrians. In addition, to prevent the potential representation collapse caused by triplet-based loss functions, we propose a diversity-consistency (DC) loss function to preserve the diversity and consistency of cross-modality feature representations by imposing variance, invariance, and covariance constraints in feature representations. Extensive benchmark experiments demonstrate that our approach outperforms the state-of-the-art methods with large margins. Yujian Feng, Feng Chen 0047, Jian Yu 0007, Yimu Ji 0001, Fei Wu 0004, Tianliang Liu, Shangdong Liu, Xiaoyuan Jing, Jiebo Luo 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Uncertainty-guided Learning for Improving Image Manipulation DetectionabstractImage manipulation detection (IMD) is of vital importance as faking images and spreading misinformation can be malicious and harm our daily life. IMD is the core technique to solve these issues and poses challenges in two main aspects: (1) Data Uncertainty, i.e., the manipulated artifacts are often hard for humans to discern and lead to noisy labels, which may disturb model training; (2) Model Uncertainty, i.e., the same object may hold different categories (tampered or not) due to manipulation operations, which could potentially confuse the model training and result in unreliable outcomes. Previous works mainly focus on solving the model uncertainty issue by designing meticulous features and networks, however, the data uncertainty problem is rarely considered. In this paper, we address both problems by introducing an uncertainty-guided learning framework, which measures data and model uncertainties by a novel Uncertainty Estimation Network (UEN). UEN is trained under dynamic supervision, and outputs estimated uncertainty maps to refine manipulation detection results, which significantly alleviates the learning difficulties. To our knowledge, this is the first work to embed uncertainty modeling into IMD. Extensive experiments on various datasets demonstrate state-of-the-art performance, validating the effectiveness and generalizability of our method. Kaixiang Ji, Feng Chen 0047, Xin Guo 0010, Yadong Xu, Jian Wang 0108, Jingdong Chen |
ICCV | 2 |
| 2023 | Improving Stock Trend Prediction with Multi-granularity Denoising Contrastive LearningabstractStock trend prediction (STP) aims to predict the price fluctuation, which is critical in financial trading. The existing STP approaches only use market data with the same granularity (such as daily market data). However, in the actual financial investment, there are a large number of more detailed investment signals contained in finer-grained data (e.g, high- frequency data). This motivates us to research how to leverage multi-granularity market data to capture more useful information and improve the accuracy in the task of STP. However, the effective utilization of multi-granularity data presents a major challenge. Firstly, the iteration of multi-granularity data with time will lead to more complex noise, which makes it difficult to identify and extract it. Secondly, the difference in granularity may lead to opposite target trends in the same time interval. Thirdly, the target trends of stocks with similar features can be quite different, and different sizes of granularity will aggravate this gap. In order to address the above three challenges, in this paper, we present a self-supervised framework of multi-granularity denoising contrast learning (MDC). Specifically, we construct a dynamic dictionary of memory, which can obtain clear and unified representations by filtering noise and aligning multi- granularity data. Moreover, we design two contrast learning objectives to solve the differences in trends by constructing additional self-supervised signals. Extensive experiments on the CSI 300 datasets show that our framework stands out from the existing top-level systems and has excellent profitability in real investing scenarios. Mingjie Wang 0001, Feng Chen 0047, Jianxiong Guo, Weijia Jia 0001 |
IJCNN | 2 |
| 2023 | Learning Implicit Entity-object Relations by Bidirectional Generative Alignment for Multimodal NERabstractThe challenge posed by multimodal named entity recognition (MNER) is mainly two-fold: (1) bridging the semantic gap between text and image and (2) matching the entity with its associated object in image. Existing methods fail to capture the implicit entity-object relations, due to the lack of corresponding annotation. In this paper, we propose a bidirectional generative alignment method named BGA-MNER to tackle these issues. Our BGA-MNER consists of image2text and text2image generation with respect to entity-salient content in two modalities. It jointly optimizes the bidirectional reconstruction objectives, leading to aligning the implicit entity-object relations under such direct and powerful constraints. Furthermore, image-text pairs usually contain unmatched components which are noisy for generation. A stage-refined context sampler is proposed to extract the matched cross-modal content for generation. Extensive experiments on two benchmarks demonstrate that our method achieves state-of-the-art performance without image input during inference. Feng Chen 0047, Jiajia Liu 0002, Kaixiang Ji, Wang Ren, Jian Wang 0108, Jingdong Chen |
ACM Multimedia | 1 |
| 2023 | Visible-Infrared Person Re-Identification via Cross-Modality Interaction TransformerabstractVisible-infrared person re-identification (VI Re-ID) is designed to match person images of the same identity from visible and infrared cameras. Transformer structures have been successfully applied in the field of VI Re-ID. However, previous Transformer-based methods were mainly designed to capture global content information in a single modality, and could not simultaneously perceive semantic information between two modalities from a global perspective. To solve this problem, we propose a novel framework named the cross-modality interaction Transformer (CMIT). It has strong abilities in modeling spatial and sequential features that can capture dependencies between long-range features, and explicitly improves the discriminativeness of features by exchanging information across modalities, thus contributing to obtaining modality-invariant representations. Specifically, CMIT utilizes a cross-modality attention mechanism to enrich the feature representations of each patch token by interacting with the patch tokens of the other modality, and aggregates local features of the CNN structure and global information of the Transformer structure to mine feature saliency representation. Furthermore, the modality-discriminative (MD) loss function is proposed to learn potential consistency between modalities to encourage intra-modality compactness within class and inter-modality separation between classes. Extensive experiments on two benchmarks demonstrate that our approach outperforms state-of-the-art methods. Yujian Feng, Jian Yu 0007, Feng Chen 0047, Yimu Ji 0001, Fei Wu 0004, Shangdong Liu, Xiaoyuan Jing |
IEEE Trans. Multim. | 3 |
| 2022 | JSPNet: Learning joint semantic & instance segmentation of point clouds via feature self-similarity and cross-task probability
Feng Chen 0047, Fei Wu 0004, Guangwei Gao, Yimu Ji 0001, Guoping Jiang, Xiaoyuan Jing |
Pattern Recognit. | 1 |
| 2021 | Local Aggressive Adversarial Attacks on 3D Point CloudabstractDeep neural networks are found to be prone to adversarial examples which could deliberately fool the model to make mistakes. Recently, a few of works expand this task from 2D image to 3D point cloud by using global point cloud optimization. However, the perturbations of global point are not effective for misleading the victim model. First, not all points are important in optimization toward misleading. Abundant points account considerable distortion budget but contribute trivially to attack. Second, the multi-label optimization is suboptimal for adversarial attack, since it consumes extra energy in finding multi-label victim model collapse and causes instance transformation to be dissimilar to any particular instance. Third, the independent adversarial and perceptibility losses, caring misclassification and dissimilarity separately, treat the updating of each point equally without a focus. Therefore, once perceptibility loss approaches its budget threshold, all points would be stock in the surface of hypersphere and attack would be locked in local optimality. Therefore, we propose a local aggressive adversarial attacks (L3A) to solve above issues. Technically, we select a bunch of salient points, the high-score subset of point cloud according to gradient, to perturb. Then a flow of aggressive optimization strategies are developed to reinforce the unperceptive generation of adversarial examples toward misleading victim models. Extensive experiments on PointNet, PointNet++ and DGCNN demonstrate the state-of-the-art performance of our method against existing adversarial attack methods. Feng Chen 0047, Zhiyu Chen 0004, Mingjie Wang 0001 |
ACML | 2 |
| 2021 | Adaptive deformable convolutional network
Feng Chen 0047, Fei Wu 0004, Guangwei Gao, Qi Ge, Xiaoyuan Jing |
Neurocomputing | 1 |
| 2021 | Efficient Cross-Modality Graph Reasoning for RGB-Infrared Person Re-IdentificationabstractThe modality and pose variance between RGB and infrared (IR) images are two key challenges for RGB-IR person re-identification. Existing methods mainly focus on leveraging pixel or feature alignment to handle the intra-class variations and cross-modality discrepancy. However, these methods are hard to keep semantic identity consistency between global and local representation, which the consistency is important for the cross-modality pedestrian re-identification task. In this work, we propose a novel cross-modality graph reasoning method (CGRNet) to globally model and reason over relations between modalities and context, and to keep semantic identity consistency between global and local representation. Specifically, we propose a local modality-similarity module to put the distribution of modality-specific features into a common subspace without losing identity information. Besides, we squeeze the input feature of RGB and IR images into a channel-wise global vector, and through graph reasoning, the identity relationship and modality relationship in each vector are inferred. Extensive experiments on two datasets demonstrate the superior performance of our approach over the existing state-of-the-art. The code is available athttps://github.com/fegnyujian/CGRNet. Yujian Feng, Feng Chen 0047, Yimu Ji 0001, Fei Wu 0004 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Dynamic attention network for semantic segmentation
Fei Wu 0004, Feng Chen 0047, Xiaoyuan Jing, Changhui Hu 0001, Qi Ge, Yimu Ji 0001 |
Neurocomputing | 2 |
| 2020 | Multi-view semantic learning network for point cloud based 3D object detection
Yongguang Yang, Feng Chen 0047, Fei Wu 0004, Deliang Zeng, Yimu Ji 0001, Xiaoyuan Jing |
Neurocomputing | 2 |
| 2020 | SiENet: Siamese Expansion Network for Image ExtrapolationabstractDifferent from image inpainting, image outpainting has relatively less context in the image center to capture and more content at the image border to predict. Therefore, classical encoder-decoder pipeline of existing methods may not predict the outstretched unknown content perfectly. In this paper, a novel two-stage siamese adversarial model for image extrapolation, named Siamese Expansion Network (SiENet) is proposed. Specifically, in two stages, a novel border sensitive convolution named adaptive filling convolution is designed for allowing encoder to predict the unknown content, alleviating the burden of decoder. Besides, to introduce prior knowledge to network and reinforce the inferring ability of encoder, siamese adversarial mechanism is designed to enable our network to model the distribution of covered long range feature as that of uncovered image feature. The results on four datasets has demonstrated that our method outperforms existing state-of-the-arts and could produce realistic results. Our code is released on https://github.com/nanjingxiaobawang/SieNet-Image-extrapolation. Feng Chen 0047, Cailing Wang, Ming Tao 0002, Guoping Jiang |
IEEE Signal Process. Lett. | 2 |