EDBT 2026 Demo / reviewers in the wild / expert
Xu Chen 0053
dblp:83/6331-53
· DBLP profile ↗
10ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-1805-5435ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Coupling the Generator with Teacher for Effective Data-Free Knowledge Distillation
Xu Chen 0053, Yang Li 0251, Yahong Han, Guangquan Xu, Jialie Shen 0001 |
ICCV | 1 |
| 2025 | Ex Pede Herculem, Predicting Global Actionness Curve from Local ClipsabstractDense multi-label action detection in untrimmed long videos is a formidable task, with end-to-end training particularly challenging due to computational constraints, typically involving separate stages of off-the-shelf feature extraction and subsequent global modeling for action prediction.Existing methods fail to optimize all modules jointly for better performance. We introduce FreETAD, a Frequency-based End-to-end Temporal Action Detection approach, which shifts the focus from local actionness scores to frequency component estimation. Using the short-term Fourier Transform, FreETAD reconstructs the global action curve seamlessly. With a DETR-like decoder and frequency-encoded vectors for queries, it enhances multi-scale time-frequency interactions. FreETAD leverages end-to-end training effectively, boosting the mAP by 1.5% on Charades and 2.7% on MultiTHUMOS. Xu Chen 0053, Yang Li 0251, Yahong Han, Jialie Shen 0001 |
ACM Multimedia | 1 |
| 2025 | Information disentanglement for unsupervised domain adaptive Oracle Bone Inscriptions detection
Yongge Liu, Deng Li 0003, Xu Chen 0053, Runhua Jiang, Yahong Han |
Signal Process. Image Commun. | 4 |
| 2025 | A Static-Dynamic Composition Framework for Efficient Action RecognitionabstractThe dynamic inference, which adaptively allocates computational budgets for different samples, is a prevalent approach for achieving efficient action recognition. Current studies primarily focus on a data-efficient regime that reduces spatial or temporal redundancy, or their combination, by selecting partial video data, such as clips, frames, or patches. However, these approaches often utilize fixed and computationally expensive networks. From a different perspective, this article introduces a novel model-efficient regime that addresses network redundancy by dynamically selecting a partial network in real time. Specifically, we acknowledge that different channels of the neural network inherently contain redundant semantics either spatially or temporally. Therefore, by decreasing the width of the network, we can enhance efficiency while compromising the feature capacity. To strike a balance between efficiency and capacity, we propose the static-dynamic composition (SDCOM) framework, which comprises a static network with a fixed width and a dynamic network with a flexible width. In this framework, the static network extracts the primary feature with essential semantics from the input frame and simultaneously evaluates the gap toward achieving a comprehensive feature representation. Based on these evaluation results, the dynamic network activates a minimal width to extract a supplementary feature that fills the identified gap. We optimize the dynamic feature extraction through the employment of the slimmable network mechanism and a novel meta-learning scheme introduced in this article. Empirical analysis reveals that by combining the primary feature with an extremely lightweight supplementary feature, we can accurately recognize a large majority of frames (76%~92%). As a result, our proposed SDCOM significantly enhances recognition efficiency. For instance, on ActivityNet, FCVID, and Mini-Kinetics datasets, SDCOM saves 90% of the baseline's floating point operations (FLOPs) while achieving comparable or superior accuracy when compared with state-of-the-art methods. Xu Chen 0053, Yahong Han, Xiaojun Chang, Yifan Sun 0003, Yi Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Linking unknown characters via oracle bone inscriptions retrieval
Xu Chen 0053, Bang Li, Yongge Liu, Runhua Jiang, Yahong Han |
Multim. Syst. | 2 |
| 2023 | OraclePoints: A Hybrid Neural Representation for Oracle CharacterabstractOracle Bone Inscriptions (OBI) are ancient hieroglyphs originated in China and are considered one of the most famous writing systems in the world. Up to now, thousands of OBIs have been discovered, which require deciphering by experts to understand their contents. Experts typically need to restore, classify, and compare each character with previous inscriptions. Although existing research can assist with one of these operations, their performance falls short of practical requirements. In this work, we propose the OraclePoints framework, which represents OBI images as hybrid neural representations comprising features of images and point sets. The image representation provides inscription appearance and character structure, while the point representation makes it easy and effective to distinguish characters and noises. In addition, we demonstrate that OraclePoints can be easily integrated with existing models in a plug-and-play manner. Comprehensive experiments demonstrate that the proposed hybrid neural representation framework supports a range of OBI tasks, including character image retrieval, recognition, and denoising. It is also demonstrated that OraclePoints is helpful for deciphering OBIs by linking ancient characters to modern Chinese characters. Our codes are available at https://ddghjikle.github.io/. Runhua Jiang, Yongge Liu, Boyuan Zhang 0003, Xu Chen 0053, Deng Li 0003, Yahong Han |
ACM Multimedia | 4 |
| 2022 | Action Keypoint Network for Efficient Video RecognitionabstractReducing redundancy is crucial for improving the efficiency of video recognition models. An effective approach is to select informative content from the holistic video, yielding a popular family of dynamic video recognition methods. However, existing dynamic methods focus on either temporal or spatial selection independently while neglecting a reality that the redundancies are usually spatial and temporal, simultaneously. Moreover, their selected content is usually cropped with fixed shapes (e.g., temporally-cropped frames, spatially-cropped patches), while the realistic distribution of informative content can be much more diverse. With these two insights, this paper proposes to integrate temporal and spatial selection into an Action Keypoint Network (AK-Net). From different frames and positions, AK-Net selects some informative points scattered in arbitrary-shaped regions as a set of "action keypoints" and then transforms the video recognition into point cloud classification. More concretely, AK-Net has two steps, i.e., the keypoint selection and the point cloud classification. First, it inputs the video into a baseline network and outputs a feature map from an intermediate layer. We view each pixel on this feature map as a spatial-temporal point and select some informative keypoints using self-attention. Second, AK-Net devises a ranking criterion to arrange the keypoints into an ordered 1D sequence. Since the video is represented with a 1D sequence after the specified layer, AK-Net transforms the subsequent layers into a point cloud classification sub-net by compacting the original 2D convolutional kernels into 1D kernels. Consequentially, AK-Net brings two-fold benefits for efficiency: The keypoint selection step collects informative content within arbitrary shapes and increases the efficiency for modeling spatial-temporal dependencies, while the point cloud classification step further reduces the computational cost by compacting the convolutional kernels. Experimental results show that AK-Net can consistently improve the efficiency and performance of baseline methods on several video recognition benchmarks. Xu Chen 0053, Yahong Han, Yifan Sun 0003, Yi Yang 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Infrared Action Detection in the Dark via Cross-Stream Attention MechanismabstractAction detection plays an important role in the field of video understanding and attracts considerable attention in the last decade. However, current action detection methods are mainly based on visible videos, and few of them consider scenes with low-light, where actions are difficult to be detected by existing methods, or even by human eyes. Compared with visible videos, infrared videos are more suitable for the dark environment and resistant to background clutter. In this paper, we investigate the temporal action detection problem in the dark by using infrared videos, which is, to the best of our knowledge, the first attempt in the action detection community. Our model takes the whole video as input, a Flow Estimation Network (FEN) is employed to generate the optical flow for infrared data, and it is optimized with the whole network to obtain action-related motion representations. After feature extraction, the infrared stream and flow stream are fed into a Selective Cross-stream Attention (SCA) module to narrow the performance gap between infrared and visible videos. The SCA emphasizes informative snippets and focuses on the more discriminative stream automatically. Then we adopt a snippet-level classifier to obtain action scores for all snippets and link continuous snippets into final detection results. All these modules are trained in an end-to-end manner. We collect an Infrared action Detection (InfDet) dataset obtained in the dark and conduct extensive experiments to verify the effectiveness of the proposed method. Experimental results show that our proposed method surpasses the state-of-the-art temporal action detection methods designed for visible videos, and it also achieves the best performance compared with other infrared action recognition methods on both InfAR and Infrared-Visible datasets. Xu Chen 0053, Chenqiang Gao, Chaoyu Li, Yi Yang 0001, Deyu Meng |
IEEE Trans. Multim. | 1 |
| 2021 | Video-to-Image Casting: A Flatting Method for Video AnalysisabstractPrevious mainstream video analysis methods, especially 3D CNNs-based models, mainly aim to transfer frameworks from the image domain to the video domain, and they follow the regime which has been succeeded in image processing, i.e., large-scale benchmarks and deep networks. However, processing videos is still time-consuming due to the increased computational cost. In this paper, we propose to flat the video and construct a Spatio-temporal Image (STI), i.e., squeezing the temporal dimension into a spatial plane. To pursuit the video-level modeling and efficient architecture, we devise a Collective Convolution (CoConv) operation to replace the 2D convolution. With the holistic sampling strategy, this novel operation can extract the video-level spatio-temporal representation. Moreover, we ensure that each CoConv operation has the same number of parameters as the original 2D filter, thus we can utilize a 2D network equipped with CoConv to analyze videos without additional computations. To verify the effectiveness of our method for the general video analysis, we evaluate it on three typical tasks, i.e., supervised action recognition, self-supervised action recognition, and dynamic texture recognition. Extensive experimental results show that our method can achieve comparable or state-of-the-art performances on these benchmarks while using much fewer computations compared with its 3D counterpart. Xu Chen 0053, Chenqiang Gao, Feng Yang 0015, Yi Yang 0001, Yahong Han |
ACM Multimedia | 1 |
| 2019 | Pose detection in complex classroom environment based on improved Faster R-CNNabstractPose detection of small targets in poor imaging conditions like heavy occlusion and low resolution is still an open and challenging task in computer vision. For instance, detection of students' poses in classrooms that are even indistinguishable to human eyes remains a rather difficult task. Motivated by the success of convolutional feature merging and locality preserving, the authors propose a pose detection framework combining merged region of interest (ROI) pooling and locality preserving learning. Unlike usual object detection algorithms which use general top‐level convolutional features as inputs, their method uses a merged ROI pooling structure to merge semantic feature and high‐resolution feature from the last two levels of convolutional feature maps, so that this merged feature is made more expressive than the single‐level feature. In addition, the locality feature‐preserving learning is used in the last fully‐connected layer. Through locality preserving learning, features belonging to the same class would be forced to be closer in the feature space, which enables the model with stronger classification ability. Experimental results show that the proposed method outperforms the state‐of‐the‐art methods. Chenqiang Gao, Xu Chen 0053, Yue Zhao 0012 |
IET Image Process. | 3 |