VLDB 2026 Research / reviewers in the wild / expert
Chen Ju
dblp:221/1300
· DBLP profile ↗
20ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FishFlow: A LLM-Empowered Dynamic Pricing Framework for Online Fleamarket Platform
Kakam Chong, Shuai Xiao 0002, Chen Ju, Fei Huang 0002, Yuantao Gu, Shuguang Han, Jufeng Chen |
WWW | 5 |
| 2026 | Disentangling Representations from Search Behaviors for Recommendation via Counterfactual LearningabstractFor recommender systems in internet platforms, search activities provide additional insights into user interest through query-click interactions with items, and are thus widely used for enhancing personalized recommendation. However, these interacted items have not only transferable features that match users’ interests and are beneficial to the recommendation domain, but also have features related to users’ unique intents in the search domain. Such a domain gap of item features is neglected by most current search-enhanced recommendation methods. They directly incorporate these search behaviors into recommendation, and thus introduce partial negative transfer. Tackling this problem is challenging due to the lack of explicit supervision signals to disentangle features matching search-specific intent or general interest. To address this, we propose ClardRec, a c ounterfactual l e a rning-driven r epresentation d isentanglement framework for search-enhanced recommendation, based on the common belief that a user would click an item under a query not solely because of the item-query match but also due to the item’s query-independent general features (e.g., color or style) that interest the user. These general features exclude the reflection of search-specific intents contained in queries, ensuring a pure match to users’ underlying interests to complement recommendation. We perform the disentanglement based on a counterfactual thinking idea, how would user preferences and query match change for items if we removed their query-related features in search. Specifically, we leverage search queries to construct counterfactual signals to disentangle item representations, isolating only query-independent general features. These representations subsequently enable feature augmentation and data augmentation for the recommendation scenario. Comprehensive experiments on real datasets demonstrate that ClardRec is effective in both collaborative filtering and sequential recommendation scenarios. The source code is available at https://github.com/JJCui96/ClardRec . Jiajun Cui, Xu Chen 0026, Shuai Xiao 0002, Chen Ju, Jinsong Lan, Jianyong Wang 0001, Wei Zhang 0056 |
ACM Trans. Inf. Syst. | 4 |
| 2025 | A Pattern-Aware Finite Element Matrix Assembly Method on GPUsabstractThe Finite Element Method (FEM) is a fundamental technique for solving large-scale and complex engineering problems. During the construction of the system equations, the efficiency of finite element matrix assembly plays a crucial role in the overall performance. However, existing approaches often overlook the sensitivity of assembly algorithm performance to mesh characteristics, making it difficult to achieve optimal performance across diverse problems. In this work, we propose a novel pattern-aware FEM matrix assembly method on GPUs. To this end, we thoroughly analyze the key factors affecting performance and extract a set of potentially influential mesh features and density representations. Based on this, we construct a Deep learning-based prediction model that fully captures the input mesh characteristics to predict the performance-optimal assembly strategy. Experimental results on mesh datasets with a wide range of feature variations demonstrate that our method achieves remarkable prediction accuracy and delivers up to$7.34 \times$speedup in execution time compared to state-of-the-art approaches. To the best of our knowledge, this is the first work that introduces auto-tuning for the FEM matrix assembly process. Changyou Zhang, Zhuo Tian, Guangzhao Li, Chen Ju |
CLUSTER | 6 |
| 2025 | Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-trainingabstractIn rapidly evolving field of vision-language models (VLMs), contrastive language-image pre-training (CLIP) has made significant strides, becoming foundation for various downstream tasks. However, relying on one-to-one (image, text) contrastive paradigm to learn alignment from large-scale messy web data, CLIP faces a serious myopic dilemma, resulting in biases towards monotonous short texts and shallow visual expressivity. To overcome these issues, this paper advances CLIP into one novel holistic paradigm, by updating both diverse data and alignment optimization. To obtain colorful data with low cost, we use image-to-text captioning to generate multi-texts for each image, from multiple perspectives, granularities, and hierarchies. Two gadgets are proposed to encourage textual diversity. To match such (image, multi-texts) pairs, we modify the CLIP image encoder into multi-branch, and propose multi-to-multi contrastive optimization for image-text part-to-part matching. As a result, diverse visual embeddings are learned for each image, bringing good interpretability and generalization. Extensive experiments and ablations across over ten benchmarks indicate that our holistic CLIP significantly outperforms existing myopic CLIP, including image-text retrieval, open-vocabulary classification, and dense visual tasks. Project page is available to further promote the prosperity of VLMs: https://voide1220.github.io/Holism/. Haicheng Wang, Chen Ju, Weixiong Lin, Shuai Xiao 0002, Mingshuai Yao, Jinsong Lan, Ying Chen 0011, Qingwen Liu 0002 |
CVPR | 2 |
| 2025 | FreeSegDiff: Annotation-free Saliency Segmentation with Diffusion ModelsabstractLearning from a large corpus of data, pre-trained models have achieved impressive progress nowadays. As a popular generative pre-training method, diffusion models stand out by capturing both low-level visual knowledge and high-level semantic relations. In this paper, we propose to exploit such knowledgeable pre-trained diffusion models for mainstream discriminative tasks such as annotation-free saliency segmentation. However, a notable structural discrepancy between generative and discriminative models poses a significant challenge to diffusion models’ direct application. Furthermore, the absence of explicit manually labeled data is a substantial barrier in annotation-free settings. To tackle these issues, we introduce FreeSegDiff, one novel synthesis-exploitation framework containing two-stage strategies. In the first synthesis stage, to alleviate data insufficiency, we synthesize abundant images, and propose a novel training-free DiffusionCut to produce masks. In the second exploitation stage, to bridge the structural gap, we employ the inversion technique to convert given images back to diffusion features. These features seamlessly integrate with downstream architectures. Extensive experiments and ablation studies demonstrate the superiority of adapting diffusion for annotation-free saliency segmentation. Chaofan Ma, Yuhuan Yang, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001 |
ICASSP | 3 |
| 2025 | Contrast-Unity for Partially-Supervised Temporal Sentence GroundingabstractTemporal sentence grounding aims to detect event timestamps described by the natural language query from given untrimmed videos. The existing fully-supervised setting achieves great results but requires expensive annotation costs; while the weakly-supervised setting adopts cheap labels but performs poorly. To pursue high performance with less annotation costs, this paper introduces an intermediate partially-supervised setting, i.e., only short-clip is available during training. To make full use of partial labels, we specially design one contrast-unity framework, with the two-stage goal of implicit-explicit progressive grounding. In the implicit stage, we align event-query representations at fine granularity using comprehensive quadruple contrastive learning: event-query gather, event-background separation, intra-cluster compactness and inter-cluster separability. Then, high-quality representations bring acceptable grounding pseudo-labels. In the explicit stage, to explicitly optimize grounding objectives, we train one fully-supervised model using obtained pseudo-labels for grounding refinement and denoising. Extensive experiments and thoroughly ablations on Charades-STA and ActivityNet Captions demonstrate the significance of partial supervision, as well as our superior performance. Haicheng Wang, Chen Ju, Weixiong Lin, Chaofan Ma, Ya Zhang 0002, Yanfeng Wang 0001 |
ICASSP | 2 |
| 2025 | FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformanceabstractRecently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their abilities to generate and understand cross-modal data. However, processing long sequences of visual tokens extracted from visual backbones poses a challenge for deployment in real-time applications. To address this issue, we introduce FOLDER, a simple yet effective plug-and-play module designed to reduce the length of the visual token sequence, mitigating both computational and memory demands during training and inference. Through a comprehensive analysis of the token reduction process, we analyze the information loss introduced by different reduction strategies and develop FOLDER to preserve key information while removing visual redundancy. We showcase the effectiveness of FOLDER by integrating it into the visual backbone of several MLLMs, significantly accelerating the inference phase. Furthermore, we evaluate its utility as a training accelerator or even performance booster for MLLMs. In both contexts, FOLDER achieves comparable or even better performance than the original models, while dramatically reducing complexity by removing up to 70% of visual tokens. Haicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju, Victor Quétu, Shuai Xiao 0002, Enzo Tartaglione |
ICCV | 4 |
| 2024 | Audio-Visual Segmentation via Unlabeled Frame ExploitationabstractAudio-visual segmentation (AVS) aims to segment the sounding objects in video frames. Although great progress has been witnessed, we experimentally reveal that current methods reach marginal performance gain within the use of the unlabeled frames, leading to the underutilization issue. To fully explore the potential of the unlabeled frames for AVS, we explicitly divide them into two categories based on their temporal characteristics, i.e., neighboring frame (NF) and distantframe (DF). NFs, temporally adjacent to the labeled frame, often contain rich motion information that assists in the accurate localization of sounding objects. Contrary to NFs, DFs have long temporal distaaces from the labeled frame, which share semantic-similar objects with appearance variations. Considering their unique characteristics, we propose a versatile framework that effectively leverages them to tackle AVS. Specifically, for NFs, we exploit the motion cues as the dynamic guidance to improve the objectness localization. Besides, we exploit the semantic cues in DFs by treating them as valid augmentations to the labeled frames, which are then used to enrich data diversity in a self-training manner. Extensive experimental results demonstrate the versatility and superiority of our method, unleashing the power of the abundant unlabeled frames. Jinxiang Liu, Fei Zhang 0016, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001 |
CVPR | 4 |
| 2024 | Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment
Zhonghua Zhai, Chen Ju, Xuewen Hong, Jinsong Lan |
ECCV (14) | 4 |
| 2024 | Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
Chen Ju, Haicheng Wang, Haozhe Cheng, Xu Chen 0026, Zhonghua Zhai, Jinsong Lan, Shuai Xiao 0002, Bo Zheng 0007 |
ECCV (46) | 1 |
| 2024 | Annotation-free Audio-Visual SegmentationabstractThe objective of Audio-Visual Segmentation (AVS) is to localise the sounding objects within visual scenes by accurately predicting pixel-wise segmentation masks. To tackle the task, it involves a comprehensive consideration of both the data and model aspects. In this paper, first, we initiate a novel pipeline for generating artificial data for the AVS task without extra manual annotations. We leverage existing image segmentation and audio datasets and match the image-mask pairs with its corresponding audio samples using category labels in segmentation datasets, that allows us to effortlessly compose (image, audio, mask) triplets for training AVS models. The pipeline is annotation-free and scalable to cover a large number of categories. Additionally, we introduce a lightweight model SAMA-AVS which adapts the pre-trained segment anything model (SAM) to the AVS task. By introducing only a small number of trainable parameters with adapters, the proposed model can effectively achieve adequate audio-visual fusion and interaction in the encoding stage with vast majority of parameters fixed. We conduct extensive experiments, and the results show our proposed model remarkably surpasses other competing methods. Moreover, by using the proposed model pretrained with our synthetic data, the performance on real AVSBench data is further improved, achieving 83.17 mIoU on S4 subset and 66.95 mIoU on MS3 set. The project page is https://jinxiang-liu.github.io/anno-free-AVS/. Jinxiang Liu, Yu Wang 0027, Chen Ju, Chaofan Ma, Ya Zhang 0002, Weidi Xie |
WACV | 3 |
| 2024 | Multi-modal Prototypes for Open-World Semantic Segmentation
Yuhuan Yang, Chaofan Ma, Chen Ju, Fei Zhang 0016, Jiangchao Yao, Ya Zhang 0002, Yanfeng Wang 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Distilling Vision-Language Pre-Training to Collaborate with Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization (WTAL) learns to detect and classify action instances with only category labels. Most methods widely adopt the off-the-shelf Classification-Based Pre-training (CBP) to generate video features for action localization. However, the different optimization objectives between classification and localization, make temporally localized results suffer from the serious incomplete issue. To tackle this issue without additional annotations, this paper considers to distill free action knowledge from Vision-Language Pre-training (VLP), as we surprisingly observe that the localization results of vanilla VLP have an over-complete issue, which is just complementary to the CBP results. To fuse such complementarity, we propose a novel distillation-collaboration framework with two branches acting as CBP and VLP respectively. The framework is optimized through a dual-branch alternate training strategy. Specifically, during the B step, we distill the confident background pseudo-labels from the CBP branch; while during the F step, the confident foreground pseudo-labels are distilled from the VLP branch. As a result, the dualbranch complementarity is effectively fused to promote one strong alliance. Extensive experiments and ablation studies on THUMOS14 and ActivityNet1.2 reveal that our method significantly outperforms state-of-the-art methods. Chen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao, Ya Zhang 0002, Jianlong Chang, Qi Tian 0001, Yanfeng Wang 0001 |
CVPR | 1 |
| 2023 | Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregation
Chaofan Ma, Yuhuan Yang, Chen Ju, Fei Zhang 0016, Ya Zhang 0002, Yanfeng Wang 0001 |
NeurIPS | 3 |
| 2023 | Adaptive Mutual Supervision for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization aims to localize actions from untrimmed long videos with only video-level category labels. Most previous methods ignore the incompleteness issue of Class Activation Sequences (CAS), suffering from trivial detection results. To tackle this issue, we propose a novel Adaptive Mutual Supervision (AMS) framework with two branches, where the base branch detects the most discriminative action regions, while the supplementary branch localizes the less discriminative action regions through an adaptive sampler. The sampler dynamically updates the inputs for the supplementary branch using a sampling weight sequence negatively correlated with the CAS from the base branch, thus encouraging the supplementary branch to localize the action regions underestimated by the base branch. To promote mutual enhancement between two branches, we further construct mutual location supervision. Each branch adopts the location pseudo-labels generated from the other branch as the localization supervision. By alternately optimizing two branches for multiple iterations, we progressively complete action regions. Extensive experiments on THUMOS14 and ActivityNet1.2 demonstrate that the proposed AMS method significantly outperforms state-of-the-art methods. Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang 0002, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Prompting Visual-Language Models for Efficient Video Understanding
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang 0002, Weidi Xie |
ECCV (35) | 1 |
| 2022 | Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationabstractWe present a simple yet effective self-supervised framework for audio-visual representation learning, to localize the sound source in videos. To understand what enables to learn useful representations, we systematically investigate the effects of data augmentations, and reveal that (1) composition of data augmentations plays a critical role, i.e. explicitly encouraging the audio-visual representations to be invariant to various transformations (transformation invariance); (2) enforcing geometric consistency substantially improves the quality of learned representations, i.e. the detected sound source should follow the same transformation applied on input video frames (transformation equivariance). Extensive experiments demonstrate that our model significantly outperforms previous methods on two sound localization benchmarks, namely, Flickr-SoundNet and VGG-Sound. Additionally, we also evaluate audio retrieval and cross-modal retrieval tasks. In both cases, our self-supervised models demonstrate superior retrieval performances, even competitive with the supervised approach in audio retrieval. This reveals the proposed framework learns strong multi-modal representations that are beneficial to sound localisation and generalization to further applications. The project page is https://jinxiang-liu.github.io/SSL-TIE. Jinxiang Liu, Chen Ju, Weidi Xie, Ya Zhang 0002 |
ACM Multimedia | 2 |
| 2022 | MePark: Using Meters as Sensors for Citywide On-Street Parking Availability PredictionabstractReal-time parking availability prediction is of great value to optimize the on-street parking resource utilization and improve traffic conditions, while the expensive costs of the existing parking availability sensing systems have limited their large-scale applications in more cities and areas. This paper presents the MePark system to predict real-time citywide on-street parking availability at fine-grained temporal level based on the readily accessible parking meter transactions data and other context data, together with the parking events data reported from a limited number of specially deployed sensors. We design an iterative mechanism to effectively integrate the aggregated inflow prediction and individual parking duration prediction for adequately exploiting the transactions data. Meanwhile, we extract discriminative features from the multi-source data, and combine the multiple-graph convolutional neural network (MGCN) and the long short-term memory (LSTM) network for capturing complex spatio-temporal correlations. The extensive experimental results based on a four-month real-world on-street parking dataset in Shenzhen, China demonstrate the advantages of our approach over various baselines. Dong Zhao 0001, Chen Ju, Guanzhou Zhu, Desheng Zhang 0002, Huadong Ma |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Divide and Conquer for Single-frame Temporal Action LocalizationabstractSingle-frame temporal action localization (STAL) aims to localize actions in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the one-stage framework but couple the counting goal and the localization goal. This paper proposes a novel two-stage framework for the STAL task with the spirit of divide and conquer. The instance counting stage leverages the location supervision to determine the number of action instances and divide a whole video into multiple video clips, so that each video clip contains only one complete action instance; and the location estimation stage leverages the category supervision to localize the action instance in each video clip. To efficiently represent the action instance in each video clip, we introduce the proposal-based representation, and design a novel differentiable mask generator to enable the end-to-end training supervised by category labels. On THUMOS14, GTEA, and BEOID datasets, our method outperforms state-of-the-art methods by 3.5%, 2.7%, 4.8% mAP on average. And extensive experiments verify the effectiveness of our method. Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001 |
ICCV | 1 |
| 2020 | Bottom-Up Temporal Action Localization with Mutual Regularization
Peisen Zhao, Lingxi Xie, Chen Ju, Ya Zhang 0002, Yanfeng Wang 0001, Qi Tian 0001 |
ECCV (8) | 3 |