EDBT 2026 Demo / reviewers in the wild / expert
Junliang Chen 0002
dblp:63/6051-2
· DBLP profile ↗
21ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-7516-9546ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FineXtrol: Controllable Motion Generation via Fine-Grained TextabstractRecent works have sought to enhance the controllability and precision of text-driven motion generation. Some approaches leverage large language models (LLMs) to produce more detailed texts, while others incorporate global 3D coordinate sequences as additional control signals. However, the former often introduces misaligned details and lacks explicit temporal cues, and the latter incurs significant computational cost when converting coordinates to standard motion representations. To address these issues, we propose FineXtrol, a novel control framework for efficient motion generation guided by temporally-aware, precise, user-friendly and fine-grained textual control signals that describe specific body part movements over time. In support of this framework, we design a hierarchical contrastive learning module that encourages the text encoder to produce more discriminative embeddings for our novel control signals, thereby improving motion controllability. Quantitative results show that FineXtrol achieves strong performance in controllable motion generation, while qualitative analysis demonstrates its flexibility in directing specific body part movements. Keming Shen, Bizhu Wu, Junliang Chen 0002, LinLin Shen |
AAAI | 3 |
| 2026 | Ranking-Based Self-Supervised Representation Learning for Skeleton-Based Action RecognitionabstractRecently, researchers have achieved significant results in the skeleton-based action recognition. To better model the skeleton sequences, we drive the encoder to learn more discriminative representations in the self-supervised setting. We find that instead of clustering feature vectors to assign pseudo labels for samples as in DeepCluster, ranking them is a more reasonable, reliable, and efficient way to learn more effective feature representations. With this intuition, we propose a novel self-supervised learning framework,DeepRank. Specifically, we rank triplets of skeleton sequences with the ranking labels, obtained from the relative distances among them. Besides, to deeply mine complementary discriminative information that exists in different modalities of skeleton sequences, we further proposeMulti-ViewDeepRank(MV-DeepRank) to enable encoders to comprehensively learn complementary features from multiple modalities. Extensive experimental results on the NTU RGB+D, NTU RGB+D 120, PKU-MMD I, and PKU-MMD II datasets under various evaluation settings demonstrate the generality, transferability, and superiority of our proposed self-supervised learning frameworks. Notably, our frameworks surpass the previous methods that employ the same backbone networks as ours by at least 1.8% (ST-GCN) and 2.1% (STTFormer) under the finetuning setting. Additionally, DeepRank gains a significant advantage on computational complexities,$O(1)$, over the contrastive learning-based methods,$O(\rm{batch size})$, and the clustering-based methods,$O(\rm{number of clusters})$. Bizhu Wu, Junliang Chen 0002, Jinheng Xie, Qiufu Li, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen |
IEEE Trans. Multim. | 2 |
| 2025 | FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMsabstractMultimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we introduce FaceBench, a dataset featuring hierarchical multi-view and multi-level attributes specifically designed to assess the comprehensive face perception abilities of MLLMs. Initially, we construct a hierarchical facial attribute structure, which encompasses five views with up to three levels of attributes, totaling over 210 attributes and 700 attribute values. Based on the structure, the proposed FaceBench consists of 49,919 visual questionanswering (VQA) pairs for evaluation and 23,841 pairs for fine-tuning. Moreover, we further develop a robust face perception MLLM baseline, Face-LLaVA, by training with our proposed face VQA data. Extensive experiments on various mainstream MLLMs and Face-LLaVA are conducted to test their face perception ability, with results also compared against human performance. The results reveal that, the existing MLLMs are far from satisfactory in understanding the fine-grained facial attributes, while our Face-LLaVA significantly outperforms existing open-source models with a small amount of training data and is comparable to commercial ones like GPT-4o and Gemini. The dataset will be released at https://github.com/CVI-SZU/FaceBench Xusen Ma, Xianxu Hou, Meidan Ding, Yudong Li 0001, Junliang Chen 0002, Wenting Chen, Xiaoyang Peng, LinLin Shen |
CVPR | 6 |
| 2025 | DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face ParsingabstractFace parsing aims to segment facial images into key components such as eyes, lips, and eyebrows. While existing methods rely on dense pixel-level annotations, such annotations are expensive and labor-intensive to obtain. To reduce annotation cost, we introduce Weakly Supervised Face Parsing (WSFP), a new task setting that performs dense facial component segmentation using only weak supervision, such as image-level labels and natural language descriptions. WSFP introduces unique challenges due to the high co-occurrence and visual similarity of facial components, which lead to ambiguous activations and degraded parsing performance. To address this, we propose DisFaceRep, a representation disentanglement framework designed to separate co-occurring facial components through both explicit and implicit mechanisms. Specifically, we introduce a co-occurring component disentanglement strategy to explicitly reduce dataset-level bias, and a text-guided component disentanglement loss to guide component separation using language supervision implicitly. Extensive experiments on CelebAMask-HQ, LaPa, and Helen demonstrate the difficulty of WSFP and the effectiveness of DisFaceRep, which significantly outperforms existing weakly supervised semantic segmentation methods. The code will be released at https://github.com/CVI-SZU/DisFaceRep. Xianxu Hou, Meidan Ding, Junliang Chen 0002, Kaijun Deng, Jinheng Xie, LinLin Shen |
ACM Multimedia | 4 |
| 2024 | Hygiea+: Toward Energy-Efficient and Highly Accurate Toothbrushing Monitoring via Wrist-Worn Gesture SensingabstractProper and effective toothbrushing technique is crucial for maintaining oral health. However, there are often limited opportunities for individuals to receive specific training in toothbrushing posture in their daily lives. In this article, we propose Hygiea+, a convenient, energy-efficient, and highly accurate toothbrushing monitoring system based on wrist-worn wearables. By leveraging inertial measurement units (IMUs) in wrist-worn devices for gesture sensing, Hygiea+ enables users to accurately and efficiently monitor their toothbrushing activities without any modifications to the toothbrush. We propose a number of novel techniques to achieve the goal of high sensing accuracy and energy efficiency. To reduce the energy consumption of continuous IMU sampling, we model the sensing problem as a Markov process and design a partially observable Markov decision process (POMDP)-based adaptive sampling strategy to dynamically adjust the sampling frequency. To achieve high sensing accuracy, we first propose a novel signal preprocessing method to mitigate variations resulting from different toothbrush types and user habits. Then, we propose a deep reinforcement learning-based data distillation mechanism to extract key segments from continuous toothbrushing actions, thus reducing the impact of redundant data and noise. In the classification stage, we design an attention-based long short-term memory (AT-LSTM) network for fine-grained toothbrushing posture recognition. In addition, to address the accuracy degradation of new users, we adopt the common but effective fine-tuning method to alleviate the data collection burden on new users. Finally, we connect advanced large language models (LLMs) to provide users with necessary feedback on toothbrushing behavior and health recommendations. Extensive experiments using both manual and electric toothbrushes demonstrate Hygiea+ achieves up to 98.8% accuracy in toothbrushing posture recognition while maintaining superior energy efficiency. Xingyu Feng 0001, Chengwen Luo 0001, Junliang Chen 0002, Jianqiang Li 0001, Zahir Tari, Weitao Xu |
IEEE Internet Things J. | 3 |
| 2024 | Enhancing Unsupervised Semantic Segmentation Through Context-Aware ClusteringabstractDespite the great progress of semantic segmentation with supervised learning, annotating large amounts of pixel-wise labels is, however, very expensive and time-consuming. To this end, Unsupervised Semantic Segmentation(USS) has been proposed to learn semantic segmentation, without any form of annotations. This approach involves dense prediction of semantics which is however challenging due to the unreliable nature of local representations. To solve this problem, we propose a newly context-aware unsupervised semantic segmentation framework, which aims to enhance the unsupervised semantic segmentation by leveraging contextual knowledge within and across images. In particular, we introduce a training strategy based on our Pyramid Semantic Guidance (PSG), which utilizes holistic semantics on pyramid views to guide pixel clustering with a siamese network-based framework. Additionally, we introduce a Context-Aware Embedding (CAE) module to fuse global features with low-level geometrical and appearance representations. We evaluate our method on the COCO-Stuff dataset and achieved competitive results compared to both the convolutional and ViT-based USS methods. Specifically, we attain significant improvements of +4.5% and +5% mIoU for Stuff and all class segmentation respectively, compared to previous approaches that employ unsupervised convolutional backbones. Yuan Wang 0083, Junliang Chen 0002, Songhe Deng, Zhi Wang 0001, LinLin Shen, Wenwu Zhu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | ADATS: Adaptive RoI-Align based Transformer for End-to-End Text SpottingabstractScene text spotting has attracted great attention in recent years. Compared with two-stage approaches that locate scene texts in the first stage and recognize them in the second stage, the advantages of joint location and recognition training are not fully explored. In this paper, we present an ADaptive RoI-Align based transformer for end-to-end Text Spotting (ADATS), which simultaneously locates and recognizes text with a single forward pass. By employing an Adaptive RoI-Align, the text features are extracted from the feature extraction network with the original aspect ratio, such that less information is lost during the alignment of arbitrarily-shaped scene text. Attention-based segmentation and recognition heads allow us to simultaneously optimize detection and recognition. Experiments on ICDAR 2015, MSRA-TD500, Total-Text, and CTW1500 demonstrate the effectiveness of our method. Zepeng Huang, Qi Wan, Junliang Chen 0002, Kai Ye 0004, LinLin Shen |
ICME | 3 |
| 2023 | JLInst: Boundary-Mask Joint Learning for Instance Segmentation
Junliang Chen 0002, Zepeng Huang, LinLin Shen |
PRCV (12) | 2 |
| 2023 | Adversarial Learning of Object-Aware Activation Map for Weakly-Supervised Semantic SegmentationabstractRecent years have witnessed impressive advances in the area of weakly-supervised semantic segmentation (WSSS). However, most of existing approaches are based on class activation maps (CAMs), which suffer from the under-segmentation problem (i.e., objects of interest are segmented partially). Although a number of literature works have been proposed to tackle this under-segmentation problem, we argue that these solutions built on CAMs may not be optimal for the WSSS task. Instead, in this paper we propose a network based on the object-aware activation map (OAM). The proposed network, termed OAM-Net, consists of four loss functions (foreground loss, background loss, average pixel and consistency loss) which ensure exactness, completeness, compactness and consistency of segmented objects via adversarial training. Compared to conventional CAM-based methods, our OAM-Net overcomes the under-segmentation drawback and significantly improves segmentation accuracy with negligible computational cost. A thorough comparison between OAM-Net and CAM-based approaches is carried out on the PASCAL VOC2012 dataset, and experimental results show that our network outperforms state-of-the-art approaches by a large margin. The code will be available soon. Junliang Chen 0002, Weizeng Lu, Yuexiang Li, LinLin Shen, Jinming Duan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | C2 AM: Contrastive learning of Class-agnostic Activation Map for Weakly Supervised Object Localization and Semantic SegmentationabstractWhile class activation map (CAM) generated by image classification network has been widely used for weakly su-pervised object localization (WSOL) and semantic segmentation (WSSS), such classifiers usually focus on discriminative object regions. In this paper, we propose Contrastive learning for Class-agnostic Activation Map (C2AM) generation only using unlabeled image data, without the involvement of image-level supervision. The core idea comes from the observation that i) semantic information of fore-ground objects usually differs from their backgrounds; ii) foreground objects with similar appearance or background with similar color/texture have similar representations in the feature space. We form the positive and negative pairs based on the above relations and force the network to disentangle foreground and background with a class-agnostic activation map using a novel contrastive loss. As the network is guided to discriminate cross-image foreground-background, the class-agnostic activation maps learned by our approach generate more complete object regions. We successfully extracted from C2AM class-agnostic object bounding boxes for object localization and background cues to refine CAM generated by classification network for semantic segmentation. Extensive experiments on CUB-200-2011, ImageNet-1K, and PASCAL VOC2012 datasets show that both WSOL and WSSS can benefit from the proposed C2AM. Code will be available at https://github.com/CVI-SZUICCAM. Jinheng Xie, Jianfeng Xiang, Junliang Chen 0002, Xianxu Hou, LinLin Shen |
CVPR | 3 |
| 2022 | RamGAN: Region Attentive Morphing GAN for Region-Level Makeup Transfer
Jianfeng Xiang, Junliang Chen 0002, Wenshuang Liu, Xianxu Hou, LinLin Shen |
ECCV (22) | 2 |
| 2022 | Sample Hardness Based Gradient Loss for Long-Tailed Cervical Cell Detection
Minmin Liu, Xuechen Li 0001, Xiangbo Gao, Junliang Chen 0002, LinLin Shen, Huisi Wu |
MICCAI (2) | 4 |
| 2021 | Selective Multi-scale Learning for Object Detection
Junliang Chen 0002, Weizeng Lu, LinLin Shen |
ICANN (2) | 1 |
| 2021 | Multi-scale Attention-Based Feature Pyramid Networks for Object Detection
Junliang Chen 0002, Minmin Liu, Kai Ye 0004, LinLin Shen |
ICIG (1) | 2 |
| 2021 | Delving into the Scale Variance Problem in Object DetectionabstractObject detection has made substantial progress in the last decade, due to the capability of convolution in extracting local context of objects. However, the scales of objects are diverse and current convolution can only process single-scale input. The capability of traditional convolution with a fixed receptive field in dealing with such a scale variance problem, is thus limited. Multiscale feature representation has been proven to be an effective way to mitigate the scale variance problem. Recent researches mainly adopt partial connection with certain scales, or aggregate features from all scales and focus on the global information across the scales. However, the information across spatial and depth dimensions is ignored. Inspired by this, we propose the multi-scale convolution (MSConv) to handle this problem. Taking into consideration scale, spatial and depth information at the same time, MSConv is able to process multi-scale input more comprehensively. MSConv is effective and computationally efficient, with only a small increase of computational cost. For most of the single-stage object detectors, replacing the traditional convolutions with MSConvs in the detection head can bring more than 2.5% improvement in AP (on COCO 2017 dataset), with only 3% increase of FLOPs. MSConv is also flexible and effective for two-stage object detectors. When extended to the mainstream two-stage object detectors, MSConv can bring up to 3.0% improvement in AP. Our best model under single-scale testing achieves 48.9% AP on COCO 2017 test-dev split, which surpasses many state-of-the-art methods. Junliang Chen 0002, LinLin Shen |
ICTAI | 1 |
| 2020 | Feature map masking based single-stage face detectionabstractAlthough great progress has been made in face detection, a trade-off between speed and accuracy is still a great challenge. We propose in this paper a feature map masking based approach for single-stage face detection. As feature maps extracted from feature pyramid network might contain face unrelated features, we propose a mask generation branch to predict those significant units for face detection. The masked feature maps, where only important features are left, are then passed through the following detection process. Ground truth masks, directly generated from the training images, based on the face bounding boxes, are used to train the feature mask generation module. A mask constrained dropout module has also been proposed to drop out significant units of the shared feature maps, such that the detection performance can be further improved. The proposed approach is extensively tested using the WIDER FACE dataset. The results suggest that our detector with ResNet-152 backbone, achieves the best precision-recall performance among competing methods. As high as 95.4%, 94.0% and 86.9% accuracies have been achieved on the easy, medium and hard subsets, respectively. Junliang Chen 0002, Weicheng Xie 0001, LinLin Shen |
IJCB | 2 |
| 2019 | Brush like a Dentist: Accurate Monitoring of Toothbrushing via Wrist-Worn Gesture SensingabstractOral health has significant impact on people’s over-all well-being. While many activity recognition systems exist in the literature, accurately sensing toothbrushing activities remains an unsolved challenging problem due to the diversity of tooth-brushing habits among different users and subtle distinctions between different brushing actions. In this work, we propose Hygiea, an energy-efficient and highly-accurate toothbrushing monitoring system which exploits IMU-based wrist-worn gesture sensing using unmodified toothbrushes. To address toothbrushing variety, Hygiea incorporates a number of novel signal preprocessing techniques to automatically transform the sensory input during arbitrary toothbrushing activities to the consistent user coordinate system. To distinguish different brushing actions, Hygiea leverages an emerging deep learning model (e.g., AT-LSTM) to achieve fine-grained activity recognitions. Moreover, a POMDP model is incorporated for sampling control to balance activity detection and energy efficiency. Extensive real-world experiments show that the Hygiea system achieves a 11.7% accuracy gain compared to the state-of-the-art while maintaining energy-efficiency and zero modification on the toothbrushes. Chengwen Luo 0001, Xingyu Feng 0001, Junliang Chen 0002, Jianqiang Li 0001, Weitao Xu, Wei Li 0058, Zahir Tari, Albert Y. Zomaya |
INFOCOM | 3 |
| 2013 | Throughput Enhancement through Selective Time Sharing and Dynamic GroupingabstractSpace sharing approaches are widely used in job scheduling for HPC systems. The main drawback of these approaches is the blocking of short jobs, which results in low throughput. The research on gang scheduling has shown the potential of time sharing in improving throughput. However, traditional gang scheduling adds jobs for time sharing without selection, which may cause a higher performance degradation of existing running jobs than the performance gain of waiting jobs. Moreover, gang scheduling often adopts a contiguous buddy allocation scheme which has problems of fragmentation and low resource utilization. We design a selective time sharing technique that allows waiting jobs to be co-scheduled with existing running jobs only if the overall throughput can be improved. To alleviate the fragmentation problem, we present a dynamic grouping resource allocation mechanism that relaxes the contiguous allocation requirement imposed on gang scheduling. By integrating these techniques, our new job co-scheduling algorithm is able to simultaneously take system throughput and resource utilization into consideration. The experimental results demonstrate that our approach significantly outperforms both EASY backfilling and traditional gang scheduling in terms of both average turnaround time and bounded slowdown. Junliang Chen 0002, Bing Bing Zhou, Chen Wang 0008, Peng Lu 0004, Penghao Wang 0002, Albert Y. Zomaya |
IPDPS | 1 |
| 2012 | Workload Characteristic Oriented Scheduler for MapReduceabstractApplications in many areas are increasingly developed and ported using the Map Reduce framework (more specifically, Hadoop) to exploit (data) parallelism. The application scope of Map Reduce has been extended beyond the original design goal which was large-scale data processing. This extension inherently makes a need for scheduler to explicitly take into account characteristics of job for two main goals of efficient resource use and performance improvement. In this paper, we study Map Reduce scheduling strategies to effectively deal with different workload characteristics CPU intensive and I/O intensive. We present the Workload Characteristic Oriented Scheduler (WCO), which strives for co-locating tasks of possibly different Map Reduce jobs with complementing resource usage characteristics. WCO is characterized by its essentially dynamic and adaptive scheduling decisions using information obtained from its characteristic estimator. Workload characteristics of tasks are primarily estimated by sampling with the help of some static task selection strategies, e.g., Java byte code analysis. Results obtained from extensive experiments using 11 benchmarks in a 4-node local cluster and a 51-node Amazon EC2 cluster show 17% performance improvement on average in terms of throughput in the situation of co-existing diverse workloads. Peng Lu 0004, Young Choon Lee, Chen Wang 0008, Bing Bing Zhou, Junliang Chen 0002, Albert Y. Zomaya |
ICPADS | 5 |
| 2012 | Just Satisfactory Resource Provisioning for Parallel Applications in the CloudabstractThe paper discusses a resource management method at the cloud tenant level. It concerns efficiently running a profitable service on resources leased from a cloud infrastructure provider. Particularly, the paper focuses on services that handle parallelizable client requests, e.g, such a request can be processed using a MPI or a MapReduce program. The objective is to manage resource use to just satisfy the performance requirements of clients and avoid the common over-provisioning or under-provisioning problem. The proposed resource management method makes initial resource leasing plan based on client request profile and performance targets in service level agreements (SLAs). It then dynamically adjusts resource allocation based on monitoring data. Our extensive experiments show that the method is able to provision and efficiently use resources to just satisfy the SLA targets. Chen Wang 0008, Junliang Chen 0002, Bing Bing Zhou, Albert Y. Zomaya |
SERVICES | 2 |
| 2011 | Profiling Applications for Virtual Machine Placement in CloudsabstractApplication profiling is an important technique for efficient resource management. The decision making of scheduling and resource allocation typically takes great advantage of such a technique primarily for improving resource utilization. With the advent of cloud computing as a multitenant virtualized platform, diverse applications are increasingly deployed onto the cloud and they more than often share physical resources. The background load (other applications running on the same physical machine) is therefore an important factor for profiling an application in this cloud computing scenario. In this paper, we present a novel application profiling technique using the canonical correlation analysis (CCA) method, which identifies the relationship between application performance and resource usage. We further devise a performance prediction model based on application profiles generated using CCA. Clearly, our profiling technique with this prediction model has a lot of potentials particularly in virtual machine (VM) placement with performance awareness. Our experimental results demonstrate the capability of our profiling technique and the accuracy of our prediction model. Anh Vu Do, Junliang Chen 0002, Chen Wang 0008, Young Choon Lee, Albert Y. Zomaya, Bing Bing Zhou |
IEEE CLOUD | 2 |