Bin Kang

dblp:16/10594 · DBLP profile ↗
← Back
49ranked-venue papers
17as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 16 · 6 first-author · 14 since 2021Computer networks · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Graph-aware multi-agent learning for cross-domain deterministic routing and scheduling
Jiming Yao, Yun Hong, Taojie Zhu, Bin Kang
Comput. Networks4
2026 Robust unsupervised visual tracking via image-to-video identity knowledge transferring
Bin Kang, Zongyu Wang, Dong Liang 0008, Tianyu Ding, Songlin Du
Pattern Recognit.1
2026 Parallel consensus transformer for local feature matching
Xiaoyong Lu, Bin Kang, Songlin Du
Pattern Recognit.3
2026 Semantic NeRF-Oriented 3-D Object Counting for Industrial Fruit Harvesting
abstract
Object counting is the core of perception components in the industry system. Nevertheless, real-world occlusion and dense clustering pose significant challenges to fruit harvesting. Existing methods closely rely on physics-agnostic 2D clustering with rigid thresholds, neglecting critical 3D geometric cues in complex scenes. Thus, they suffer from the multi-view double-counting issue. In this paper, we propose a 3D object counting framework that integrates semantic Neural Radiance Fields (NeRF) with a physics-guided adaptive clustering algorithm. In particular, we employ a semantic NeRF to achieve implicit 3D scene reconstruction, which effectively isolates target objects from the background. Based on the semantic NeRF, we introduce a physics-guided adaptive clustering algorithm that exploits physically interpretable features, including surface normals and elevation gradients, for accurate unsupervised point cloud segmentation. Subsequently, an energy function optimization mechanism is utilized to autonomously aggregate multi-dimensional features for dynamically adjusting clustering thresholds, which enables the 3D counting to adapt to diverse fruit morphologies without manual parameter tuning. Extensive experiments on both synthetic and real-world datasets demonstrate the superiority of the proposed framework, which improves counting accuracy by an average of 9.0 percentage points (6.9 pp on synthetic, 7.2 pp on real-world, and 12.9 pp on Fuji) over the state-of-the-art 3-D baseline.
Yimo Wang, Bin Kang, Jian Liu 0006, Changyin Sun 0001
IEEE Trans Autom. Sci. Eng.2
2026 DASS-Net: Degradation-Adaptive Semi-Symmetric Network for Robust Image Hiding
Junzhi Zhao, Hongjie He 0005, Fan Chen 0003, Lingfeng Qu, Bin Kang
IEEE Trans. Circuits Syst. Video Technol.5
2026 Towards General Cross-Modal Visual Coding for Emergency Communications
abstract
Multi-modal visual signals are prevalent in emergency communications. To ensure high reliability of signal transmission under bandwidth constraints, it is crucial to compress redundant information both within and between modalities as much as possible, and ensure the fidelity of the reconstructed signals. Most existing studies depend exclusively on single-modal coding schemes and fail to effectively leverage the semantic correlations between modalities. In this paper, we introduce an end-to-end general cross-modal visual coding scheme, namely CMVC, which aims to jointly compress multi-modal visual signals (such as visible and infrared signals). First, we propose a cross-modal asynchronous entropy module that extracts common features using a cross-attention mechanism. Additionally, we enhance the accuracy of common features extraction by maximizing mutual information loss. This module further compresses multi-modal visual signals by compressing only the residual features between modalities. Second, we propose a cascaded enhancement module based on cross-modal Mamba that fuses complementary information to enhance the reconstruction quality of multi-modal visual signals. Finally, extensive experimental results demonstrate that our scheme significantly outperforms other advanced methods on visible-infrared datasets. Even at low bitrates, multi-modal visual signals can still achieve excellent reconstruction quality. Additionally, our scheme exhibits outstanding compression and reconstruction performance when applied to visible-depth signals, effectively demonstrating its robustness and generalizability.
Lindong Zhao, Ang Li 0012, Bin Kang, Dan Wu 0001, Liang Zhou 0002
IEEE Trans. Multim.5
2026 Robust Fine-Grained Visual Categorization via Cyclical Attention
abstract
Fine-grained visual categorization (FGVC) in open-world settings frequently encounters heavy occlusion (HO) samples that compromise discriminative features. However, effectively addressing heavy occlusion remains a challenge. Existing methods often either discard the occluded parts or utilize them through additional techniques such as image inpainting or multimodel strategies, each with its own set of advantages and limitations. In this article, we propose a novel approach inspired by human self-regulated learning (SRL) behavior: cyclical attention that leverages occluded regions through the attention recalibration in the feedback loop. In particular, we introduce a new multi-instance model where occluded parts are essential due to a special feedback structure at the basis of a cooperative game mechanism. This mimics SRL to re-evaluate the previous attention-based image patch selection strategy. We then embed the proposed multi-instance model into a transformer architecture, creating an SRL-FGVC transformer. The key innovation of this design is the cyclical attention, with the forward and feedback self-attention formulating a cooperative union to mitigate attention bias. Extensive experiments on six public datasets and an additional dataset we established demonstrate that the SRL-FGVC transformer consistently outperforms existing approaches in HO scenarios. This work presents a promising new direction for robust FGVC in challenging real-world conditions.
Bin Kang, Dong Liang 0008, Daoyuan Chen, Tianyu Ding, Mingqiang Wei
IEEE Trans. Neural Networks Learn. Syst.1
2025 OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision
abstract
Open-vocabulary detection aims to detect objects from novel categories beyond the base categories on which the detector is trained. However, existing open-vocabulary detectors trained on base category data tend to assign higher confidence to trained categories and confuse novel categories with the background. To resolve this, we propose OV-DQUO, an Open-Vocabulary DETR with Denoising text Query training and open-world Unknown Objects supervision. Specifically, we introduce a wildcard matching method. This method enables the detector to learn from pairs of unknown objects recognized by the open-world detector and text embeddings with general semantics, mitigating the confidence bias between base and novel categories. Additionally, we propose a denoising text query training strategy. It synthesizes foreground and background query-box pairs from open-world unknown objects to train the detector through contrastive learning, enhancing its ability to distinguish novel objects from the background. We conducted extensive experiments on the OV-COCO and OV-LVIS benchmarks, achieving new state-of-the-art results of 45.6 AP50 and 39.3 mAP on novel categories, respectively.
Bin Chen 0022, Bin Kang, Yulin Li 0003, Weizhi Xian, Yichi Chen 0002
AAAI3
2025 DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception
abstract
Dense visual prediction tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have shown promise in open-vocabulary tasks, their direct application to dense prediction often leads to suboptimal performance due to limitations in local feature representation. In this work, we present our observation that CLIP’s image tokens struggle to effectively aggregate information from spatially or semantically related regions, resulting in features that lack local discriminability and spatial consistency. To address this issue, we propose DeCLIP, a novel framework that enhances CLIP by decoupling the self-attention module to obtain "content" and "context" features respectively. The "content" features are aligned with image crop representations to improve local discriminability, while "context" features learn to retain the spatial correlations under the guidance of vision foundation models, such as DINO. Extensive experiments demonstrate that DeCLIP significantly outperforms existing methods across multiple open-vocabulary dense prediction tasks, including object detection and semantic segmentation. Code is available at https://github.com/xiaomoguhz/DeCLIP.
Bin Chen 0022, Yulin Li 0003, Bin Kang, Yichi Chen 0002, Zhuotao Tian
CVPR4
2025 CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval
abstract
Existing Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggregation process and suppressing the discriminative features in text-driven image retrieval tasks. To address this, we introduce \textbf{CalibCLIP}, a training-free method designed to calibrate the suppressive effect of dominant tokens. Specifically, in the visual space, we propose the Contrastive Visual Enhancer (CVE), which decouples visual features into target and low information regions. Subsequently, it identifies dominant tokens and dynamically suppresses their representations.In the textual space, we introduce the Discriminative Concept Calibrator (DCC), which aims to differentiate between general and discriminative concepts within the text query. By mitigating the challenges posed by generic concepts and improving the representations of discriminative concepts, DCC strengthens the differentiation among similar samples. Finally, extensive experiments demonstrate consistent improvements across seven benchmarks spanning three image retrieval tasks, underscoring the effectiveness of CalibCLIP. Code is available at: https://github.com/kangbin98/CalibCLIP
Bin Kang, Bin Chen 0022, Yulin Li 0003, Junzhi Zhao, Junle Wang, Zhuotao Tian
ACM Multimedia1
2025 Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
abstract
Recent advances in Video Large Language Models (VLLMs) have achieved remarkable video understanding capabilities, yet face critical efficiency bottlenecks due to quadratic computational growth with lengthy visual token sequences of long videos. While existing keyframe sampling methods can improve temporal modeling efficiency, additional computational cost is introduced before feature encoding, and the binary frame selection paradigm is found suboptimal. Therefore, in this work, we propose **Dy**namic **To**ken compression via LLM-guided **K**eyframe prior (**DyToK**), a training-free paradigm that enables dynamic token compression by harnessing VLLMs' inherent attention mechanisms. Our analysis reveals that VLLM attention layers naturally encoding query-conditioned keyframe priors, by which DyToK dynamically adjusts per-frame token retention ratios, prioritizing semantically rich frames while suppressing redundancies. Extensive experiments demonstrate that DyToK achieves state-of-the-art efficiency-accuracy tradeoffs. DyToK shows plug-and-play compatibility with existing compression methods, such as VisionZip and FastV, attaining 2.5x faster inference while preserving accuracy across multiple VLLMs, such as LLaVA-OneVision and Qwen2.5-VL. Code and models will be made publicly available.
Yulin Li 0003, Haokun Gui, Ziyang Fan, Bin Kang, Bin Chen 0022, Zhuotao Tian
NeurIPS5
2025 Progressive Masking Oriented Self-Taught Learning for Occluded Facial Expression Recognition
abstract
Self-taught learning (STL) is a promising solution that reduces the performance gap between weakly supervised and fully supervised learning for easily accessible, label-free images. The success of traditional STL solutions relies on the assumption that the target appearance is completely visible and well-defined. In real-world facial expression recognition scenarios, however, saliency regions are often partially occluded, which significantly hampers the generalization capability of STL methods. Nevertheless, few studies have investigated the impact of occlusion on STL. In this paper, we propose an interweaved autoencoder network for weakly supervised facial expression recognition in occlusion scenarios. The key innovation of our network lies in the Residual Connection Union (RCU) blocks that can integrate the Convolutional Neural Network (CNN) and Transformer layers into a multi-scale structure. The RCU enables a progressive masking strategy to accurately identify and focus on contributive yet often overlooked image patches by analyzing the relationships among region-level target representations. In addition, we introduce a self-knowledge distillation module for the effective training of the proposed autoencoder network. Extensive experiments are conducted on four public datasets to demonstrate the superiority of our method over related works.
Bin Kang, Shuangshuang Wang, Zongyu Wang, Haie Dou, Lei Wang 0009, Zhijie Xia
IEEE Trans. Affect. Comput.1
2025 DPNet: Dual-Path Network for Real-Time Object Detection With Lightweight Attention
abstract
The recent advances in compressing high-accuracy convolutional neural networks (CNNs) have witnessed remarkable progress in real-time object detection. To accelerate detection speed, lightweight detectors always have few convolution layers using a single-path backbone. Single-path architecture, however, involves continuous pooling and downsampling operations, always resulting in coarse and inaccurate feature maps that are disadvantageous to locate objects. On the other hand, due to limited network capacity, recent lightweight networks are often weak in representing large-scale visual data. To address these problems, we present a dual-path network, named DPNet, with a lightweight attention scheme for real-time object detection. The dual-path architecture enables us to extract in parallel high-level semantic features and low-level object details. Although DPNet has a nearly duplicated shape with respect to single-path detectors, the computational costs and model size are not significantly increased. To enhance representation capability, a lightweight self-correlation module (LSCM) is designed to capture global interactions, with only a few computational overheads and network parameters. In the neck, LSCM is extended into a lightweight cross correlation module (LCCM), capturing mutual dependencies among neighboring scale features. We have conducted exhaustive experiments on MS COCO, Pascal VOC 2007, and ImageNet datasets. The experimental results demonstrate that DPNet achieves a state-of-the-art trade off between detection accuracy and implementation efficiency. More specifically, DPNet achieves 31.3% AP on MS COCO test-dev, 82.7% mAP on Pascal VOC 2007 test set, and 41.6% mAP on ImageNet validation set, together with nearly 2.5M model size, 1.04 GFLOPs, and 164 and 196 frames/s (FPS) FPS for input images of three datasets.
Quan Zhou 0004, Huimin Shi, Weikang Xiang, Bin Kang, Longin Jan Latecki
IEEE Trans. Neural Networks Learn. Syst.4
2024 Multi-Attribute Consistency Driven Visual Language Framework for Surface Defect Detection
abstract
Visual Language Pre-training models encounter significant challenges stemming from the scarcity of data and the presence of ambiguous cues in industrial defect detection tasks. In this work, we propose a multi-attribute consistency-driven defect detection (MACD) framework to optimize text prompts in a coarse-to-fine trajectory. To bridge differences in domain knowledge, we build a structured attribute repository that contains descriptions of various defects’ inherent attributes. Based on this, we propose a multi-attribute consistency (MAC) module that can adequately model the global alignment between sentences with multiple attributes and defect images. Furthermore, we design a refined cross-alignment (RCA) module to determine the fine-grained correspondence between each attribute and the region within the image. Finally, the proposed method is experimentally validated on two benchmarks, resulting in significant performance improvements in a wide range of defective scenarios.
Bin Kang, Bin Chen 0022, Weizhi Xian, Huifeng Chang
ICME1
2024 Fine-grained recognition via submodular optimization regulated progressive training
Bin Kang, Songlin Du, Dong Liang 0008, Xin Li 0086
Pattern Recognit.1
2024 Boundary-Guided Lightweight Semantic Segmentation With Multi-Scale Semantic Context
abstract
Lightweight semantic segmentation plays an essential role in image signal processing that is beneficial to many multimedia applications, such as self-driving, robotic vision, and virtual reality. Due to the powerful capability to encode image details and semantics, many lightweight dual-resolution networks have been proposed in recent years for semantic segmentation. In spite of achieving remarkable progresses, they often ignore semantic context ranged from different scales. Furthermore, most of them always neglect the object boundaries, serving as a significant assistance for lightweight semantic segmentation. To alleviate these problems, this paper develops a Boundary-guide dual-resolution lightweight network with multi-scale Semantic Context, called BSCNet, for semantic segmentation. Specifically, to enhance the capability of feature representation, an Extremely Lightweight Pyramid Pooling Module (ELPPM) is designed to capture multi-scale semantic context at the top of low-resolution branch of BSCNet. In addition, to increase feature similarity of the same object while keeping feature discrimination of different objects, pixel information is propagated throughout the entire object area using a simple Boundary Auxiliary Fusion Module (BAFM), where the predicted object boundaries are served as high-level guidance to refine low-level convolutional features. The comprehensive experimental results have demonstrated that our BSCNet is simple and effective, achieving state-of-the-art trade-off in terms of segmentation accuracy and running efficiency on CityScapes, CamVid, and KITTI datasets.
Quan Zhou 0004, Guangwei Gao, Bin Kang, Weihua Ou, Huimin Lu 0001
IEEE Trans. Multim.4
2023 ParaFormer: Parallel Attention Transformer for Efficient Feature Matching
abstract
Heavy computation is a bottleneck limiting deep-learning-based feature matching algorithms to be applied in many real-time applications. However, existing lightweight networks optimized for Euclidean data cannot address classical feature matching tasks, since sparse keypoint based descriptors are expected to be matched. This paper tackles this problem and proposes two concepts: 1) a novel parallel attention model entitled ParaFormer and 2) a graph based U-Net architecture with attentional pooling. First, ParaFormer fuses features and keypoint positions through the concept of amplitude and phase, and integrates self- and cross-attention in a parallel manner which achieves a win-win performance in terms of accuracy and efficiency. Second, with U-Net architecture and proposed attentional pooling, the ParaFormer-U variant significantly reduces computational complexity, and minimize performance loss caused by downsampling. Sufficient experiments on various applications, including homography estimation, pose estimation, and image matching, demonstrate that ParaFormer achieves state-of-the-art performance while maintaining high efficiency. The efficient ParaFormer-U variant achieves comparable performance with less than 50% FLOPs of the existing attention-based models.
Xiaoyong Lu, Yaping Yan, Bin Kang, Songlin Du
AAAI3
2023 CFFMixer: Multi-Dimensional Feature Fusion for Object Detection
abstract
Object detection is a fundamental task in the field of computer vision, and one of its essential requirements is high-quality feature fusion. Previous works have made various efforts in this regard: CNN-based detectors use convolutional blocks to fuse local features and dense prior knowledge to predict objects, while query-based detectors fuse global features by self-attention then decode features with object queries. However, their feature fusion methods are relatively monotonous. Considering that different modules are applicable to different dimensions, we proposed an object detector named CFFMixer which used hybrid architecture to achieve multi-dimensional feature fusion. The sampling strategy to extract abundant local and global features was first introduced then the Comprehensive Feature Fusion Network (CFFN) was proposed to integrate them. CFFN not only achieved local and global features interaction in the spatial dimension, but also fused semantics in the channel dimension. Furthermore, we conducted experiments and made a comparison with competitive models, our model finally got 43.0 mAP on COCO 2017 dataset within 12 epochs. Experimental results showed that the model’s accuracy benefits from the powerful feature fusion capability of CFFN. Besides, we performed ablation studies on our modules to evaluate their effectiveness.
Weizhe Yuan, Bin Kang, Songlin Du
ICASSP3
2023 Robust RGB-T Tracking via Consistency Regulated Scene Perception
abstract
RGB-T tracking has received increasing attention due to its significant advantage under severe weather conditions. Existing RGB-T tracking methods pay close attention to the representation of target appearance, ignoring the importance of scene information. In this paper, we propose a global reasoning-oriented method for RGB-T tracking. In particular, within a multi-task learning framework, our approach adopts a nested global reasoning model to regulate the consistency of scene perception (reasoning the relation between targets and the surrounding semantic regions) in different image domains. Moreover, a meta-unsupervised learning strategy is designed to enforce the nested global reasoning model to utilize partial multi-domain target information for the updating of scene perception. Extensive experiments on GTOT, RGBT210 and LasHeR datasets show the superior performance of our method when compared with related works.
Bin Kang, Songlin Du
ICIP1
2023 Global Information-Assisted Fine-Grained Visual Categorization in Internet of Things
abstract
In fine-grained visual categorization (FGVC), most part-based frameworks do not work effectively in some extremely challenging scenarios such as partial occlusion. This limitation is due to the heavy disorder of local features extracted from such occluded targets. To address this issue, we propose a global information-assisted network (GIAN), where auxiliary global information can search the useful elements of local information and integrate with them for an efficient unified feature representation. In particular, in order to acquire the global information, we design a global attention-concentrated convolutional neural network (GAC-CNN) by extending a convolutional neural network with a nonlocal GCN module. Then, the unified feature representation is produced by two strategies. On the one hand, a global–local aggregation strategy is developed to selectively integrate global features with local features through consistency evaluation and reweighting method. On the other hand, an alternative knowledge distillation strategy is developed to help generate more powerful global and local features. Two strategies collaboratively make the unified features more robust and more discriminative than traditional part-based features. Experimental results show that the proposed GIAN can achieve accuracies of 92.8%, 93.8%, and 95.7% on CUB-200-2011, FGVC Aircraft, and Stanford Cars, respectively.
Ang Li 0012, Bin Kang, Dan Wu 0001, Liang Zhou 0002
IEEE Internet Things J.2
2023 IVF-Net: An Infrared and Visible Data Fusion Deep Network for Traffic Object Enhancement in Intelligent Transportation Systems
abstract
Infrared and visible data fusion (IVF) aims to generate a fused output that simultaneously highlights salient thermal radiation features and preserves texture information, which can not only grasp the necessary information for traffic movement, but also highlight the invisible objects that need to be dodged in intelligent transportation system (ITS). Therefore, IVF is capable of improving the environmental perception ability for various challenging traffic situations, e.g., foggy scenarios, rainy environments, and low-light illumination. However, current available IVF algorithms cannot offer a theoretical manner to integrate a priori knowledge and the network structure into a unified model. Moreover, they always fail to handle infrared and visible data pairs with different resolutions, which is a common occurrence in real ITS scenarios. To this end, this study develops a novel model-inspired unsupervised network termed IVF-Net. Specifically, an enhanced IVF model (IVFM), which pays more attention on detailed texture information and salient objects, is first established. According to proximal gradient theory, then we map this model into a deep network with learnable feature extraction parameters, aiming to draw on the strengths of the fusion model and deep learning to better describe the IVF task. Finally, a multiple task-driven loss function is designed to train the mapped network. Unlike previous work, our IVF-Net is motivated by IVFM, each layer in which has a semantic interpretability and a clear mission, thereby leading to a significantly enhanced fusion effect. Another advantage is that it is only composed of simple convolution-based structures, which ensures its lightweight and efficiency. Experiments demonstrate that IVF-Net can have a stronger ability to capture the key traffic information and highlight the salient feature of imperceptible objects, which makes it an excellent candidate to improve the reliability of subsequent applications in ITS.
Mingye Ju, Chunming He, Juping Liu, Bin Kang, Jian Su 0001, Dengyin Zhang
IEEE Trans. Intell. Transp. Syst.4
2023 Robust RGB-T Tracking via Graph Attention-Based Bilinear Pooling
abstract
RGB-T tracker possesses strong capability of fusing two different yet complementary target observations, thus providing a promising solution to fulfill all-weather tracking in intelligent transportation systems. Existing convolutional neural network (CNN)-based RGB-T tracking methods often consider the multisource-oriented deep feature fusion from global viewpoint, but fail to yield satisfactory performance when the target pair only contains partially useful information. To solve this problem, we propose a four-stream oriented Siamese network (FS-Siamese) for RGB-T tracking. The key innovation of our network structure lies in that we formulate multidomain multilayer feature map fusion as a multiple graph learning problem, based on which we develop a graph attention-based bilinear pooling module to explore the partial feature interaction between the RGB and the thermal targets. This can effectively avoid uninformed image blocks disturbing feature embedding fusion. To enhance the efficiency of the proposed Siamese network structure, we propose to adopt meta-learning to incorporate category information in the updating of bilinear pooling results, which can online enforce the exemplar and current target appearance obtaining similar sematic representation. Extensive experiments on grayscale-thermal object tracking (GTOT) and RGBT234 datasets demonstrate that the proposed method outperforms the state-of-the-art methods for the task of RGB-T tracking.
Bin Kang, Dong Liang 0008, Junxi Mei, Xiaoyang Tan, Dengyin Zhang
IEEE Trans. Neural Networks Learn. Syst.1
2022 Progressive Training Enabled Fine-Grained Recognition
abstract
Organizing training samples in a meaningful order is beneficial for accelerating the convergence rate and enhancing the recognition performance in the CNN model. However, achieving reasonable sample ranking for fine-grained recognition datasets is very challenging because the intra and inter class relation in those datasets is opposite to that in public recognition datasets. In this paper, we propose a general framework for the progressive training of fine-grained recognition models. In particular, we first formulate the training subset selection as a group ranking-oriented submodular optimization problem, where the submodularity is adopted to evaluate the benefit of selected training subsets. This can give theoretical guidance for the consecutive discrimination of difficult and ordinary training subsets. Secondly, we design a training strategy to dynamically adjust the ratio of difficult and ordinary training subsets according to the recognition performance. Extensive experiments on CUB-200-2011 and Stanford Dogs datasets demonstrate that the proposed method outperforms the state-of-the-art curriculum learning methods.
Bin Kang
ICIP1
2022 PointTAD: Multi-Label Temporal Action Detection with Learnable Query Points
abstract
Traditional temporal action detection (TAD) usually handles untrimmed videos with small number of action instances from a single label (e.g., ActivityNet, THUMOS). However, this setting might be unrealistic as different classes of actions often co-occur in practice. In this paper, we focus on the task of multi-label temporal action detection that aims to localize all action instances from a multi-label untrimmed video. Multi-label TAD is more challenging as it requires for fine-grained class discrimination within a single video and precise localization of the co-occurring instances. To mitigate this issue, we extend the sparse query-based detection paradigm from the traditional TAD and propose the multi-label TAD framework of PointTAD. Specifically, our PointTAD introduces a small set of learnable query points to represent the important frames of each action instance. This point-based representation provides a flexible mechanism to localize the discriminative frames at boundaries and as well the important frames inside the action. Moreover, we perform the action decoding process with the Multi-level Interactive Module to capture both point-level and instance-level action semantics. Finally, our PointTAD employs an end-to-end trainable framework simply based on RGB input for easy deployment. We evaluate our proposed method on two popular benchmarks and introduce the new metric of detection-mAP for multi-label TAD. Our model outperforms all previous methods by a large margin under the detection-mAP metric, and also achieves promising results under the segmentation-mAP metric.
Jing Tan 0002, Xintian Shi, Bin Kang, Limin Wang 0002
NeurIPS4
2022 Contextual ensemble network for semantic segmentation
Quan Zhou 0004, Xiaofu Wu, Suofei Zhang, Bin Kang, ZongYuan Ge, Longin Jan Latecki
Pattern Recognit.4
2022 Exploring the Benefits of Cross-Modal Coding
abstract
Multi-modal services, typically integrating such signals as audio, video, and haptic, will become an inevitable application trend of the 5G and beyond. However, due to the essential differences among the haptic and audio/video signals, the existing coding schemes usually fail to satisfy the critical requirements in terms of the rate distortion performance. Inspired by the phenomenon that hearing, sight and touch are highly correlated, we provide an affirmative answer by proposing the framework of cross-modal coding, which compresses multi-modal signals aided by their semantic correlation. In particular, the highlights of this work lie in addressing three fundamental technical problems: i) how to exploit the semantic correlation among different modalities, ii) to what extent of benefit we can get from cross-modal coding, and iii) how to design a general cross-modal codec. On the theoretical end, we determine the minimum number of bits required to compress haptic signals under the rate conditions of video streams through investigating their semantic correlation. On the technical end, we design a general cross-modal codec to approach the optimal compression limit by using the AI-enabled cross-modal prediction and channel coding. Numerical results demonstrate that the proposed cross-modal coding can achieve significant benefits relative to the existing schemes, especially when multi-modal signals have strong semantic correlation.
Bin Kang, Xin Wei 0001, Liang Zhou 0002
IEEE Trans. Circuits Syst. Video Technol.2
2021 Cross Scene Video Foreground Segmentation Via Co-Occurrence Probability Oriented Supervised and Unsupervised Model Interaction
abstract
Using only one deep model for cross scene video foreground segmentation is still very challenging because existing methods are scene-dependent, which restricts the consistent segmentation. In this paper, we propose a cross scene video foreground segmentation framework to extend the generalization capability of those supervised model depending on scene-specific training. The proposed framework flexibly utilizes three well-trained supervised models as guidance to yield a coarse segmentation mask. The co-occurrence probability-based unsupervised background subtraction model is introduced to achieve scene adaptation in the plug and play style without any fine-tuning and labels. Experimental results on LIMU and CDNet2014 datasets validate our framework outperforms the state-of-the-art supervised/unsupervised approaches that participate in the comparison. Experiments also show the training efficiency-related improvements – when introducing the guidance models, the demand for quantity and quality of training samples to train the unsupervised model is reduced. Codes https://github.com/MeteoorLiu/Venus/tree/MeteoorLiu-SUMC
Dong Liang 0008, Bin Kang, Liyan Zhang 0001, Ningzhong Liu
ICASSP2
2021 Noise robust intuitionistic fuzzy c-means clustering algorithm incorporating local information
abstract
Abstract The human brain magnetic resonance image (MRI) is always contaminated by noise and has uncertainty on the boundary between different tissues. These characteristics bring challenges to the human brain image segmentation. To handle these limitations, many variants of standard fuzzy c‐means (FCM) algorithm have been proposed. Some methods attempt to incorporate the local spatial information in the standard FCM algorithm. However, they can't solve the problem of data uncertainty very well. And some other methods can handle the problem of data uncertainty, but they are sensitive to noise since it doesn't incorporate any local spatial information. In this paper, we propose a noise robust intuitionistic fuzzy c‐means (NR‐IFCM) algorithm, which can handle noise and uncertainty problems simultaneously. In order to process the human brain MRI with noise better, we introduce a noise robust intuitionistic fuzzy set (NR‐IFS) which is noise robust in this NR‐IFCM algorithm. Meanwhile, in order to handle the data uncertainty, we also introduce a new intuitionistic fuzzy factor to this NR‐IFCM algorithm which combine the local gray‐level and the spatial information together. A large number of experimental results on human brain MRI validate the effectiveness and the superiority of our proposed NR‐IFCM algorithm.
Zhenzhen Yang, Bin Kang
IET Image Process.4
2021 Extreme Learning Machine for Accurate Indoor Localization Using RSSI Fingerprints in Multifloor Environments
abstract
A new extreme learning machine (ELM) localization technique that uses received signal strength indicator fingerprints only is proposed for multifloor environments. This structured scheme forms multiple individual ELMs for the floors as well as for the geographically formed data clusters of each floor. Multifloor environments often have huge amount of training and online measurement data. To maximize efficiency, we develop a data preprocessing algorithm, aiming to: 1) efficiently extract out only the essential information from the vast amount of data sets and reduce the data dimension and 2) transform the floor-level data sets and positioning data sets of each floor into a proper structure that is suitable for the proposed ensemble ELM technique. The proposed solution is unique in that its offline phase exploits multiple individual ELMs for all floors to generate a set of floor-level classification functions with the preprocessed training data sets, and for each floor, it exploits multiple ELMs for the data clusters to generate a set of position regression functions. The online phase executes a coarse localization step to estimate the floor by using the floor-level classification functions and a refined step to estimate the position on the floor by using the position regression functions. The proposed algorithm and several existing algorithms are implemented to perform localization using the same measured datasets in a multistory building. For both floor estimation and localization on the floor, it outperforms existing schemes. For most cases, the performance gap is substantial.
Jun Yan 0006, Guowen Qi, Bin Kang, Xiaohuan Wu, Huaping Liu 0002
IEEE Internet Things J.3
2021 Cross-scene foreground segmentation with supervised and unsupervised model communication
Dong Liang 0008, Bin Kang, Pan Gao 0001, Xiaoyang Tan, Shun'ichi Kaneko
Pattern Recognit.2
2020 TEA: Temporal Excitation and Aggregation for Action Recognition
abstract
Temporal modeling is key for action recognition in videos. It normally considers both short-range motions and long-range aggregations. In this paper, we propose a Temporal Excitation and Aggregation (TEA) block, including a motion excitation (ME) module and a multiple temporal aggregation (MTA) module, specifically designed to capture both short- and long-range temporal evolution. In particular, for short-range motion modeling, the ME module calculates the feature-level temporal differences from spatiotemporal features. It then utilizes the differences to excite the motion-sensitive channels of the features. The long-range temporal aggregations in previous works are typically achieved by stacking a large number of local temporal convolutions. Each convolution processes a local temporal window at a time. In contrast, the MTA module proposes to deform the local convolution to a group of sub-convolutions, forming a hierarchical residual architecture. Without introducing additional parameters, the features will be processed with a series of sub-convolutions, and each frame could complete multiple temporal aggregations with neighborhoods. The final equivalent receptive field of temporal dimension is accordingly enlarged, which is capable of modeling the long-range temporal relationship over distant frames. The two components of the TEA block are complementary in temporal modeling. Finally, our approach achieves impressive results at low FLOPs on several action recognition benchmarks, such as Kinetics, Something-Something, HMDB51, and UCF101, which confirms its effectiveness and efficiency.
Yan Li 0043, Xintian Shi, Jianguo Zhang 0001, Bin Kang, Limin Wang 0002
CVPR5
2020 FDDWNet: A Lightweight Convolutional Neural Network for Real-Time Semantic Segmentation
abstract
This paper introduces a lightweight convolutional neural network, called FDDWNet, for real-time accurate semantic segmentation. In contrast to recent advances of lightweight networks that prefer to utilize shallow structure, FDDWNet makes an effort to design more deeper network architecture, while maintains faster inference speed and higher segmentation accuracy. Our network uses factorized dilated depth-wise separable convolutions (FDDWC) to learn feature representations from different scale receptive fields with fewer model parameters. Additionally, FDDWNet has multiple branches of skipped connections to gather context cues from intermediate convolution layers. The experiments show that FDDWNet only has 0.8M model size, while achieves 60 FPS running speed on a single RTX 2080Ti GPU with a 1024 × 512 input image. The comprehensive experiments demonstrate that our model achieves state-of-the-art results in terms of available speed and accuracy trade-off on CityScapes and CamVid datasets.
Quan Zhou 0004, Yong Qiang, Bin Kang, Xiaofu Wu, Baoyu Zheng
ICASSP4
2020 Deep-broad Learning System for Traffic Flow Prediction toward 5G Cellular Wireless Network
abstract
Nowadays, accurate traffic flow prediction toward 5G cellular wireless network has become an indispensable part for future artificial intelligence (AI)-assisted network. Meanwhile, low delay communication is also the essential part in the upcoming 5G era. However, traditional deep learning models applied in traffic flow prediction have many drawbacks, such as too much running time and computational resources. To tackle these issues, especially jointly considering effectiveness and efficiency, we design a deep-broad learning system (DBLS) for traffic flow prediction. Specifically, based on broad learning system (BLS), we firstly adopt deep representative learning to extract meaningful information from raw data in mapped feature nodes. Then, to further improve the performance of prediction, we add some other nodes i.e., enhancement nodes generated from mapped features as extra inputs to enhance the representative capability. Finally, taking mapped features nodes and enhancement nodes as inputs of the last-layer neural network, we apply ridge regression to compute the final weights quickly. Experimental results demonstrate that our proposed DBLS can make full use of advantages of both deep neural network and traditional BLS to increase the accuracy of traffic flow prediction, meanwhile, maintaining low complexity and running time.
Mingzi Chen, Xin Wei 0001, Liqi Huang, Mingkai Chen 0001, Bin Kang
IWCMC6
2020 Federated Quantile Regression over Networks
abstract
In order to solve the issue of isolated data islands and data security and personal privacy in the development of artificial intelligence, federated machine learning effectively solves the problem of sharing knowledge while protecting user privacy and data security. In a wireless sensor network, a secure learning framework is particularly needed, so that each sensor node can jointly learn knowledge without leaking local node data. Compared to traditional regression analysis algorithms, quantile regression can more fully describe the relationship between response values and its covariates by estimating conditional quantile sequences rather than a single value (such as the mean). In this paper, we propose a quantile regression federated learning framework that applies quantile regression with federated learning frameworks to wireless sensor networks, studies the performance of the algorithms, and the effectiveness of the algorithms is verified by simulations.
Liqi Huang, Xin Wei 0001, Peikang Zhu, Mingkai Chen 0001, Bin Kang
IWCMC6
2020 MEC-enabled video streaming in device-to-device networks
abstract
By offloading video streaming from the centralised cloud to the edge, mobile edge computing (MEC) servers offer new opportunities for real‐time video transmission. Deploying on the edge of users can ensure low latency transmission, however, the limited storage and computing ability cannot adapt to the currently used video transmission technologies such as video transcoding or simulcast. To solve this problem, a more flexible video transmission architecture needs to be considered. Under this motivation, the authors propose a device‐to‐device (D2D) assisted video streaming scheme, which fuses the technical advantages of MEC and scalable video coding. Specifically, they first construct a novel architecture for delay‐sensitive live video streaming services in edge‐enabled wireless heterogeneous networks named MEC‐enabled goodput‐aware (MEGA) model. Then they present a mathematical formulation for optimising the aggregation goodput performance of video traffic including both cellular and D2D links. Finally, they derive a three‐step solution based on a distributed heuristic algorithm. Numerical simulation results show that MEGA outperforms existing models in terms of goodput, end‐to‐end delay, effective loss rate, and users' quality‐of‐experience.
Huangda Lin, Mingkai Chen 0001, Bin Kang, Lei Wang 0009
IET Commun.4
2020 Grayscale-Thermal Tracking via Inverse Sparse Representation-Based Collaborative Encoding
abstract
Grayscale-thermal tracking has attracted a great deal of attention due to its capability of fusing two different yet complementary target observations. Existing methods often consider extracting the discriminative target information and exploring the target correlation among different images as two separate issues, ignoring their interdependence. This may cause tracking drifts in challenging video pairs. This paper presents a collaborative encoding model called joint correlation and discriminant analysis based inver-sparse representation (JCDA-InvSR) to jointly encode the target candidates in the grayscale and thermal video sequences. In particular, we develop a multi-objective programming to integrate the feature selection and the multi-view correlation analysis into a unified optimization problem in JCDA-InvSR, which can simultaneously highlight the special characters of the grayscale and thermal targets through alternately optimizing two aspects: the target discrimination within a given image and the target correlation across different images. For robust grayscale-thermal tracking, we also incorporate the prior knowledge of target candidate codes into the SVM based target classifier to overcome the overfitting caused by limited training labels. Extensive experiments on GTOT and RGBT234 datasets illustrate the promising performance of our tracking framework.
Bin Kang, Dong Liang 0008, Wan Ding, Huiyu Zhou 0001, Wei-Ping Zhu 0001
IEEE Trans. Image Process.1
2019 Grayscale-thermal Tracking via Canonical Correlation Analysis Based Inverse Sparse Representation
abstract
The grayscale-thermal tracking has attracted increasing attention due to the fact that it can make thermal information complement with grayscale information. Since there exists a large gap between the grayscale and the thermal video sequences, how to exploit the intrinsic relation between the grayscale and the thermal targets has become the key point. To address this issue, in this paper, we propose an inverse sparse representation based framework for the grayscale-thermal tracking, in which a canonical correlation analysis based inverse sparse representation model is adopted to jointly encode the target candidates in the grayscale and the thermal video sequences. The target coding process can explore the similarity between the grayscale and the thermal appearance in a common subspace, which can highlight the useful and discriminative information in both grayscale and thermal targets. The experiments on OSU-CT dataset can illustrate the promising performance of our tracking framework.
Wan Ding, Bin Kang, Quan Zhou 0004, Min Lin 0001, Suofei Zhang
ICASSP2
2019 Score-specific Non-maximum Suppression and Coexistence Prior for Multi-scale Face Detection
abstract
Face detection is an ultimate component to support various visual facial related tasks. However, detecting faces with extremely low resolution or high occlusion is still an open problem. In this paper, we propose a two-step general approach to refine the performance of modern face detectors according to human's high-level context-aware ability. First, we propose Score-specific Non-Maximum Suppression (SNMS) to preserve overlapped faces. Second, we consider the coexistence prior among faces in the scene, which could raise the sensitivity of face detection in the crowd. When integrating our approach to the existing face detectors, most of them have better results on a challenging benchmark (WIDER FACE) and a newly proposed dataset (Faces in Crowd, FIC) made by us. Codes are available on https://github.com/AIoTP/SNMSandCoexistence.
Tianpeng Wu, Dong Liang 0008, Jiaxing Pan, Bin Kang, Shun'ichi Kaneko, Huiyu Zhou 0001
ICASSP5
2019 Intelligent Content Sharing Based on Cooperative Crowdsensing
abstract
Mobile crowdsensing (MCS) has become a promising solution to support the location-based content sharing applications. To meet users' demand on personalized content sharing, a general region of interest (RoI) distribution model that allows each user to have its specific RoI needs to be considered. In this context, how to deal with the asymmetry of cooperation caused by different RoI distribution is of significance for achieving the full benefits of personalized content sharing. Thus motivated, we propose an intelligent content sharing scheme based on cooperative crowdsensing, which ensures both efficiency and fairness. Specifically, users' decision-making of whether to participate in MCS is cast as a MCS participation game (MPG). The game captures the impact of different RoI distributions on the collective cooperation of MCS. By computing the Nash equilibrium of MPG with desirable properties, we develop a cooperation scheme that maximizes the overall system utility and is acceptable to all users. The system efficiency of the proposed scheme is further quantified by numerical simulations over various parameters.
Lindong Zhao, Lei Wang 0009, Mingkai Chen 0001, Bin Kang, Baoyu Zheng
ICC4
2019 Delay Constrainted-Rate Allocation for SVC over Device-to-Device Networks
abstract
Device-to-Device (D2D) multicast content sharing is becoming a promising technology to alleviate video traffic overload and can improve the quality of local area services. Whereas existing studies mainly focus on the delay or throughput performance. However, for delay sensitive real-time video traffic, throughput as an indicator of the network-layer cannot properly indicate the benefits of upper-layer applications. Thus, in this paper we propose a video multicast scheme for D2D cooperative scalable video coding (SVC) streaming distribution to cope with the difference between multicast channels firstly. Then, we have provided analytical expressions of the goodput in heterogeneous multicast networks based on D2D collaboration, and a distributed heuristic algorithm is proposed to solve this NP-hard optimization problem. Our results show that the proposed scheme can effectively reduce end-to-end delay, effective loss rate and improve the goodput in the system.
Lei Wang 0009, Huangda Lin, Mingkai Chen 0001, Bin Kang, Wenqin Zhuang
IWCMC4
2019 Robust visual tracking via nonlocal regularized multi-view sparse representation
Bin Kang, Wei-Ping Zhu 0001, Dong Liang 0008
Pattern Recognit.1
2019 Visual Tracking Via Multi-Layer Factorized Correlation Filter
abstract
Pruning the parameters of basis filters can effectively eliminate the negative effect of redundant deep features in discriminative correlation filter based trackers. However, traditional methods often treat feature maps in Convolutional Neural Networks (CNN) as isolate observations, ignore the intrinsic correlation between partially attentional feature maps in multiple convolutional layers, when basis filter pruning is pursued. In this letter, we propose a multi-layer factorized discriminant correlation filter (MLF-DCF) for visual tracking. By integrating the multi-view discriminant learning and the discriminative correlation filter into a unified optimization problem, we can explore the correlation between different target sub-regions from multi-layer viewpoint, thus can effectively prune multi-layer basis filters. To enhance the efficiency of MLF-DCF in terms of speed and accuracy, we not only adopt alternating direction method of multipliers (ADMM) to solve the unified optimization, but also employ a mask estimation strategy to eliminate the background noise in deep features. A large number of experiments on challenging video sequences are given to illustrate the superiority of our tracking method.
Bin Kang, Gaowei Chen, Quan Zhou 0004, Jun Yan 0006, Min Lin 0001
IEEE Signal Process. Lett.1
2018 Social-Aware Cooperative Video Distribution via SVC Streaming Multicast
abstract
Scalable Video Coding (SVC) streaming multicast is considered as a promising solution to cope with video traffic overload and multicast channel differences. To solve the challenge of delivering high‐definition SVC streaming over burst‐loss prone channels, we propose a social‐aware cooperative SVC streaming multicast scheme. The proposed scheme is the first attempt to enable D2D cooperation for SVC streaming multicast to conquer the burst‐loss, and one salient feature of it is that it takes fully into account the hierarchical encoding structure of SVC in scheduling cooperation. By using our scheme, users form groups to share video packets among each other to restore incomplete enhancement layers. Specifically, a cooperative group formation method is designed to stimulate effective cooperation, based on coalitional game theory; and an optimal D2D links scheduling scheme is devised to maximize the total decoded enhancement layers, based on potential game theory. Extensive simulations using real video traces corroborate that the proposed scheme leads to a significant gain on the received video quality.
Lindong Zhao, Lei Wang 0009, Bin Kang
Wirel. Commun. Mob. Comput.4
2017 Robust visual tracking via multi-view discriminant based sparse representation
abstract
In traditional sparse representation based visual tracking, the particles are densely sampled, the appearance of some candidates may be very similar, hence the particle observations can be divided into disjointed groups. Existing methods only exploit the group similarity in a certain feature space. In this paper we propose a multi-view discriminant based multitask sparse representation method to exploit the group similarity in a multi-feature space. The proposed method can discriminate the reliability of observation groups and achieve a proper multi-view fusion by using a multi-view discriminant matrix to project multi-feature observation groups into a common subspace. Experiment results show that our method can achieve a better tracking performance than state-of-the-art tracking methods do.
Bin Kang, Dong Liang 0008, Suofei Zhang
ICIP1
2017 Adaptive local spatial modeling for online change detection under abrupt dynamic background
abstract
Change detection is an important theme in video processing. To provide reliable detection results in challenging scenes, traditional methods introduced sophisticated statistical distributions and handcraft spatial features to build background models. In this paper, we develop an intuitive background model based on simple statistical distribution and adaptive spatial correlation among pixels: For each observed pixel, we select a group of supporting pixels with high correlation, and then employ a single Gaussian to model the intensity deviations of each pixel pair. To compensate camera motion and fast adapt to dynamic pattern that coming afterwards, a randomized multichannel on-line updating mechanism is introduced. This observation is robust to abrupt illumination variation and dynamic background. Experimental results using all the video sequences provided by three challenging benchmarks (CDW-2012, CDW-2014 and SABS) validate it outperforms many state-of-the-art methods under various situations.
Dong Liang 0008, Shun'ichi Kaneko, Bin Kang
ICIP4
2017 Robust multi-feature visual tracking via multi-task kernel-based sparse learning
abstract
Feature selection and fusion is of crucial importance in multi‐feature visual tracking. This study proposes a multi‐task kernel‐based sparse learning method for multi‐feature visual tracking. The proposed sparse learning method can discriminate the reliable and unreliable features for optimal multi‐feature fusion through using a Fisher discrimination criterion‐based multi‐objective model to adaptively train the kernel weights of different features such as pixel intensity, edge and texture. To guarantee a robustness of the sparse representation method, a mixed norm is employed in the sparse leaning method to adaptively select correlated particle observations for multi‐task sparse reconstruction. Experimental results show that the proposed sparse learning method can achieve a better tracking performance than state‐of‐the‐art tracking methods do.
Bin Kang, Wei-Ping Zhu 0001, Dong Liang 0008
IET Image Process.1
2015 Robust moving object detection using compressed sensing
abstract
Moving object detection plays a key role in video surveillance. A number of object detection methods have been proposed in the spatial domain. In this study, the authors propose a compressed sensing‐based algorithm for the detection of moving object. They first use a practical three‐dimensional circulant sampling method to yield sampled measurements. Then, they propose an object detection model to simultaneously reconstruct the foreground support, background and video sequence using the sampled measurements directly. Experimental results show that the proposed moving object detection algorithm outperforms the state‐of‐the‐art approaches and it is robust to the movement turbulence, camera motion and video noise.
Bin Kang, Wei-Ping Zhu 0001
IET Image Process.1
2013 Fusion framework for multi-focus images based on compressed sensing
abstract
In this study, an efficient image fusion framework for multi‐focus images is proposed based on compressed sensing. The new fusion framework consists of three parts: image sampling, measurement fusion and image reconstruction. First, the dual‐channel pulse coupled neural network model is used in the image sampling part as an important weighting factor in the fusion scheme. Second, the result from the measurement fusion part is reconstructed through a new reconstruction algorithm called self‐adaptively modified Landwebber filter. Finally, computer simulation‐based experiment is conducted, showing that the novel fusion framework is capable of saving computational resource and enhancing the fusion result and is easy to implement.
Bin Kang, Wei-Ping Zhu 0001, Jun Yan 0006
IET Image Process.1
2008 Modeling the Effects of Cell Cycle M-phase Transcriptional Inhibition on Circadian Oscillation
abstract
Circadian clocks are endogenous time-keeping systems that temporally organize biological processes. Gating of cell cycle events by a circadian clock is a universal observation that is currently considered a mechanism serving to protect DNA from diurnal exposure to ultraviolet radiation or other mutagens. In this study, we put forward another possibility: that such gating helps to insulate the circadian clock from perturbations induced by transcriptional inhibition during the M phase of the cell cycle. We introduced a periodic pulse of transcriptional inhibition into a previously published mammalian circadian model and simulated the behavior of the modified model under both constant darkness and light-dark cycle conditions. The simulation results under constant darkness indicated that periodic transcriptional inhibition could entrain/lock the circadian clock just as a light-dark cycle does. At equilibrium states, a transcriptional inhibition pulse of certain periods was always locked close to certain circadian phases where inhibition on Per and Bmal1 mRNA synthesis was most balanced. In a light-dark cycle condition, inhibitions imposed at different parts of a circadian period induced different degrees of perturbation to the circadian clock. When imposed at the middle- or late-night phase, the transcriptional inhibition cycle induced the least perturbations to the circadian clock. The late-night time window of least perturbation overlapped with the experimentally observed time window, where mitosis is most frequent. This supports our hypothesis that the circadian clock gates the cell cycle M phase to certain circadian phases to minimize perturbations induced by the latter. This study reveals the hidden effects of the cell division cycle on the circadian clock and, together with the current picture of genome stability maintenance by circadian gating of cell cycle, provides a more comprehensive understanding of the phenomenon of circading gating of cell cycle.
Bin Kang, Xiao Chang
PLoS Comput. Biol.1