EDBT 2026 Demo / reviewers in the wild / expert
Chaochen Gu
dblp:156/8442
· DBLP profile ↗
29ranked-venue papers
1as first author
19since 2021 · last 2026
0000-0002-9748-7139ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What drives attention sinks? A study of massive activations and rotational positional encoding in large vision-language models
Xiaofeng Zhang 0006, Yuanchao Zhu, Chaochen Gu, Hao Cheng 0004, Kaijie Wu 0002 |
Inf. Process. Manag. | 3 |
| 2025 | Shallow Focus, Deep Fixes: Enhancing Shallow Layers Vision Attention Sinks to Alleviate Hallucination in LVLMsabstractXiaofeng Zhang, Yihao Quan, Chen Shen, Chaochen Gu, Xiaosong Yuan, Shaotian Yan, Jiawei Cao, Hao Cheng, Kaijie Wu, Jieping Ye. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xiaofeng Zhang 0006, Yihao Quan, Chen Shen 0003, Chaochen Gu, Xiaosong Yuan, Shaotian Yan, Hao Cheng 0004, Kaijie Wu 0002, Jieping Ye |
EMNLP | 4 |
| 2025 | EFDTR: Learnable Elliptical Fourier Descriptor Transformer for Instance SegmentationabstractPolygon-based object representations efficiently model object boundaries but are limited by high optimization complexity, which hinders their adoption compared to more flexible pixel-based methods.
In this paper, we introduce a novel vertex regression loss grounded in Fourier elliptic descriptors, which removes the need for rasterization or heuristic approximations and resolves ambiguities in boundary point assignment through frequency-domain matching.
To advance polygon-based instance segmentation, we further propose EFDTR (\textbf{E}lliptical \textbf{F}ourier \textbf{D}escriptor \textbf{Tr}ansformer), an end-to-end learnable framework that leverages the expressiveness of Fourier-based representations.
The model achieves precise contour predictions through a two-stage approach: the first stage predicts elliptical Fourier descriptors for global contour modeling, while the second stage refines contours for fine-grained accuracy. Experimental results on the COCO dataset show that EFDTR outperforms existing polygon-based methods, offering a promising alternative to pixel-based approaches. Code is available at \url{https://github.com/chrisclear3/EFDTR}. Chaochen Gu, Hao Cheng 0004, Xiaofeng Zhang 0006, Kaijie Wu 0002, Changsheng Lu |
ICML | 2 |
| 2025 | Longitudinal MRI-Clinical Multimodal Fusion for pCR Prediction in Breast Cancer
Dingrui Ma, Hao Cheng 0004, Xiaofeng Zhang 0006, Kaijie Wu 0002, Chaochen Gu, Xin-Ping Guan |
MICCAI (15) | 8 |
| 2025 | From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning TasksabstractXiaofeng Zhang, Yihao Quan, Chen Shen, Xiaosong Yuan, Shaotian Yan, Liang Xie, Wenxiao Wang, Chaochen Gu, Hao Tang, Jieping Ye. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Xiaofeng Zhang 0006, Yihao Quan, Chen Shen 0003, Xiaosong Yuan, Shaotian Yan, Liang Xie 0003, Wenxiao Wang 0001, Chaochen Gu, Hao Tang 0005, Jieping Ye |
NAACL (Long Papers) | 8 |
| 2025 | Simignore: Exploring and enhancing multimodal large model complex reasoning via similarity computation
Xiaofeng Zhang 0006, Fanshuo Zeng, Chaochen Gu |
Neural Networks | 3 |
| 2025 | Distribution Learning Based on Evolutionary Algorithm-Assisted Deep Neural Networks for Imbalanced Image ClassificationabstractImbalanced image classification faces critical challenges in balancing the quality and diversity of synthetic minority samples. This article proposes the improved estimation distribution algorithm-based latent feature distribution evolution (MEDA_LUDE) algorithm, an evolutionary algorithm-assisted deep distribution learning framework that optimizes latent feature distributions through a multivariate Gaussian mixture (GM) assumption and a novel four-phase training strategy. We introduce a large-margin GM (L-GM) loss to dynamically model covariances for feature learning and design a MEDA that evolves latent features via a similarity-guided fitness function, thus enhancing diversity while preserving synthesis quality. Extensive experiments demonstrate significant improvements: MEDA_LUDE achieves 95.9% accuracy on MNIST (imbalanced ratio-IR:100), surpassing state-of-the-art methods by 1.26% on CIFAR-10. For industrial fabric defect data sets, it elevates accuracy by 1.45% on DHU-FD and 0.92% on ALIYUN-FD, especially with precision and G-mean improvements of 2.5% and 1.17%, respectively, on DHU-FD. Visualizations confirm that MEDA_LUDE generates minority samples with superior quality-diversity tradeoffs. The framework's success in real-world fabric defect classification underscores its practical value in addressing imbalanced learning challenges. Yudi Zhao, Kuangrong Hao, Chaochen Gu, Bing Wei 0003, Xin-Ping Guan |
IEEE Trans. Cybern. | 3 |
| 2025 | Wakeup-Darkness: When Multimodal Meets Unsupervised Low-Light Image EnhancementabstractLow-light image enhancement is a crucial visual task, and many unsupervised methods overlook the degradation of visible information in low-light scenes, adversely affecting the fusion of complementary information and hindering the generation of satisfactory results. To address this, we introduce Wakeup-Darkness, a multimodal enhancement framework that innovatively enriches user interaction through voice and textual commands. This approach signifies a technical leap and represents a paradigm shift in user engagement. We introduce a Cross-Modal Feature Fusion (CMFF) that synergizes semantic and depth context with low-light enhancement operations. Moreover, we propose a Gated Residual Block (GRB) and a channel-aware Look-Up Table (LUT) to adjust the intensity distribution of each channel. Crucially, the proposed Wakeup-Darkness scheme demonstrates remarkable generalization in unsupervised scenarios. The source code can be accessed from https://github.com/zhangbaijin/Wakeup-Dakness . Xiaofeng Zhang 0006, Zishan Xu, Hao Tang 0005, Chaochen Gu, Wei Chen 0036, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Shadclips: When Parameter-Efficient Fine-Tuning with Multimodal Meets Shadow RemovalabstractSegment Anything Model (SAM), an advanced universal image segmentation model trained on an expansive visual dataset, has set a new benchmark in image segmentation and computer vision. However, it faced challenges when it came to distinguishing between shadows and their backgrounds. To address this, we proposed ShadClips, which consists of SAM-optimizer and SONet. It has dramatically enhanced SAM’s ability to segment shadow images, differentiating between the background and both soft and hard shadows adeptly. Due to its dependence on pixel point inputs, the SAM-Optimizer interface could do better. This method presents challenges, especially when dealing with long, extended shadows. To make the user experience more intuitive and effective, we incorporated the capabilities of CLIPs. Therefore, simple text descriptions like “A photo of a shadow” can be used to guide the SAM-Optimizer, allowing it to select the most relevant shadow mask from SAM’s comprehensive category list. Meanwhile, we introduce SONet to shadow removal. A large number of experiments on ISTD/SRD prove that the proposed method is effective and satisfactory. The source code of the ShadClips can be accessed from https://github.com/zhangbaijin/SAM-helps-Shadow . Xiaofeng Zhang 0006, Chaochen Gu, Zishan Xu, Hao Tang 0005, Hao Cheng 0004, Kaijie Wu 0002, Shanying Zhu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2023 | Self-Supervised Implicit Glyph Attention for Text RecognitionabstractThe attention mechanism has become the de facto module in scene text recognition (STR) methods, due to its capability of extracting character-level representations. These methods can be summarized into implicit attention based and supervised attention based, depended on how the attention is computed, i.e., implicit attention and supervised attention are learned from sequence-level text annotations and or character-level bounding box annotations, respectively. Implicit attention, as it may extract coarse or even incorrect spatial regions as character attention, is prone to suffering from an alignment-drifted issue. Supervised attention can alleviate the above issue, but it is character category-specific, which requires extra laborious character-level bounding box annotations and would be memory-intensive when handling languages with larger character categories. To address the aforementioned issues, we propose a novel attention mechanism for STR, self-supervised implicit glyph attention (SICA). SICA delineates the glyph structures of text images by jointly self-supervised text seg-mentation and implicit attention alignment, which serve as the supervision to improve attention correctness without extra character-level annotations. Experimental results demonstrate that SIGA performs consistently and significantly better than previous attention-based STR methods, in terms of both attention correctness and final recognition performance on publicly available context benchmarks and our contributed contextless benchmarks. Tongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang 0005, Yudi Zhao, Wei Shen 0002 |
CVPR | 2 |
| 2023 | An Oriented Object Detector towards DiatomsabstractAutomatic diatom detection refers to the task of identifying and characterizing diatoms based on artificial intelligence. It will replace traditional time-consuming and laborious manual microscopy method of diatom observation to greatly accelerate the process of diatom research and some diatomrelated studies, such as diatom abundance statistics, using diatom properties for environmental monitoring and paleoenvironmental reconstruction. However, complex background interference and the detection of slender diatoms with different integrity are two major challenges for automatic diatom detection. To solve the mentioned-above issues, we propose an oriented object detector for automatic diatom detection based on RepPoints, called OOD-RepPoints. Specifically, for encouraging the network to adaptively capture the feature of slender diatoms, we design a cascaded feature refinement head (CFRH) which consists of points generation stage and points refinement stage, to progressively optimize the extraction of slender diatom features. Furthermore, to fit the shape of diatoms well, especially for slender diatoms, we propose a tailored label assignment strategy for our CFRH, which contains a short side assigner (SSA) for points generation stage and an adaptive IoU thresholds assigner (AITA) for points refinement stage. Besides, we contribute a so called O-Diatom dataset for automatic diatom detection. The dataset has 1711 images which contains 3949 diatoms and provides finely manual oriented bounding box annotations. Extensive experiments demonstrated our method achieve state of the art performance and can reach mAP 89.9% which is highest on O-Diatom, and shows competitive results on slender categories of publicly available datasets (i.e., DOTA and HRSC2016). Song Gong, Kaijie Wu 0002, Zhiying Xia, Lihua Ran, Chaochen Gu, Changsheng Lu, Tongkun Guan, Yudi Zhao |
IJCNN | 5 |
| 2023 | SpA-Former:An Effective and lightweight Transformer for image shadow removalabstractIn this paper, we propose an Effective and lightweight Transformer for image shadow detection and removal named SpA-Former to recover a shadow-free image from a single shaded image. In contrast to conventional methods that require two stages for shadow detection and then shadow removal, the SpA-Former is a one-stage network capable of learning the mapping function between shadows and no shadows, and does not require a separate shadow detection. SpA-Former is composed of Transformer encoder and CNN decoder, where the CNN decoder contains the GAN network. In the Transformer encoding stage, Gated Feed-Forward Network(GFFN) is devised to control the information flow. In the CNN decoding stage, Two-wheel RNN joint spatial attention(TWRNN) and Fourier transform residual block (FTR) are designed to achieve satisfactory results in shadow removal. The combination of Transformer and CNN is able to feed global features from the Vision Transformer encoder into CNN to enhance the global perception of CNN branches, taking into account the complementarity of local features and the global. The SpA-Former's inference speed is 0.0459s, and the final Parameters and FLOPS are only 0.47MB and 15G, achieving the current lightweight of image shadow removal. The source code of MemoryNet can be obtained from https://github.com/zhangbaijin/SpA-Former-shadow-removal Xiaofeng Zhang 0006, Yudi Zhao, Chaochen Gu, Changsheng Lu, Shanying Zhu |
IJCNN | 3 |
| 2023 | The Cyber-Physical System of Machine Tool Monitoring: A Model-Driven Approach With Extended Kalman Filter ImplementationabstractThe condition monitoring is essential to the advanced manufacturing process in the era of the fourth industrial revolution because it ensures the prediction and optimization of machine tool conditions via data analytics or physical modeling methods. The cyber-physical system (CPS) has the property of intellectuality, scalability, adaptability, and openness, making it suitable for machine tool monitoring. The current data-driven CPS method is prone to interpretability and generalization limitations due to the empirical selection of hyper-parameters in the model and the need for heterogeneous data. On the other hand, traditional model-driven systems are difficult to adjust models to practical working conditions data due to empirical equations constructed by offline data. This article proposes a novel model-driven cyber-physical system (MDCPS) to overcome these weaknesses. First, the physical model generates a counterpart of the machining process to form a cyber world, and sensors depict the real-time state of the machining process to form a physical world. Second, for deep fusion between the cyber and physical worlds, the extended Kalman filter (EKF) approach is applied to calibrate the empirical model with online measured data. Third, the model-based diagnosis and prediction methods are used for online monitoring and control. Case studies of MDCPS for machining monitoring are presented to prove the feasibility of this model-driven system. Dezhi Yuan, Ting Luo 0003, Chaochen Gu, Kunpeng Zhu |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | A Fast Stain Normalization Network for Cervical Papanicolaou Images
Changsheng Lu, Kaijie Wu 0002, Chaochen Gu |
ICONIP (6) | 4 |
| 2022 | Segmentation based 6D pose estimation using integrated shape pattern and RGB information
Chaochen Gu, Changsheng Lu, Shuxin Zhao, Rui Xu 0010 |
Pattern Anal. Appl. | 1 |
| 2022 | Industrial Scene Text Detection With Refined Feature-Attentive NetworkabstractDetecting the marking characters of industrial metal parts remains challenging due to low visual contrast, uneven illumination, corroded surfaces, and cluttered background of metal part images. Affected by these factors, bounding boxes generated by most existing methods could not locate low-contrast text areas very well. In this paper, we propose a refined feature-attentive network (RFN) to solve the inaccurate localization problem. Specifically, we first design a parallel feature integration mechanism to construct an adaptive feature representation from multi-resolution features, which enhances the perception of multi-scale texts at each scale-specific level to generate a high-quality attention map. Then, an attentive proposal refinement module is developed by the attention map to rectify the location deviation of candidate boxes. Besides, a re-scoring mechanism is designed to select text boxes with the best rectified location. To promote the research towards industrial scene text detection, we contribute two industrial scene text datasets, including a total of 102156 images and 1948809 text instances with various character structures and metal parts. Extensive experiments on our dataset and four public datasets demonstrate that our proposed method achieves the state-of-the-art performance. Both code and dataset are available at:https://github.com/TongkunGuan/RFN. Tongkun Guan, Chaochen Gu, Changsheng Lu, Jingzheng Tu, Kaijie Wu 0002, Xin-Ping Guan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Complementary Patch for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) based on image-level labels has been greatly advanced by exploiting the outputs of Class Activation Map (CAM) to generate the pseudo labels for semantic segmentation. However, CAM merely discovers seeds from a small number of regions, which may be insufficient to serve as pseudo masks for semantic segmentation. In this paper, we formulate the expansion of object regions in CAM as an increase in information. From the perspective of information theory, we propose a novel Complementary Patch (CP) Representation and prove that the information of the sum of the CAMs by a pair of input images with complementary hidden (patched) parts, namely CP Pair, is greater than or equal to the information of the baseline CAM. Therefore, a CAM with more information related to object seeds can be obtained by narrowing down the gap between the sum of CAMs generated by the CP Pair and the original CAM. We propose a CP Network (CPN) implemented by a triplet network and three regularization functions. To further improve the quality of the CAMs, we propose a Pixel-Region Correlation Module (PRCM) to augment the contextual in-formation by using object-region relations between the feature maps and the CAMs. Experimental results on the PAS-CAL VOC 2012 datasets show that our proposed method achieves a new state-of-the-art in WSSS, validating the effectiveness of our CP Representation and CPN. Fei Zhang 0016, Chaochen Gu, Chenyue Zhang, Yuchao Dai |
ICCV | 2 |
| 2021 | Integrating Classical Control into Reinforcement Learning Policy
Chaochen Gu, Xin-Ping Guan |
Neural Process. Lett. | 2 |
| 2021 | SVMs multi-class loss feedback based discriminative dictionary learning for image classification
Baoqing Yang, Xin-Ping Guan, Junwu Zhu, Chaochen Gu, Kaijie Wu 0002, Jiajie Xu 0004 |
Pattern Recognit. | 4 |
| 2020 | Automatic Curriculum Generation by Hierarchical Reinforcement Learning
Zhenghua He, Chaochen Gu, Rui Xu 0010, Kaijie Wu 0002 |
ICONIP (2) | 2 |
| 2020 | Double Attention for Pathology Image Diagnosis Network with Visual InterpretabilityabstractIn recent years, cervical cancer has been one of the most common diseases in women's cancer. The advanced diagnosis of cervical precancerous lesions is essential for preventing cervical cancer. Its effectiveness and efficiency can be greatly improved by computer aided diagnosis, while challenged by the imprecise conclusions and uninterpretable process of diagnosis. To solve this problem, we propose a novel deep learning-based interpretable diagnosis system for pathology images, consisting of three interrelated models: an image model, an attention model and a conclusion model. Computer aided diagnosis improves the effectiveness and efficiency of the proposed image model uses a convolutional neural network (CNN) to ex-tract semantic features. Combining the model with the semantic attribute attention model, it aims to capture the discriminant relationship between se-mantic attributes by predicting the conclusion label through long-term and short-term memory (LSTM). The network is trained in an end-to-end manner, with different weights for each model. Experimental results on cervical intraepithelial neoplasia images, diagnostic reports and label datasets show that the proposed method achieves a significant improvement over traditional methods with a better interpretability. Hao Cheng 0004, Kaijie Wu 0002, Kai Ma 0001, Rui Xu 0010, Chaochen Gu, Xin-Ping Guan |
IJCNN | 6 |
| 2020 | Deep transfer neural network using hybrid representations of domain discrepancy
Changsheng Lu, Chaochen Gu, Kaijie Wu 0002, Si-Yu Xia, Xin-Ping Guan |
Neurocomputing | 2 |
| 2019 | PointDoN: A Shape Pattern Aggregation Module for Deep Learning on Point CloudabstractAs point cloud is a typical and significant type of geometric 3D data, deep learning on the classification and segmentation of point cloud has received widely interests recently. However, the critical problems to process the irregularity of point cloud and feature extraction of shape pattern have not yet been fully explored. In this paper, a geometric deep learning architecture based on our PointDoN module is presented. Inspired by the Difference of Normals (DoN) in traditional point clouds processing, our PointDoN module is a feature aggregation module combining DoN shape pattern descriptor with both 3D coordinates and extra features (such as RGB colors). Our PointDoN-based architecture can be flexibly applied to multiple point cloud processing tasks such as 3D shape classification and scene semantic segmentation. Experiments demonstrate that PointDoN model achieves state-of-the-art results on multiple types of challenging benchmark datasets. Shuxin Zhao, Chaochen Gu, Changsheng Lu, Kaijie Wu 0002, Xin-Ping Guan |
IJCNN | 2 |
| 2018 | Reinforcement Learning Policy with Proportional-Integral Control
Chaochen Gu, Kaijie Wu 0002, Xin-Ping Guan |
ICONIP (3) | 2 |
| 2018 | Viewpoint Estimation for Workpieces with Deep Transfer Learning from Cold to Hot
Changsheng Lu, Chaochen Gu, Kaijie Wu 0002, Xin-Ping Guan |
ICONIP (1) | 3 |
| 2018 | A Pathology Image Diagnosis Network with Visual Interpretability and Structured Diagnostic Report
Kai Ma 0001, Kaijie Wu 0002, Hao Cheng 0004, Chaochen Gu, Rui Xu 0010, Xin-Ping Guan |
ICONIP (6) | 4 |
| 2018 | Parallel Search by Reinforcement Learning for Object Detection
Chaochen Gu, Kaijie Wu 0002, Xin-Ping Guan |
PRCV (4) | 2 |
| 2017 | Simultaneous dimensionality reduction and dictionary learning for sparse representation based classification
Baoqing Yang, Chaochen Gu, Kaijie Wu 0002, Tao Zhang 0010, Xin-Ping Guan |
Multim. Tools Appl. | 2 |
| 2016 | A novel face recognition method based on IWLD and IWBC
Baoqing Yang, Tao Zhang 0010, Chaochen Gu, Kaijie Wu 0002, Xin-Ping Guan |
Multim. Tools Appl. | 3 |