EDBT 2026 Demo / reviewers in the wild / expert
Zheng Lu 0002
dblp:15/2402-2 · also Zhen Lu 0002
· DBLP profile ↗
33ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 3 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CellMixer: Pathological image classification using dual-branch VMamba with randomly mixing gradient features data augmentationabstract• Fuses gradient-pixel features to amplify subtle pathological distinctions. • Jointly captures tissue-context and cellular details via dual-branch modeling. • Integrates multi-scale features for class uniformity-variability representation. • SOTA over datasets against costy foundational models Pathological diagnosis is crucial for patient care, and Region of Interest (ROI) analysis serves as a key pathological method for extracting local cellular details to guide precise clinical decision-making. While most of the current foundation models have shown promise in ROI pathological image classification, existing approaches often fall short in addressing the unique characteristics of pathology ROI data from three aspects simultaneously: (1) inter-class similarity, (2) complex global patterns, and (3) multi-scale granularity. To address them, we propose CellMixer, a novel framework designed to extract and integrate local-global ROI pathological image representations. The key innovation lies in the synergistic integration of three corresponding domain-aware components: (1) To amplify subtle morphological distinctions, we designed a data augmentation (GradMix), which selectively fuses gradient maps and pixel-level features to enhances low-level feature sensitivity, directly improving discrimination of visually similar classes; (2) To capture both localized patterns and global tissue structures across ROI regions, we proposed Dual-branch VMamba Block (DVB), which enhances long-range dependency modeling and simultaneously extracts cell-level fine-grained features; (3) To fuse local and global features to concurrently represent intra-class homogeneity and inter-class heterogeneity across scales, a novel feature fusion strategy (Insert-Merge (InM)). Extensive experiments on 8 public pathology ROI datasets demonstrate that CellMixer consistently outperforms existing methods, proving task-specific model, even with limited data, yields superior visual representations to generic foundation models. Enhui Chai, Zheng Lu 0002, Tianxiang Cui |
Expert Syst. Appl. | 3 |
| 2026 | Rewarding fine-grained image captioning with keyword group contrastive
Kailiang Ye, Zheng Lu 0002, LinLin Shen, Tianxiang Cui |
Expert Syst. Appl. | 2 |
| 2026 | Fine-grained facial description generation with retrieval augmentation
Kailiang Ye, Zheng Lu 0002, LinLin Shen, Tianxiang Cui |
Neurocomputing | 2 |
| 2025 | Accelerating Convergence in Bounding Box Regression with a Refined IoU Loss FunctionabstractBounding box regression (BBR) is a critical component in object detection, significantly influencing the accuracy of object localization. However, existing Intersection over Union (IoU)-based loss functions encounter two primary challenges: (i) The penalty factor configuration often results in the expansion of anchor boxes during the regression, which in turn slows the convergence rate of the loss. (ii) There is a spatial imbalance caused by the disproportionate influence of anchor boxes with minimal overlap with the ground truth boxes. To resolve these two challenges, this paper proposes a novel loss function termed Fast-IoU, designed to swiftly and precisely measure the overlap area and aspect ratio in BBR. Building upon this, a dynamic non-monotonic focusing mechanism is integrated to evaluate the quality of anchor boxes in a non-linear manner. Fast-IoU can enhance the capability to focus on anchor boxes of medium quality. By incorporating Fast-IoU into popular object detectors such as YOLOv7, YOLOv8 and YOLOv10, we achieved an increase in average precision and improved performance compared to their original loss functions on the MS COCO datasets, thus validating the effectiveness of ourproposed improvement strategies. Enhui Chai, Tianxiang Cui, Zheng Lu 0002, Fiseha B. Tesema |
ICASSP | 4 |
| 2025 | A cascaded retrieval-while-reasoning multi-document comprehension framework with incremental attention for medical question answering
Jianfeng Ren, Ruibin Bai, Zheng Lu 0002 |
Expert Syst. Appl. | 5 |
| 2024 | Seat belt detection using gated Bi-LSTM with part-to-whole attention on diagonally sampled patchesabstractOne of the high-risk behaviors leading to severe traffic injuries is not wearing a seat belt. It is therefore very important to be able to automatically detect seat belts from surveillance images, encourage drivers to wear seat belts, and enhance passenger safety. In this paper, a novel deep neural network, Gated Bi-directional Long Short-Term Memory network with part-to-whole attention (GBL-PA), is proposed for seat belt detection from surveillance images. The innovation of our model lies in its unique diagonal sampling strategy, which meticulously captures the seat belt’s fine details, typically oriented from top right to bottom left across the torso of vehicle occupants. Our framework’s novelty is further encapsulated by the part-to-whole attention mechanism, which intelligently harmonizes the detailed local information from seat belt–specific patches with the broader contextual insights from the regional proposals. The pioneering design of a Gated Bi-directional LSTM network facilitates the dynamic integration of interactions across patches to deliver an optimized final prediction. The superiority of GBL-PA is established through rigorous comparison with the state-of-the-art methods on a new, large benchmark dataset comprising 14,936 images from traffic surveillance footage. Our framework demonstrates a notable improvement, achieving a mean Average Precision (mAP) of 72.3%, which surpasses the second best, YOLOX, by 0.9% mAP. This significant and consistent outperformance across various metrics underscores the transformative potential of GBL-PA in the realm of traffic safety enforcement. The source code of our framework is available at ANONYMISED. Zheng Lu 0002, Jianfeng Ren, Qian Zhang 0018 |
Expert Syst. Appl. | 2 |
| 2024 | MITER: Medical Image-TExt joint adaptive pretRaining with multi-level contrastive learningabstractRecently multimodal medical pretraining models play a significant role in automatic medical image and text analysis that has wide social and economical impact in healthcare. Despite being able to be quickly transferred to downstream tasks, the models are greatly limited due to the fact that these models can only be pretrained with professional medical image-text datasets, which usually contain a very small number of samples. In this work We propose MITER (Medical Image-Text Joint adaptive Pretraining), a joint adaptive pretraining framework via multi-level contrastive learning to overcome this limitation by pretraining image and text models for medical domain and utilizing existing models pretrained on generic data, which contain enormous number of samples. MITER features two types of objectives to solve the problem. The first type is uni-modal objectives that pretrain the models with medical images and text separately on uni-modal tasks. The other type is a cross-modal objective that pretrains jointly, allowing the models to influence each other on cross-modal tasks. We also introduce a strategy to dynamically select hard negative samples during the training process for better performance. Experimental results over four medical tasks, image-report retrieval, multi-label image classification, visual question answering, and report generation, show that our MITER framework solves the limitation problem by greatly outperforming existing benchmark models on all the tasks. The source code of our framework is available online.2 Xiaochu Tang, Jing Xiao 0006, Youxin Chen, Xiu Li 0001, Qian Zhang 0018, Zheng Lu 0002 |
Expert Syst. Appl. | 8 |
| 2024 | Medical chief complaint classification with hierarchical structure of label descriptions
Zheng Lu 0002, Ruibin Bai |
Expert Syst. Appl. | 2 |
| 2024 | MergeTalk: Audio-Driven Talking Head Generation From Single Image With Feature MergeabstractAudio-driven talking head generation has wide real world applications but remains challenging due to the problems such as audio-lip synchronization, head poses, identity preservation, video quality, etc. We propose a novel two-stage framework that uses explicit 3D face images rendered from a 3D model based on the audio input, as intermediate features. We devise two independent 3D motion parameter generation networks to generate expression and pose parameters for the popular 3DMM model to solve the audio-lip synchronization problem and natural head poses without losing identity information. To improve the final talking head quality such as avoiding facial distortion and artifacts, we propose a novel face feature merge network to accurately extract and fuse the background, identity information, facial texture from the source image, and the lip movements and head poses from the 3D face images, and generate the final videos based on generative adversarial networks. Extensive experiments show that our framework outperforms the SOTA methods in several aspects and has good generalization ability. Ximin Zheng, Zheng Lu 0002, Nengsheng Bao |
IEEE Signal Process. Lett. | 4 |
| 2023 | Improving Visual-Semantic Embedding with Adaptive Pooling and Optimization ObjectiveabstractZijian Zhang, Chang Shu, Ya Xiao, Yuan Shen, Di Zhu, Youxin Chen, Jing Xiao, Jey Han Lau, Qian Zhang, Zheng Lu. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Ya Xiao 0006, Youxin Chen, Jing Xiao 0006, Jey Han Lau, Qian Zhang 0018, Zheng Lu 0002 |
EACL | 10 |
| 2022 | Occlusion-Invariant Representation Alignment for Entity Re-IdentificationabstractEntity re-identification is the foundation of tracking- and matching-based computer vision tasks, which are widely employed in a variety of applications. However, when trained exclusively on clear images, the models capacity to generalize is significantly affected by the presence of occlusion at referencing time, whereas data argumentation-based approaches are costly to construct without guaranteeing a test-time improvement. To tackle this problem, we propose a domain adaptation framework based on learning representations that generates occlusion-invariant feature representations by aligning the clean image embedding distribution with the occluded one, using a disparity discrepancy metric derived from the siamese network architecture. Without the need for additional processing modules during the inference stage or an expensive occlusion-augmentation-enlarged dataset during the training stage, we could obtain occlusion invariant embeddings that are free of the impact of occluders. Extensive experimental results for two tasks across three datasets indicate the proposed method’s robustness and effectiveness to a variety of occlusions at all levels. Zhanghao Jiang, Heshan Du, Huan Jin, Zheng Lu 0002, Qian Zhang 0018 |
ICIP | 5 |
| 2022 | ICAF: Iterative Contrastive Alignment Framework for Multimodal Abstractive SummarizationabstractIntegrating multimodal knowledge for abstractive summarization task is a work-in-progress research area, with present techniques inheriting fusion-then-generation paradigm. Due to semantic gaps between computer vision and natural language processing, current methods often treat multiple data points as separate objects and rely on attention mechanisms to search for connection in order to fuse together. In addition, missing awareness of cross-modal matching from many frameworks leads to performance reduction. To solve these two drawbacks, we propose an Iterative Contrastive Alignment Framework (ICAF) that uses recurrent alignment and contrast to capture the coherences between images and texts. Specifically, we design a recurrent alignment (RA) layer to gradually investigate fine-grained semantical relationships between image patches and text tokens. At each step during the encoding process, crossmodal contrastive losses are applied to directly optimize the embedding space. According to ROUGE, relevance scores, and human evaluation, our model outperforms the state-of-the-art baselines on MSMO dataset. Experiments on the applicability of our proposed framework and hyperparameters settings have been also conducted. Youxin Chen, Jing Xiao 0006, Qian Zhang 0018, Zheng Lu 0002 |
IJCNN | 6 |
| 2022 | Cross-document attention-based gated fusion network for automated medical licensing exam
Jianfeng Ren, Zheng Lu 0002, Menglin Cui, Ruibin Bai |
Expert Syst. Appl. | 3 |
| 2022 | Rain-component-aware capsule-GAN for single image de-raining
Jianfeng Ren, Zheng Lu 0002, Jialu Zhang 0003, Qian Zhang 0018 |
Pattern Recognit. | 3 |
| 2021 | You Get What You Focus on: A Weighting Factor for IoU-based Regression LossabstractLoss functions are essential to bounding box regression which plays a significant role in deep learning based object detection. Despite the effectiveness of the popular Intersection over Union (IoU) based losses, there is still an imbalance problem of high- and low-quality predicted bounding boxes, impeding the accuracy and convergence speed during bound box regression. Specifically, we observe that the huge amount of predicted bounding boxes having small overlapping regions with ground truth box overwhelms the amount of predicted bounding boxes having large overlapping regions. In this paper, we propose a simple weighting factor that is able to reshape the existing IoU-based losses according to a geometric relationship of bounding boxes. In this way, we are able to effectively down-weight the contribution of low-quality predicted boxes and focus training on high-quality ones. Extensive experiments have been carried out on popular IoU-based losses with various object detection techniques. By simply incorporating the proposed weighting factor, we are able to achieve notable performance gains on the popular MS COCO dataset. Zheng Lu 0002, Tianxiang Cui |
IJCNN | 3 |
| 2021 | A hybrid medical text classification framework: Integrating attentive rule construction and neural network
Xiang Li 0170, Menglin Cui, Jingpeng Li 0001, Ruibin Bai, Zheng Lu 0002, Uwe Aickelin |
Neurocomputing | 5 |
| 2020 | Data-Driven Regular Expressions Evolution for Medical Text Classification Using Genetic ProgrammingabstractIn medical fields, text classification is one of the most important tasks that can significantly reduce human work-load through structured information digitization and intelligent decision support. Despite the popularity of learning-based text classification techniques, it is hard for human to understand or manually fine-tune the classification for better precision and recall, due to the black box nature of learning. This study proposes a novel regular expression-based text classification method making use of genetic programming (GP) approaches to evolve regular expressions that can classify a given medical text inquiry with satisfaction. Given a seed population of regular expressions (randomly initialized or manually constructed by experts), our method evolves a population of regular expressions, using a novel regular expression syntax and a series of carefully chosen reproduction operators. Our method is evaluated with real-life medical text inquiries from an online healthcare provider and shows promising performance. More importantly, our method generates classifiers that can be fully understood, checked and updated by medical doctors, which are fundamentally crucial for medical related practices. Ruibin Bai, Zheng Lu 0002, Peiming Ge, Uwe Aickelin, Daoyun Liu |
CEC | 3 |
| 2020 | Examining the effects of social influence in pre-adoption phase and initial post-adoption phase in the healthcare context
Zheng Lu 0002, Tingru Cui, Yu Tong 0001 |
Inf. Manag. | 1 |
| 2019 | Retrieving and ranking short medical questions with two stages neural matching modelabstractInternet hospital is a rising business thanks to recent advances in mobile web technology and high demand of health care services. Online medical services become increasingly popular and active. According to US data in 2018, 80 percent of internet users have asked health-related questions online. Numerous data is generated in unprecedented speed and scale. Those representative questions and answers in medical fields are valuable raw data sources for medical data mining. Automated machine interpretation on those sheer amount of data gives an opportunity to assist doctors to answer frequently asked medical-related questions from the perspective of information retrieval and machine learning approaches. In this work, we propose a novel two-stage framework for the semantic matching of query-level medical questions, which takes advantages of sentence similarity-based search engine techniques and Siamese inspired recent recurrent neural network. The two-stage hierarchical design optimises the performance of automatic information retrieval of user queries. Compared against the classical TFIDF search technique as a single-stage, our novel soft search technique performs significantly better. Incorporating an advanced deep learning model as the second stage can improve the results further, which we believe is the new state-of-the-art in the current problem setting with the unique medical corpus from one of the largest online healthcare provider in market. Xiang Li 0170, Xinyu Fu 0001, Zheng Lu 0002, Ruibin Bai, Uwe Aickelin, Peiming Ge, Gong Liu |
CEC | 3 |
| 2018 | Capsule Based Image Synthesis for Interior Design Effect Rendering
Zheng Lu 0002, Guoping Qiu, Qian Zhang 0018 |
ACCV (5) | 2 |
| 2017 | Face super resolution based on parent patch prior for VLQ scenarios
Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Qing Li 0001, Zheng Lu 0002 |
Multim. Tools Appl. | 5 |
| 2017 | Dual Structure Constrained Multimodal Feature Coding for Social Event Detection from Flickr DataabstractIn this work, a three-stage social event detection (SED) framework is proposed to discover events from Flickr-like data. First, multiple bipartite graphs are constructed for the heterogeneous feature modalities to achieve fused features. Furthermore, considering the geometrical structures of dictionary and data, a dual structure constrained multimodal feature coding model is designed to learn discriminative feature codes by incorporating corresponding regularization terms into the objective. Finally, clustering models utilizing density or label knowledge and data recovery residual models are devised to discover real-world events. The proposed SED approach achieves the highest performance on the MediaEval 2014 SED dataset. Zhenguo Yang, Qing Li 0001, Zheng Lu 0002, Yun Ma 0001, Zhiguo Gong, Wenyin Liu |
ACM Trans. Internet Techn. | 3 |
| 2016 | Face Super Resolution for VLQ facial images via parent patch matchingabstractFace Super Resolution(FSR) is to infer High Resolution(HR) facial images from given Low Resolution(LR) ones with the assistance of LR and HR training pairs. Among existing methods, local patch based methods are superior in visual and objective quality than global based methods. These local patch based methods are based on the consistency assumption that the neighbors in HR/LR space form similar local geometry. But when LR images are Very Low Quality(VLQ), the LR space is seriously contaminated that even two distinct patches look similar, which means that the consistency assumption is not well held anymore. To this end, in this paper we use the target patch as well as the surrounding pixels, which we called parent patch, to represent the target patch. By incorporating the peripheral information, the parent patch is much more robust to noise in the LR and HR consistency learning. The effectiveness of proposed method is verified both quantitatively and qualitatively. Liang Chen 0026, Ruimin Hu, Zhen Han 0002, Zhongyuan Wang 0001, Qing Li 0001, Zheng Lu 0002 |
IJCNN | 6 |
| 2013 | Story-Driven Summarization for Egocentric VideoabstractWe present a video summarization approach that discovers the story of an egocentric video. Given a long input video, our method selects a short chain of video sub shots depicting the essential events. Inspired by work in text analysis that links news articles over time, we define a random-walk based metric of influence between sub shots that reflects how visual objects contribute to the progression of events. Using this influence metric, we define an objective for the optimal k-subs hot summary. Whereas traditional methods optimize a summary's diversity or representative ness, ours explicitly accounts for how one sub-event "leads to" another-which, critically, captures event connectivity beyond simple object co-occurrence. As a result, our summaries provide a better sense of story. We apply our approach to over 12 hours of daily activity video taken from 23 unique camera wearers, and systematically evaluate its quality compared to multiple baselines with 34 human subjects. Zheng Lu 0002, Kristen Grauman |
CVPR | 1 |
| 2013 | A 3D Imaging Framework Based on High-Resolution Photometric-Stereo and Low-Resolution Depth
Zheng Lu 0002, Yu-Wing Tai, Fanbo Deng, Moshe Ben-Ezra, Michael S. Brown |
Int. J. Comput. Vis. | 1 |
| 2013 | When Amazon Meets Google: Product Visualization by Exploring Multiple Web Sources
Meng Wang 0001, Guangda Li, Zheng Lu 0002, Yue Gao 0002, Tat-Seng Chua |
ACM Trans. Internet Techn. | 3 |
| 2012 | Synthesizing oil painting surface geometry from a single photographabstractWe present an approach to synthesize the subtle 3D relief and texture of oil painting brush strokes from a single photograph. This task is unique from traditional synthesize algorithms due to its mixed modality between the input and output; i.e., our goal is to synthesize surface normals given an intensity image input. To accomplish this task, we propose a framework that first applies intrinsic image decomposition to produce a pair of initial normal maps. These maps are combined into a conditional random field (CRF) optimization framework that incorporates additional information derived from a training set consisting of normals captured using photometric stereo on oil paintings with similar brush styles. Additional constraints are incorporated into the CRF framework to further ensures smoothness and preserve brush stroke edges. Our results show that this approach can produce compelling reliefs that are often indistinguishable from results captured using photometric stereo. Zheng Lu 0002, Xiaogang Wang 0001, Ying-Qing Xu, Moshe Ben-Ezra, Xiaoou Tang, Michael S. Brown |
CVPR | 2 |
| 2012 | Nonuniform Lattice Regression for Modeling the Camera Imaging Pipeline
Hai Ting Lin, Zheng Lu 0002, Seon Joo Kim, Michael S. Brown |
ECCV (1) | 2 |
| 2012 | A New In-Camera Imaging Model for Color Computer Vision and Its ApplicationabstractWe present a study of in-camera image processing through an extensive analysis of more than 10,000 images from over 30 cameras. The goal of this work is to investigate if image values can be transformed to physically meaningful values, and if so, when and how this can be done. From our analysis, we found a major limitation of the imaging model employed in conventional radiometric calibration methods and propose a new in-camera imaging model that fits well with today's cameras. With the new model, we present associated calibration procedures that allow us to convert sRGB images back to their original CCD RAW responses in a manner that is significantly more accurate than any existing methods. Additionally, we show how this new imaging model can be used to build an image correction application that converts an sRGB input image captured with the wrong camera settings to an sRGB output image that would have been recorded under the correct settings of a specific camera. Seon Joo Kim, Hai Ting Lin, Zheng Lu 0002, Sabine Süsstrunk, Stephen Lin 0001, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | In-video product annotation with web information miningabstractProduct annotation in videos is of great importance for video browsing, search, and advertisement. However, most of the existing automatic video annotation research focuses on the annotation of high-level concepts, such as events, scenes, and object categories. This article presents a novel solution to the annotation of specific products in videos by mining information from the Web. It collects a set of high-quality training data for each product by simultaneously leveraging Amazon and Google image search engine. A visual signature for each product is then built based on the bag-of-visual-words representation of the training images. A correlative sparsification approach is employed to remove noisy bins in the visual signatures. These signatures are used to annotate video frames. We conduct experiments on more than 1,000 videos and the results demonstrate the feasibility and effectiveness of our approach. Guangda Li, Meng Wang 0001, Zheng Lu 0002, Richang Hong, Tat-Seng Chua |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2010 | A framework for ultra high resolution 3D imagingabstractWe present an imaging framework to acquire 3D surface scans at ultra high-resolutions (exceeding 600 samples per mm2). Our approach couples a standard structured-light setup and photometric stereo using a large-format ultra-high-resolution camera. While previous approaches have employed similar hybrid imaging systems to fuse positional data with surface normals, what is unique to our approach is the significant asymmetry in the resolution between the low-resolution geometry and the ultra-high-resolution surface normals. To deal with these resolution differences, we propose a multi-resolution surface reconstruction scheme that propagates the low-resolution geometric constraints through the different frequency bands while gradually fusing in the high-resolution photometric stereo data. In addition, to deal with the ultra-high-resolution images, our surface reconstruction is performed in a patch-wise fashion and additional boundary constraints are used to ensure patch coherence. Based on this multi-resolution reconstruction scheme, our imaging framework can produce 3D scans that show exceptionally detailed 3D surfaces far exceeding existing technologies. Zheng Lu 0002, Yu-Wing Tai, Moshe Ben-Ezra, Michael S. Brown |
CVPR | 1 |
| 2009 | Directed assistance for ink-bleed reduction in old documentsabstractInk-bleed interference is a serious problem that affects the legibility of old documents. Ink-bleed can be reduced using pixel classification based on user-supplied markup that labels examples of ink-bleed, foreground-ink, and background. The main challenge is ensuring that the user's markup sufficiently captures the characteristics of the document. This is particularly troublesome for old documents that can exhibit significant change within the same page. In this paper, we address this markup problem using a “directed assistance” approach in which the user provides a small amount of initial markup. The image is then classified and regions with low classification confidence are grouped and displayed to the user for another round of markup. The key idea is to direct the user to where markup is needed. In addition, local markup can be weighted in the classification algorithm to produce better results. Zheng Lu 0002, Michael S. Brown |
CVPR | 1 |
| 2009 | Interactive degraded document binarization: An example (and case) for interactive computer visionabstractThis paper describes a user-assisted application to perform adaptive thresholding (i.e. binarization) on degraded handwritten documents. While existing adaptive thresholding techniques purport to be automatic, they in fact require the user to perform non-intuitive parameter tuning to obtain satisfactory results. In our work, we recast the problem into one where the user needs only to coarsely markup regions in the thresholded image that have unsatisfactory results. These regions are then segmented and processed locally - no parameter tuning is necessary. Our user study shows that not only do the majority of users prefer our application over parameter tuning, but our final results are better than existing algorithms due to the more targeted processing. While our main contribution is an effective user-assisted application for document binarization, we use this as an example to advocate the need to rethink how many computer vision solutions, notoriously reliant on parameter tuning, can be reworked to exploit meaningful user interaction. Zheng Lu 0002, Michael S. Brown |
WACV | 1 |