EDBT 2026 Demo / reviewers in the wild / expert
Qiang Zhai
dblp:117/3413
· DBLP profile ↗
24ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Computer networks · 8 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Progressive Knowledge Distillation for Edge-Deployable Solder Joint SegmentationabstractABSTRACT High‐precision segmentation of solder joint quality is vital for reliable electronics manufacturing, especially in flexible printed circuit ribbon cables, where even subtle defects can trigger severe functional failures. To address the fundamental challenge of model capacity gap in knowledge distillation for resource‐constrained devices, we present Solder‐Yolo, a lightweight instance segmentation framework with three core technical contributions: (1) a module‐level pruning strategy that reduces computational complexity by 32.8% and parameters by 86.2% compared to YOLOv8n; (2) a progressive knowledge distillation framework that bridges the capacity gap between heavy teachers and lightweight students through intermediate knowledge refinement—where YOLOv8l‐seg acts as a ‘knowledge translator’ to make complex features digestible for the compact model; and (3) a hierarchical context attention module that generates spatial‐channel aware feature maps as augmented knowledge to compensate for representational limitations induced by pruning. Extensive experiments on our newly curated industrial dataset show that Solder‐Yolo achieves 96.7% precision and 91.3% mAP, outperforming the baseline by 8.6% in precision while maintaining real‐time inference speed. The progressive distillation approach demonstrates a 3.4% improvement in mAP@50‐95 over direct distillation, validating its efficacy in mitigating the capacity gap problem for edge deployment. Qiang Zhai |
IET Image Process. | 5 |
| 2026 | DWDMixer: Detail-Aware Wavelet Decomposition and Mixing for Time Series Forecasting in IoT SystemsabstractAccurate time series forecasting is essential for proactive management and resource allocation in intelligent IoT systems, where data are often affected by sensor noise, multi-scale periodicity, and cross-frequency interactions. Although decomposition-based frameworks are widely adopted, most existing approaches rely on time-domain modeling with limited nonlinear capacity or treat frequency decomposition as a generic preprocessing step. Such designs fail to align model inductive bias with the intrinsic multi-scale structure of seasonal dynamics, leading to incomplete seasonal representations and degraded forecasting performance. To address this limitation, we propose Detail-aware Wavelet Decomposition and Mixing (DWDMixer), a seasonal-centric forecasting architecture built on multi-scale representation learning and seasonal-trend decomposition, where trend components serve as complementary signals. To enable fine-grained seasonal modeling, we propose a Detail-aware Wavelet Transform Hybrid (DWTH) block that performs hierarchical wavelet decomposition guided by detail-sensitive representations. DWTH captures cross-scale and cross-frequency interactions while jointly modeling temporal and spectral information, yielding expressive multi-resolution representations. Extensive experiments on 10 benchmark datasets demonstrate that DWDMixer significantly outperforms state-of-the-art models, achieving error reductions of up to 9% in MSE for long-term forecasting and 22% in MAE for short-term forecasting. Our code is available at https://github.com/zpc2002zpc/DWDMixer. Wei Huang 0037, Jie Wang 0152, Ran Peng, Qiang Zhai, Xiaocao Ouyang |
IEEE Internet Things J. | 5 |
| 2026 | SegMIC: A universal model for medical image segmentation through in-context learning
Fan Yang 0054, Xin Li 0079, Zhicheng Jiao, Qiang Zhai, Xiaomeng Li 0001, De Wu, Huazhu Fu, Hong Cheng 0002 |
Pattern Recognit. | 5 |
| 2025 | MExD: An Expert-Infused Diffusion Model for Whole-Slide Image ClassificationabstractWhole Slide Image (WSI) classification poses unique challenges due to the vast image size and numerous non-informative regions, which introduce noise and cause data imbalance during feature aggregation. To address these issues, we propose MExD, an Expert-Infused Diffusion Model that combines the strengths of a Mixture-of-Experts (MoE) mechanism with a diffusion model for enhanced classification. MExD balances patch feature distribution through a novel MoE-based aggregator that selectively emphasizes relevant information, effectively filtering noise, addressing data imbalance, and extracting essential features. These features are then integrated via a diffusion-based generative process to directly yield the class distribution for the WSI. Moving beyond conventional discriminative approaches, MExD represents the first generative strategy in WSI classification, capturing fine-grained details for robust and precise results. Our MExD is validated on three widely-used benchmarks—Camelyon16, TCGANSCLC, and BRACS—consistently achieving state-of-the-art performance in both binary and multi-class tasks. Our code and model are available at https://github.com/JWZhao-uestc/MExD. Xin Li 0079, Fan Yang 0054, Qiang Zhai, Ao Luo, Yang Zhao 0024, Hong Cheng 0002, Huazhu Fu |
CVPR | 4 |
| 2025 | Memory-Guided Transformer with group attention for knee MRI diagnosis
Rui Huang 0008, Zonghai Huang, Hantang Zhou, Qiang Zhai, Fengjun Mu, Huayi Zhan, Hong Cheng 0002 |
Pattern Recognit. | 4 |
| 2024 | Teach CLIP to Develop a Number Sense for Ordinal Regression
Qiang Zhai, Weihang Dai, Xiaomeng Li 0001 |
ECCV (85) | 2 |
| 2024 | FocusDiffuser: Perceiving Local Disparities for Camouflaged Object Detection
Xin Li 0079, Fan Yang 0054, Qiang Zhai, Ao Luo, Zicheng Jiao, Hong Cheng 0002 |
ECCV (53) | 4 |
| 2023 | MGL: Mutual Graph Learning for Camouflaged Object DetectionabstractCamouflaged object detection, which aims to detect/segment the object(s) that blend in with their surrounding, remains challenging for deep models due to the intrinsic similarities between foreground objects and background surroundings. Ideally, an effective model should be capable of finding valuable clues from the given scene and integrating them into a joint learning framework to co-enhance the representation. Inspired by this observation, we propose a novel Mutual Graph Learning (MGL) model by shifting the conventional perspective of mutual learning from regular grids to graph domain. Specifically, an image is decoupled by MGL into two task-specific feature maps - one for finding the rough location of the target and the other for capturing its accurate boundary details. Then, the mutual benefits can be fully exploited by reasoning their high-order relations through graphs recurrently. It should be noted that our method is different from most mutual learning models that model all between-task interactions with the use of a shared function. To increase information interactions, MGL is built with typed functions for dealing with different complementary relations. To overcome the accuracy loss caused by interpolation to higher resolution and the computational redundancy resulting from recurrent learning, the S-MGL is equipped with a multi-source attention contextual recovery module, called R-MGL_v2, which uses the pixel feature information iteratively. Experiments on challenging datasets, including CHAMELEON, CAMO, COD10K, and NC4K demonstrate the effectiveness of our MGL with superior performance to existing state-of-the-art methods. The code can be found at https://github.com/fanyang587/MGL. Qiang Zhai, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Ping Luo 0002, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Co-Communication Graph Convolutional Network for Multi-View Crowd CountingabstractWe study and address the multi-view crowd counting (MVCC) problem which poses more realistic challenges than single-view crowd counting for better facilitating crowd management/public safety systems. Its major challenge lies in how to fully distill and aggregate useful, complementary information among multiple camera views to create powerful ground-plane representations for wide-area crowd analysis. In this paper, we present a graph-based, multi-view learning model called Co-Communication Graph Convolutional Network (CoCo-GCN) to jointly investigate intra-view contextual dependencies and inter-view complementary relations. More specifically, CoCo-GCN builds a view-agnostic graph interaction space for each camera view to conduct efficient contextual reasoning, and extends the intra-view reasoning by using a novel Graph Communication Layer (GCL) to also take between-graph (cross-view), complementary information into account. Moreover, CoCo-GCN uses a new Co-Memory Layer (CoML) to jointly coarsen the graphs and close the ‘representational gap’ among them for further exploiting the compositional nature of graphs and learning more consistent representations. Finally, these jointly learned features of multiple views can be easily fused to create ground-plane representations for wide-area crowd counting. Experiments show that the proposed CoCo-GCN achieves state-of-the-art results on three MVCC datasets, i.e., PETS2009, DukeMTMC, and City Street, significantly improving the scene-level accuracy over previous models. Qiang Zhai, Fan Yang 0054, Xin Li 0079, Guosen Xie, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Mutual Graph Learning for Camouflaged Object DetectionabstractAutomatically detecting/segmenting object(s) that blend in with their surroundings is difficult for current models. A major challenge is that the intrinsic similarities between such foreground objects and background surroundings make the features extracted by deep model indistinguishable. To overcome this challenge, an ideal model should be able to seek valuable, extra clues from the given scene and incorporate them into a joint learning framework for representation co-enhancement. With this inspiration, we design a novel Mutual Graph Learning (MGL) model, which generalizes the idea of conventional mutual learning from regular grids to the graph domain. Specifically, MGL decouples an image into two task-specific feature maps — one for roughly locating the target and the other for accurately capturing its boundary details — and fully exploits the mutual benefits by recurrently reasoning their high-order relations through graphs. Importantly, in contrast to most mutual learning approaches that use a shared function to model all between-task interactions, MGL is equipped with typed functions for handling different complementary relations to maximize information interactions. Experiments on challenging datasets, including CHAMELEON, CAMO and COD10K, demonstrate the effectiveness of our MGL with superior performance to existing state-of-the-art methods. Code is available at https://github.com/fanyang587/MGL. Qiang Zhai, Xin Li 0079, Fan Yang 0054, Chenglizhao Chen, Hong Cheng 0002, Deng-Ping Fan |
CVPR | 1 |
| 2021 | Uncertainty-Guided Transformer Reasoning for Camouflaged Object DetectionabstractSpotting objects that are visually adapted to their surroundings is challenging for both humans and AI. Conventional generic / salient object detection techniques are suboptimal for this task because they tend to only discover easy and clear objects, while overlooking the difficult-to-detect ones with inherent uncertainties derived from indistinguishable textures. In this work, we contribute a novel approach using a probabilistic representational model in combination with transformers to explicitly reason under uncertainties, namely uncertainty-guided transformer reasoning (UGTR), for camouflaged object detection. The core idea is to first learn a conditional distribution over the backbone's output to obtain initial estimates and associated uncertainties, and then reason over these uncertain regions with attention mechanism to produce final predictions. Our approach combines the benefits of both Bayesian learning and Transformer-based reasoning, allowing the model to handle camouflaged object detection by leveraging both deterministic and probabilistic information. We empirically demonstrate that our proposed approach can achieve higher accuracy than existing state-of-the-art models on CHAMELEON, CAMO and COD10K datasets. Code is available at https://github.com/fanyang587/UGTR. Fan Yang 0054, Qiang Zhai, Xin Li 0079, Rui Huang 0008, Ao Luo, Hong Cheng 0002, Deng-Ping Fan |
ICCV | 2 |
| 2019 | Efficient health-related abnormal behavior detection with visual and inertial sensor integration
Ying Li 0138, Qiang Zhai, Sihao Ding 0001, Fan Yang 0059, Yuan F. Zheng |
Pattern Anal. Appl. | 2 |
| 2018 | MV-Sports: A Motion and Vision Sensor Integration-Based Sports Analysis SystemabstractRecently, intelligent sports analytics is becoming a hot area in both industry and academia for coaching, practicing tactic and technical analysis. With the growing trend of bringing sports analytics to live broadcasting, sports robots and common playfield, a low cost system that is easy to deploy and performs real-time and accurate sports analytics is very desirable. However, existing systems, such as Hawk-Eye, cannot satisfy these requirements due to various factors. In this paper, we present MV-Sports, a cost-effective system for real-time sports analysis based on motion and vision sensor integration. Taking tennis as a case study, we aim to recognize player shot types and measure ball states. For fine-grained player action recognition, we leverage motion signal for fast action highlighting and propose a long short term memory (LSTM)-based framework to integrate MV data for training and classification. For ball state measurement, we compute the initial ball state via motion sensing and devise an extended kalman filter (EKF)-based approach to combine ball motion physics-based tracking and vision positioning-based tracking to get more accurate ball state. We implement MV-Sports on commercial off-the-shelf (COTS) devices and conduct real-world experiments to evaluate the performance of our system. The results show our approach can achieve accurate player action recognition and ball state measurement with sub-second latency. Cheng Zhang 0014, Fan Yang 0059, Qiang Zhai, Dong Xuan |
INFOCOM | 4 |
| 2017 | EV-Matching: Bridging Large Visual Data and Electronic Data for Efficient SurveillanceabstractVisual (V) surveillance systems are extensively deployed and becoming the largest source of big data. On the other hand, electronic (E) data also plays an important role in surveillance and its amount increases explosively with the ubiquity of mobile devices. One of the major problems in surveillance is to determine human objects' identities among different surveillance scenes. Traditional way of processing big V and E datasets separately does not serve the purpose well because V data and E data are imperfect alone for information gathering and retrieval. Matching human objects in the two datasets can merge the good of the two for efficient large-scale surveillance. Yet such matching across two heterogeneous big datasets is challenging. In this paper, we propose an efficient set of parallel algorithms, called EV-Matching, to bridge big E and V data. We match E and V data based on their spatiotemporal correlation. The EV-Matching algorithms are implemented on Apache Spark to further accelerate the whole procedure. We conduct extensive experiments on a large synthetic dataset under different settings. Results demonstrate the feasibility and efficiency of our proposed algorithms. Fan Yang 0059, Guoxing Chen, Qiang Zhai, Xinfeng Li, Jin Teng, Junda Zhu 0001, Dong Xuan, Biao Chen 0002, Wei Zhao 0001 |
ICDCS | 4 |
| 2017 | BridgeLoc: Bridging Vision-Based Localization for RobotsabstractIn this paper, we study vision-based localization for robots. We anticipate that numerous mobile robots will serve or interact with humans in indoor scenarios such as healthcare, entertainment, and public service. Such scenarios entail accurate and scalable indoor visual robot localization, the subject of this work. Most existing vision-based localization approaches suffer from low localization accuracy and scalability issues due to visual environmental features' limited effective range and detection accuracy. In light of infrastructural cameras' wide indoor deployment, this paper proposes BRIDGELOC, a novel vision-based indoor robot localization system that integrates both robots' and infrastructural cameras. BRIDGELOC develops three key technologies: robot and infrastructural camera view bridging, rotation symmetric visual tag design, and continuous localization based on robots' visual and motion sensing. Our system bridges robots' and infrastructural cameras' views to accurately localize robots. We use visual tags with rotation symmetric patterns to extend scalability greatly. Our continuous localization enables robot localization in areas without visual tags and infrastructural camera coverage. We implement our system and build a prototype robot using commercial off-the-shelf hardware. Our real-world evaluation validates BRIDGELOC's promise for indoor robot localization. Qiang Zhai, Fan Yang 0059, Adam C. Champion, Chunyi Peng 0001, Jingchuan Wang, Dong Xuan, Wei Zhao 0001 |
MASS | 1 |
| 2017 | SurvSurf: human retrieval on large surveillance video data
Sihao Ding 0001, Ying Li 0138, Xinfeng Li, Qiang Zhai, Adam C. Champion, Junda Zhu 0001, Dong Xuan, Yuan F. Zheng |
Multim. Tools Appl. | 5 |
| 2016 | Flash-Loc: Flashing mobile phones for accurate indoor localizationabstractAccurate indoor localization is a key enabling technology for numerous applications such as indoor navigation, mobile social networking, and augmented reality. Despite major effort from the research community, state-of-the-art indoor localization performance remains unsatisfactory. Current approaches using radio frequency entail tedious site surveys and have limited accuracy. While vision-based localization techniques are promising, they struggle with human recognition and changing environments. This paper proposes Flash-Loc, an accurate indoor localization system leveraging flashes of light to localize people carrying mobile phones in areas with deployed surveillance cameras. A person's mobile phone emits a sequence of flashes that uniquely “represents” the person from the cameras' view. Flash-Loc develops three key mechanisms that distinguish people while avoiding long irritating flashes: adaptive-length flash coding, pulse width modulation based flash generation, and image subtraction based flash localization. Further, we design a system in which Flash-Loc cooperates with fingerprinting and dead reckoning for accurate human localization. We implement Flash-Loc on commercial off-the-shelf equipment. Our real-world experiments show Flash-Loc achieves accurate indoor localization by itself and in cooperation with other localization technology. In particular, Flash-Loc can localize a user 45 m away from the camera with sub-meter accuracy. Fan Yang 0059, Qiang Zhai, Guoxing Chen, Adam C. Champion, Junda Zhu 0001, Dong Xuan |
INFOCOM | 2 |
| 2016 | S-Mirror: Mirroring Sensing Signals for Mobile Robots in Indoor EnvironmentsabstractMany mobile robots are expected to work for or interact with humans indoors in applications such as guided shopping, policing, and senior care. Mobile robots' sensors alone are insufficient in order to realize these applications, infrastructural support is needed. Existing support for mobile robots requires heavy or expensive infrastructures with limited scalability or deployment of unsightly lines or magnetic strips. This paper presents S-Mirror, a novel approach that "reflects" various ambient signals towards mobile robots, greatly extending their sensing abilities. S-Mirror forms a network of S-Mirror nodes that mainly reflect visual signals (as well as electronic and acoustic signals) to assist mobile robots. To illustrate the advantages of S-Mirror, we develop a localization approach for mobile robots that integrates S-Mirror and robots' on-board motion sensors. We implement S-Mirror and a mobile robot prototype on commercial off-the-shelf hardware. Our real-world experimental validation shows that S-Mirror achieves accurate timely localization with low network bandwidth consumption as well as robustness and scalability to many mobile robots. Qiang Zhai, Fan Yang 0059, Adam C. Champion, Chunyi Peng 0001, Junda Zhu 0001, Dong Xuan, Biao Chen 0002, Yuan F. Zheng, Wei Zhao 0001 |
MSN | 1 |
| 2016 | Simultaneous body part and motion identification for human-following robots
Sihao Ding 0001, Qiang Zhai, Ying Li 0138, Junda Zhu 0001, Yuan F. Zheng, Dong Xuan |
Pattern Recognit. | 2 |
| 2015 | Human feet tracking guided by locomotion modelabstractFollowing a person is a fundamental requirement for human-robot interaction. In this paper we propose a novel tracking approach for robust human feet tracking which integrates human locomotion into tracking algorithms. The vertical displacement between the two feet is analyzed and we observe that this displacement during the walking cycle is close to a modulated cosine waveform. Based on this, we propose an adaptive model for the human walking pattern. We divide the motion of the human feet into local motion and global motion. The local motion is modeled by a modified cosine wave that updates along time. Global motion is estimated by the continuity between successive frames. This model is combined with particle filtering to guide the searching of the feet. A 2D Gaussian mask is generated according to the predicted position estimated by the motion model and used to modify the weight of the particles. Experiments are implemented in several human walking videos and the algorithm is evaluated against the generic particle filtering method. Results show that the feet can be tracked successfully with significant improvements compared to the generic particle filtering method. Ying Li 0138, Sihao Ding 0001, Qiang Zhai, Yuan F. Zheng, Dong Xuan |
ICRA | 3 |
| 2015 | VM-tracking: Visual-motion sensing integration for real-time human trackingabstractHuman tracking in video has many practical applications such as visual guided navigation, assisted living, etc. In such applications, it is necessary to accurately track multiple humans across multiple cameras, subject to real-time constraints. Despite recent advances in visual tracking research, the tracking systems purely relying on visual information fail to meet the accuracy and real-time requirements at the same time. In this paper, we present a novel accurate and real-time human tracking system called VM-Tracking. The system aggregates the information of motion (M) sensor on human, and integrates it with visual (V) data based on physical locations. The system has two key features, i.e. location-based VM fusion and appearance-free tracking, which significantly distinguish itself from other existing human tracking systems. We have implemented the VM-Tracking system and conducted comprehensive experiments on challenging scenarios. Qiang Zhai, Sihao Ding 0001, Xinfeng Li, Fan Yang 0059, Jin Teng, Junda Zhu 0001, Dong Xuan, Yuan F. Zheng, Wei Zhao 0001 |
INFOCOM | 1 |
| 2013 | EV-Human: Human localization via visual estimation of body electronic interferenceabstractHuman localization is an enabling technology for many mobile applications. As more and more people carry mobile phones with them, we can now localize a person by localizing his mobile phone. However, it is observed that presence of human bodies introduces heavy interference to mobile phone signals. This has been one of the major causes of inaccurate wireless localization for humans. In this paper, we propose using video cameras to help estimate human body's interference on mobile device's signals. We combine human orientation detection and human/phone/AP relative position inference estimation to better measure how a human blocks or reflects wireless signals. We have also developed a signal distortion compensation model. Based on these technologies, we have implemented a human localization system called EV-Human. Real world experiments show that our EV-system can accurately and robustly localize humans. Xinfeng Li, Jin Teng, Qiang Zhai, Junda Zhu 0001, Dong Xuan, Yuan F. Zheng, Wei Zhao 0001 |
INFOCOM | 3 |
| 2013 | Side-view face authentication based on wavelet and random forest with subsetsabstractThis paper provides a novel side-view face authentication method based on discrete wavelet transform and random forest. A subset selection method that increases the number of training samples and allows subsets to preserve the global information is presented. The authentication method can be summarized to have the following steps: profile extraction, wavelet decomposition, subset splitting and random forest verification. The new method takes the advantage of wavelet's localization property in both frequency and spatial domains, while maintaining the generalized properties of random forest. The implementation of the proposed method is computationally feasible and the experimental results show that the performance is satisfactory. Future improvements are discussed in the paper. Sihao Ding 0001, Qiang Zhai, Yuan F. Zheng, Dong Xuan |
ISI | 2 |
| 2012 | Enclave: Promoting Unobtrusive and Secure Mobile Communications with a Ubiquitous Electronic World
Adam C. Champion, Xinfeng Li, Qiang Zhai, Jin Teng, Dong Xuan |
WASA | 3 |