EDBT 2026 Demo / reviewers in the wild / expert
Zhili Liu
dblp:03/10297
· DBLP profile ↗
30ranked-venue papers
5as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorSystems, architecture and hardware · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AtomThink: Multimodal Slow Thinking With Atomic Step ReasoningabstractIn this paper, we address the challenging task of multimodal reasoning by incorporating the notion of "slow thinking" into multimodal large language models (MLLMs). Our core idea is that models can learn to adaptively use different levels of reasoning to tackle questions of varying complexity. We propose a novel paradigm of Self-structured Chain of Thought (SCoT), which consists of minimal semantic atomic steps. Unlike existing methods that rely on structured templates or free-form paradigms, our method not only generates flexible CoT structures for various complex tasks but also mitigates the phenomenon of overthinking for easier tasks. To introduce structured reasoning into visual cognition, we design a novel AtomThink framework with four key modules: (i) a data engine to generate high-quality multimodal reasoning paths; (ii) a supervised fine-tuning (SFT) process with serialized inference data; (iii) a policy-guided multi-turn inference method; and (iv) an atomic capability metric to evaluate the single-step utilization rate. Extensive experiments demonstrate that the proposed AtomThink significantly improves the performance of baseline MLLMs, achieving more than 10% average accuracy gains on MathVista and MathVerse. Compared to state-of-the-art structured CoT approaches, our method not only achieves higher accuracy but also improves data utilization by 5 × and boosts inference efficiency by 85.3%. Kun Xiang, Zhili Liu, Terry Jingchen Zhang, Yinya Huang, Yunshuang Nie, Kaixin Cai, Yiyang Yin, Runhui Huang, Yihan Zeng, Yu-Jie Yuan, Jianhua Han, Lanqing Hong, Hang Xu 0004, Xiaodan Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Mixture of Cluster-Conditional LoRA Experts for Vision-Language Instruction TuningabstractInstruction tuning of Large Vision-language Models (LVLMs) has revolutionized the development of versatile models with zero-shot generalization across a wide range of downstream vision-language tasks. However, the diversity of different training tasks from various sources and formats would lead to inevitable task conflicts, where different tasks conflict for the same set of model parameters, resulting in sub-optimal instruction-following abilities. To address that, we propose the Mixture of Cluster-conditional LoRA Experts (MoCLE), a novel Mixture of Experts (MoE) architecture designed to activate task-customized model parameters based on instruction clusters. A separate universal expert is further incorporated to improve generalization abilities of MoCLE for novel instructions. Extensive experiments on InstructBLIP and LLaVA demonstrate the effectiveness of MoCLE. Yunhao Gou, Zhili Liu, Kai Chen 0023, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, Yu Zhang 0006 |
IEEE Trans. Image Process. | 2 |
| 2025 | Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-AlignmentabstractZhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, James Kwok. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhili Liu, Yunhao Gou, Kai Chen 0023, Lanqing Hong, Jiahui Gao 0002, Fei Mi, Yu Zhang 0006, Zhenguo Li, Xin Jiang 0002, Qun Liu 0001, James T. Kwok |
ACL (1) | 1 |
| 2025 | EMOVA: Empowering Language Models to See, Hear and Speak with Vivid EmotionsabstractGPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging for the open-source community. Existing vision-language models rely on external tools for speech processing, while speech-language models still suffer from limited or totally without vision-understanding capabilities. To address this gap, we propose the EMOVA (EMotionally Omni-present Voice Assistant), to enable Large Language Models with end-to-end speech abilities while maintaining the leading vision-language performance. With a semantic-acoustic disentangled speech tokenizer, we surprisingly notice that omni-modal alignment can further enhance vision-language and speech abilities compared with the bi-modal aligned counterparts. Moreover, a lightweight style module is introduced for the flexible speech style controls including emotions and pitches. For the first time, EMOVA achieves state-of-the-art performance on both the vision-language and speech benchmarks, and meanwhile, supporting omni-modal spoken dialogue with vivid emotions. Yunhao Gou, Runhui Huang, Zhili Liu, Daxin Tan, Chunwei Wang, Yihan Zeng, Dingdong Wang, Kun Xiang, Haoli Bai, Jianhua Han, Weike Jin, Nian Xie, James T. Kwok, Hengshuang Zhao, Xiaodan Liang, Dit-Yan Yeung, Zhenguo Li, Qun Liu 0001, Lanqing Hong, Lu Hou 0002 |
CVPR | 4 |
| 2025 | Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction TuningabstractYunhao Gou, Hansi Yang, Zhili Liu, Kai Chen, Yihan Zeng, Lanqing Hong, Zhenguo Li, Qun Liu, Bo Han, James Kwok, Yu Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yunhao Gou, Hansi Yang, Zhili Liu, Kai Chen 0023, Yihan Zeng, Lanqing Hong, Zhenguo Li, Qun Liu 0001, Bo Han 0003, James T. Kwok, Yu Zhang 0006 |
EMNLP | 3 |
| 2025 | Image-level Memorization Detection via Inversion-based Inference PerturbationabstractRecent studies have discovered that widely used text-to-image diffusion models can replicate training samples during image generation, a phenomenon known as memorization. Existing detection methods primarily focus on identifying memorized prompts. However, in real-world scenarios, image owners may need to verify whether their proprietary or personal images have been memorized by the model, even in the absence of paired prompts or related metadata. We refer to this challenge as image-level memorization detection, where current methods relying on original prompts fall short. In this work, we uncover two characteristics of memorized images after perturbing the inference procedure: lower similarity of the original images and larger magnitudes of TCNP.
Building on these insights, we propose Inversion-based Inference Perturbation (IIP), a new framework for image-level memorization detection. Our approach uses unconditional DDIM inversion to derive latent codes that contain core semantic information of original images and optimizes random prompt embeddings to introduce effective perturbation. Memorized images exhibit distinct characteristics within the proposed pipeline, providing a robust basis for detection. To support this task, we construct a comprehensive setup for the image-level memorization detection, carefully curating datasets to simulate realistic memorization scenarios. Using this setup, we evaluate our IIP framework across three different memorization settings, demonstrating its state-of-the-art performance in identifying memorized images in various settings, even in the presence of data augmentation attacks. Haokun Lin, Bo Peng 0002, Zhili Liu, Yueming Lyu, Xing Zheng, Jing Dong 0003 |
ICLR | 5 |
| 2025 | TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion ModelsabstractDespite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world modeling, compounded by the necessity to manage appearance and disappearance, drastic scale changes, and ensure consistency for instances across frames. These challenges hinder the development of video generation that can faithfully mimic real-world complexity, limiting utility for applications requiring high-level realism and controllability, including advanced scene simulation and training of perception systems. To address that, we propose TrackDiffusion, a novel video generation framework affording fine-grained trajectory-conditioned motion control via diffusion models, which facilitates the precise manipulation of the object trajectories and interactions, overcoming the prevalent limitation of scale and continuity disruptions. A pivotal component of TrackDiffusion is the instance enhancer, which explicitly ensures inter-frame consistency of multiple objects, a critical factor overlooked in the current literature. More-over, we demonstrate that generated video sequences by our TrackDiffusion can be used as training data for visual per-ception models. To the best of our knowledge, this is the first work to apply video diffusion models with tracklet conditions and demonstrate that generated frames can be beneficial for improving the performance of object trackers. 1 Kai Chen 0023, Zhili Liu, Ruiyuan Gao 0001, Lanqing Hong, Dit-Yan Yeung, Huchuan Lu, Xu Jia 0012 |
WACV | 3 |
| 2025 | Location privacy protection scheme of user collaborative probabilistic indistinguishability based on Hyperledger Fabric
Lei Zhang 0052, Yongbo Bai, Shuaishuai Lian, Yijia Geng, Zhili Liu |
J. Supercomput. | 6 |
| 2024 | ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language ModelsabstractHaochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu, Zhili Liu, Yunlong Feng, Xiaoguang Li, Yasheng Wang, Lifeng Shang, Qun Liu, Linqi Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Haochen Tan, Zhijiang Guo, Zhan Shi 0001, Zhili Liu, Yunlong Feng, Yasheng Wang, Lifeng Shang, Qun Liu 0001, Linqi Song |
ACL (1) | 5 |
| 2024 | MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-Wise Pruning Error MetricabstractVision-language pretrained models have achieved impressive performance on various downstream tasks. However, their large model sizes hinder their utilization on platforms with limited computational resources. We find that directly using smaller pretrained models and applying magnitude-based pruning on CLIP models leads to in-flexibility and inferior performance. Recent efforts for VLP compression either adopt uni-modal compression metrics resulting in limited performance or involve costly mask-search processes with learnable masks. In this paper, we first propose the Module-wise Pruning Error (MoPE) met-ric, accurately assessing CLIP module importance by performance decline on cross-modal tasks. Using the MoPE metric, we introduce a unified pruning framework applica-ble to both pretraining and task-specific fine-tuning compression stages. For pretraining, MoPE-CLIP effectively leverages knowledge from the teacher model, significantly reducing pretraining costs while maintaining strong zero-shot capabilities. For fine-tuning, consecutive pruning from width to depth yields highly competitive task-specific models. Extensive experiments in two stages demonstrate the effectiveness of the MoPE metric, and MoPE-CLIP outperforms previous state-of-the-art VLP compression methods. Haokun Lin, Haoli Bai, Zhili Liu, Lu Hou 0002, Muyi Sun, Linqi Song, Ying Wei 0001, Zhenan Sun |
CVPR | 3 |
| 2024 | Eyes Closed, Safety on: Protecting Multimodal LLMs via Image-to-Text Transformation
Yunhao Gou, Kai Chen 0023, Zhili Liu, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung, James T. Kwok, Yu Zhang 0006 |
ECCV (17) | 3 |
| 2024 | Implicit Concept Removal of Diffusion Models
Zhili Liu, Kai Chen 0023, Jianhua Han, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung, James T. Kwok |
ECCV (21) | 1 |
| 2024 | Towards Robust Visual Localization Using Multi-View Images and HD Vector MapabstractRobust and accurate localization is highly desired in intelligent driving and robotic navigation. Existing methods highly rely on feature maps and complex parameter tuning, while suffering from ineffective data association, heavy computation, high dependency on training data and low robustness. In this paper, we propose a high-robust and cost-effective visual localization system, which jointly exploits the semantic information of Bird’s-Eye-View (BEV) representation from multi-view images and the vectorized High Definition (HD) map. We formulate the visual localization as cross-modal data association issue and innovatively project the vectorized landmarks of HD map into BEV semantic map. Finally, the highly accurate vehicle’s pose can be estimated by pose optimization based on direct image alignment. Extensive simulations experimented on nuScenes dataset show that the proposed method can deliver robust and accurate localization results under various scenarios. In addition, the proposed system is convenient for large-scale deployment and has been tested on the commercial test car. Lili Zhao 0001, Zhili Liu, Qian Yin 0002, Lei Yang 0063, Meng Guo 0007 |
ICIP | 2 |
| 2023 | Mixed Autoencoder for Self-Supervised Visual Representation LearningabstractMasked Autoencoder (MAE) has demonstrated superior performance on various vision tasks via randomly masking image patches and reconstruction. However, effective data augmentation strategies for MAE still remain open questions, different from those in contrastive learning that serve as the most important part. This paper studies the prevailing mixing augmentation for MAE. We first demonstrate that naïve mixing will in contrast degenerate model performance due to the increase of mutual information (MI). To address, we propose homologous recognition, an auxiliary pretext task, not only to alleviate the MI increasement by explicitly requiring each patch to recognize homologous patches, but also to perform object-aware self-supervised pre-training for better downstream dense perception performance. With extensive experiments, we demonstrate that our proposed Mixed Autoencoder (MixedAE) achieves the state-of-the-art transfer results among masked image modeling (MIM) augmentations on different downstream tasks with significant efficiency. Specifically, our MixedAE outperforms MAE by +0.3% accuracy, +1.7 mIoU and +0.9 AP on ImageNet-1K, ADE20K and COCO respectively with a standard ViT-Base. Moreover, MixedAE surpasses iBOT, a strong MIM method combined with instance discrimination, while accelerating training by 2×. To our best knowledge, this is the very first work to consider mixing for MIM from the perspective of pretext task design. Code will be made available. Kai Chen 0023, Zhili Liu, Lanqing Hong, Hang Xu 0004, Zhenguo Li, Dit-Yan Yeung |
CVPR | 2 |
| 2023 | DiffFit: Unlocking Transferability of Large Diffusion Models via Simple Parameter-Efficient Fine-TuningabstractDiffusion models have proven to be highly effective in generating high-quality images. However, adapting large pre-trained diffusion models to new domains remains an open challenge, which is critical for real-world applications. This paper proposes DiffFit, a parameter-efficient strategy to fine-tune large pre-trained diffusion models that enable fast adaptation to new domains. DiffFit is embarrassingly simple that only fine-tunes the bias term and newly-added scaling factors in specific layers, yet resulting in significant training speed-up and reduced model storage costs. Compared with full fine-tuning, DiffFit achieves 2× training speed-up and only needs to store approximately 0.12% of the total model parameters. Intuitive theoretical analysis has been provided to justify the efficacy of scaling factors on fast adaptation. On 8 downstream datasets, DiffFit achieves superior or competitive performances compared to the full fine-tuning while being more efficient. Remarkably, we show that DiffFit can adapt a pre-trained low-resolution generative model to a high-resolution one by adding minimal cost. Among diffusion-based methods, DiffFit sets a new state-of-the-art FID of 3.02 on ImageNet 512×512 benchmark by fine-tuning only 25 epochs from a public pre-trained ImageNet 256×256 checkpoint while being 30× more training efficient than the closest competitor. Enze Xie, Lewei Yao, Zhili Liu, Daquan Zhou, Zhaoqiang Liu, Zhenguo Li |
ICCV | 4 |
| 2023 | Your Contrastive Learning Is Secretly Doing Stochastic Neighbor Embedding
Tianyang Hu 0001, Zhili Liu, Fengwei Zhou, Weiran Huang 0001 |
ICLR | 2 |
| 2023 | Task-customized Masked Autoencoder via Mixture of Cluster-conditional Experts
Zhili Liu, Kai Chen 0023, Jianhua Han, Lanqing Hong, Hang Xu 0004, Zhenguo Li, James T. Kwok |
ICLR | 1 |
| 2023 | Structured Term Pruning for Computational Efficient Neural Networks InferenceabstractThe state-of-the-art convolutional neural network accelerators are showing a growing interest in exploiting the bit-level sparsity and eliminating the ineffectual computations of zero bits. However, the excessive redundancy and the irregular distribution of nonzero bits limit the real speedup in the accelerators. To address this, we propose an algorithm-architecture codesign, named structured term pruning (STP), to boost the computation efficiency of neural networks inference. Specifically, we enhance the bit sparsity by guiding the weights toward the value with fewer power-of-two terms. Then, we structure the terms with layer-wise group budgets. Retraining is adopted to recover the accuracy drop. We also design the hardware of the group processing element and the fast signed-digital encoder for efficient implementation of STP networks. The system design of STP is realized with some easy alterations on an input stationary systolic array design. Extensive evaluation results demonstrate that STP can reduce significant inference computation costs, and achieve$2.35\times $computational energy saving for the ResNet18 network on the ImageNet dataset. Kai Huang 0002, Bowen Li 0017, Siang Chen, Luc Claesen, Wei Xi 0001, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong, Xiaolang Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | Structured Dynamic Precision for Deep Neural Networks QuantizationabstractDeep Neural Networks (DNNs) have achieved remarkable success in various Artificial Intelligence applications. Quantization is a critical step in DNNs compression and acceleration for deployment. To further boost DNN execution efficiency, many works explore to leverage the input-dependent redundancy with dynamic quantization for different regions. However, the sensitive regions in the feature map are irregularly distributed, which restricts the real speed up for existing accelerators. To this end, we propose an algorithm-architecture co-design, named Structured Dynamic Precision (SDP). Specifically, we propose a quantization scheme in which the high-order bit part and the low-order bit part of data can be masked independently. And a fixed number of term parts are dynamically selected for computation based on the importance of each term in the group. We also present a hardware design to enable the algorithm efficiently with small overheads, whose inference time mainly scales with the precision proportionally. Evaluation experiments on extensive networks demonstrate that compared to the state-of-the-art dynamic quantization accelerator DRQ, our SDP can achieve 29% performance gain and 51% energy reduction for the same level of model accuracy. Kai Huang 0002, Bowen Li 0017, Dongliang Xiong, Haitian Jiang, Xiaowen Jiang 0001, Xiaolang Yan, Luc Claesen, Dehong Liu, Junjian Chen, Zhili Liu |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2022 | Task-Customized Self-Supervised Pre-training with Scalable Dynamic RoutingabstractSelf-supervised learning (SSL), especially contrastive methods, has raised attraction recently as it learns effective transferable representations without semantic annotations. A common practice for self-supervised pre-training is to use as much data as possible. For a specific downstream task, however, involving irrelevant data in pre-training may degenerate the downstream performance, observed from our extensive experiments. On the other hand, for existing SSL methods, it is burdensome and infeasible to use different downstream-task-customized datasets in pre-training for different tasks. To address this issue, we propose a novel SSL paradigm called Scalable Dynamic Routing (SDR), which can be trained once and deployed efficiently to different downstream tasks with task-customized pre-trained models. Specifically, we construct the SDRnet with various sub-nets and train each sub-net with only one subset of the data by data-aware progressive training. When a downstream task arrives, we route among all the pre-trained sub-nets to get the best along with its corresponding weights. Experiment results show that our SDR can train 256 sub-nets on ImageNet simultaneously, which provides better transfer performance than a unified model trained on the full ImageNet, achieving state-of-the-art (SOTA) averaged accuracy over 11 downstream classification tasks and AP on PASCAL VOC detection task. Zhili Liu, Jianhua Han, Lanqing Hong, Hang Xu 0004, Kai Chen 0023, Chunjing Xu, Zhenguo Li |
AAAI | 1 |
| 2022 | Structured precision skipping: Accelerating convolutional neural networks with budget-aware dynamic precision selection
Kai Huang 0002, Siang Chen, Bowen Li 0017, Luc Claesen, Hao Yao, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong |
J. Syst. Archit. | 8 |
| 2022 | Acceleration-Aware Fine-Grained Channel Pruning for Deep Neural Networks via Residual GatingabstractDeep neural networks have achieved remarkable advancement in various intelligence tasks. However, the massive computation and storage consumption limit applications on resource-constrained devices. While channel pruning has been widely applied to compress models, it is challenging to reach very deep compressions for such a coarse-grained pruning structure without significant performance degradation. In this article, we propose an acceleration-aware fine-grained channel pruning (AFCP) framework for accelerating neural networks, which optimizes trainable gate parameters by estimating residual errors between pruned and original channels with hardware characteristics. Our fine-grained concept consists of both algorithm and structure levels. Different from existing methods that leverage a predefined pruning criterion, AFCP explicitly considers both zero-out and similar criteria for each channel, and adaptively selects the suitable one via residual gate parameters. For structure level, AFCP adopts a fine-grained channel pruning strategy for residual neural networks and a decomposition-based structure, which further extends the pruning optimization space. Moreover, instead of using theoretical computation costs, such as floating-point operations, we propose the hardware predictor that bridges the gap between realistic acceleration and pruning procedure to guide the learning of pruning, which improves the efficiency of model pruning when deployed on accelerators. Extensive evaluation results demonstrate that AFCP outperforms state-of-the-art methods, and achieves a favorable balance between model performance and computation cost. Kai Huang 0002, Siang Chen, Bowen Li 0017, Luc Claesen, Hao Yao, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2022 | Real-Time Scene-Aware LiDAR Point Cloud Compression Using Semantic Prior RepresentationabstractExisting LiDAR point cloud compression (PCC) methods tend to treat compression as afidelityissue, without sufficiently addressing itsmachine perceptionaspect. The latter issue is often encountered by the decoder agents that might aim to conduct scene-understanding related tasks only, such as computing the localization information. For tackling this challenge, a novel LiDAR PCC system is proposed to compress the point cloud geometry, which contains aback channelfor allowing the decoder to initiate such request to the encoder. The key success of our PCC method lies in our proposedsemantic prior representation(SPR) and its lossy encoding algorithm with variable precision to generate the final bitstream; the entire process is fast and achieves real-time performance. Note that our SPR is a compact and effective representation of three-dimensional (3D) input point clouds, and it consists oflabels, predictions, andresiduals. These information can be generated by first exploiting ascene-aware object segmentationto a set of 2D range images (frames) individually, which were generated from the 3D point clouds via a projection process. Based on the generated labels, the pixels associated with those moving objects are considered as noisy information and should be removed for not only saving bit budget on transmission but also, most importantly, improving the accuracy of localization computed at the decoder. Experimental results conducted on the commonly-used test dataset have shown that our proposed system outperforms the MPEG’s G-PCC (TMC13-v14.0) in a large bitrate range. In fact, the performance gap will become even larger when more and/or large moving objects are involved in the input point clouds. Lili Zhao 0001, Kai-Kuang Ma, Zhili Liu, Qian Yin 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | EHSOD: CAM-Guided End-to-End Hybrid-Supervised Object Detection with Cascade RefinementabstractObject detectors trained on fully-annotated data currently yield state of the art performance but require expensive manual annotations. On the other hand, weakly-supervised detectors have much lower performance and cannot be used reliably in a realistic setting. In this paper, we study the hybrid-supervised object detection problem, aiming to train a high quality detector with only a limited amount of fully-annotated data and fully exploiting cheap data with image-level labels. State of the art methods typically propose an iterative approach, alternating between generating pseudo-labels and updating a detector. This paradigm requires careful manual hyper-parameter tuning for mining good pseudo labels at each round and is quite time-consuming. To address these issues, we present EHSOD, an end-to-end hybrid-supervised object detection system which can be trained in one shot on both fully and weakly-annotated data. Specifically, based on a two-stage detector, we proposed two modules to fully utilize the information from both kinds of labels: 1) CAM-RPN module aims at finding foreground proposals guided by a class activation heat-map; 2) hybrid-supervised cascade module further refines the bounding-box position and classification with the help of an auxiliary head compatible with image-level data. Extensive experiments demonstrate the effectiveness of the proposed method and it achieves comparable results on multiple object detection benchmarks with only 30% fully-annotated data, e.g. 37.5% mAP on COCO. We will release the code and the trained models. Linpu Fang, Hang Xu 0004, Zhili Liu, Sarah Parisot, Zhenguo Li |
AAAI | 3 |
| 2018 | Application of Photon Recollision Probability Theory for Compatibility Check Between Foliage Clumping and Leaf Area Index Products Obtained from Earth Observation DataabstractClumping index (CI) is a measure of foliage aggregation relative to a random distribution of leaves in space. The CI can help with estimating fractions of sunlit and shaded leaves for a given value of leaf area index (LAI). Both the CI and LAI can be obtained from global Earth Observing (EO) systems such as the Moderate Resolution Imaging Spectrometer (MODIS). Here, the compatibility between CI and LAI products derived from EO data is examined independently using the theory of spectral invariants, also referred to as photon recollision probability theory (i.e. ` p-theory'), along with raw LAI-2000/2200 Plant Canopy Analyzer data from 75 sites distributed across a range of plant functional types (PFTs). The p-theory describes the probability (p-value) that a photon, having intercepted an element in the canopy, will recollide with another canopy element rather than escape the canopy. Our results indicate that the integration of empirically-based CI maps with the MODIS LAI product is feasible, providing a potential means to improve the accuracy of LAI EO data products. Given the strong results for the large range of PFTs explored here, we demonstrate the capacity to obtain p-values for any location solely from EO data. This is relevant for future applications of the photon recollision probability concept for global and local monitoring of vegetation using EO data. Jan Pisek, Henning Buddenbaum, Fernando Camacho, Joachim Hill, Jennifer L. R. Jensen, Holger Lange, Zhili Liu, Arndt Piayda, Yonghua Qu, Olivier Roupsard, Shawn P. Serbin, Svein Solberg, Oliver Sonnentag, Anne Thimonier, Francesco Vuolo |
IGARSS | 7 |
| 2003 | Comparison analysis of AVHRR albedo temporal changes and dust TSP dataabstractChinese and Japanese researchers established a joint project in 2000 to set up ground observation stations along dust source areas, transportation roads, and precipitation areas by collecting TSP (dry dust precipitation) and AVHRR data to retrieve albedo (surface energy). The selected data from retrieved albedo temporal imagery are used to construct LST/TSP/albedo curves. Finally a comparison was made between albedo curves and TSP curves. The result showed that there were good correlation between these two kinds of curves. It was proved that the albedo could be one of the physical parameters for predicting dust storm in future monitoring systems. Xiuzhen Han, Jianwen Ma, Zhili Liu, Hasibagan, Qiqing Li |
IGARSS | 3 |
| 2003 | Spectral and spatial feature integrated method for edge information extraction from high resolution remote sensing imageabstractWe introduce a four-stage process for urban construction edge detection using IKONOS images. The four stages include: (1) binary image processing, (2) pixel swapping by using different kennels to separate pixels according to the gray level of the image, in this paper 8 adjacent kennels are used, (3) interactive analyzing and selecting numbers representing edges, (4) using many edge images based on image gray level, in this paper we introduce three gray level edge detection processes. Qiqing Li, Jianwen Ma, Hasibagan, Xiuzhen Han, Zhili Liu |
IGARSS | 5 |
| 2003 | Comprehensive analysis of in situ measurement data and satellite data during dust storm in spring 2002abstractIn 2001, some in situ measurement stations were set up by the Asia dust storm project, ADEC, in Korea, Japan and Xinjiang, Gansu, Shanxi, Neimenggu, Beijing, and Qindao in China. Many instruments were fixed in those places, for example an aerosol detector, dust particle collector, wind speed measuring apparatus, and so on. Combined with remote sensing technology, the instruments were used to measure the emission transport and deposition of the dust storm. In April and May 2002, through these instruments, the wind speed, total suspended particulate matter and land surface temperature data were obtained. In this article, taking the dust storm in March and April 2002 as an example and combining satellite data with in situ measurement data, some basic data may be provided for the general analysis and forecast of dust storms. Zhili Liu, Jianwen Ma, Xiaoye Zhang, Wanghong, Buheaosir, Anjin |
IGARSS | 1 |
| 2003 | The endangered rare plant coverage change detection in twelve years by using TM/ETM dataabstractThe endangered rare plants, named Tetraena mongolica, Helianthemum soongolicum, Reaumuria trigyna, Reaumuria soongorica, Potaninia mongolica, Ammopiptanthus mongolicus, are distributed in West Ordos Plateau National Protected Region in Inner Mongolia, China. During last ten years coverage area of the plants changed a lot. Recently two periods TM data were used to detect the coverage changes of the plants which provides some new evidence for current situation of the plants. Jianwen Ma, Qiqing Li, Hasi Bagan, Zhili Liu |
IGARSS | 4 |
| 2002 | The use of wavelet fusion method to improve multi-spectral imagery for land cover change monitoringabstractIHS transform was one of typical method for remote sensing data fusion. In recent years, newly developed method taking the advantage of IHS and Wavelet algorithms makes image fusion. In this case after the Wavelet substitution based on pixels or features, and then transforms inversely. In this paper we introduces a high frequency substitution method to improve spatial resolutions. The procedure of the method introduced as flowchart, in which the dot line area is our newly added method. The result was used in making 1:50000 scale NDVI imagery for monitoring land cover change in Minjiang River, Sichuan province, China and for providing information monitoring of Return Farmland Back to Forest or Grassland Project. Jianwen Ma, Hasi Bagan, Chaofei Ma, Xiuzhen Han, Zhili Liu |
IGARSS | 5 |