EDBT 2026 Demo / reviewers in the wild / expert
Dacheng Li
dblp:61/2812
· DBLP profile ↗
23ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FDCFusion: frequency-domain CAFormer-CNN fusion network for infrared and visible image fusion
Huifang Kong, Zixiang Dong, Dacheng Li |
Pattern Anal. Appl. | 3 |
| 2025 | NVILA: Efficient Frontier Visual Language ModelsabstractVisual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention. This paper introduces NVILA, a family of open VLMs designed to optimize both efficiency and accuracy. Building on top of VILA, we improve its model architecture by first scaling up the spatial and temporal resolutions, and then compressing visual tokens. This "scale-then-compress" approach enables NVILA to efficiently process high-resolution images and long videos. We also conduct a systematic investigation to enhance the efficiency of NVILA throughout its entire lifecycle, from training to deployment. NVILA matches or surpasses the accuracy of many leading open and proprietary VLMs across a wide range of image and video benchmarks. At the same time, it reduces training costs by 1.9-5.1×, prefilling latency by 1.6-2.2×, and decoding latency by 1.2-2.8×. Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Haotian Tang, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Jinyi Hu, Sifei Liu, Ranjay Krishna, Pavlo Molchanov 0001, Jan Kautz, Hongxu Yin, Song Han 0003, Yao Lu 0006 |
CVPR | 10 |
| 2025 | LongVILA: Scaling Long-Context Visual Language Models for Long VideosabstractLong-context capability is critical for multi-modal foundation models, especially for long video understanding. We introduce LongVILA, a full-stack solution for long-context visual-language models by co-designing the algorithm and system. For model training, we upgrade existing VLMs to support long video understanding by incorporating two additional stages, i.e., long context extension and long video supervised fine-tuning. However, training on long video is computationally and memory intensive. We introduce the long-context Multi-Modal Sequence Parallelism (MM-SP) system that efficiently parallelizes long video training and inference, enabling 2M context length training on 256 GPUs without any gradient checkpointing. LongVILA efficiently extends the number of video frames of VILA from 8 to 2048, achieving 99.8% accuracy in 6,000-frame (more than 1 million tokens) video needle-in-a-haystack. LongVILA-7B demonstrates strong accuracy on 9 popular video benchmarks, e.g., 65.1% VideoMME with subtitle. Besides, MM-SP is 2.1x - 5.7x faster than ring style sequence parallelism and 1.1x - 1.4x faster than Megatron with a hybrid context and tensor parallelism. Moreover, it seamlessly integrates with Hugging Face Transformers. Yukang Chen, Fuzhao Xue, Dacheng Li, Qinghao Hu 0004, Ligeng Zhu, Xiuyu Li, Yunhao Fang, Haotian Tang, Shang Yang, Yihui He, Hongxu Yin, Pavlo Molchanov 0001, Jan Kautz, Linxi Fan, Yuke Zhu, Yao Lu 0006, Song Han 0003 |
ICLR | 3 |
| 2025 | VILA-U: a Unified Foundation Model Integrating Visual Understanding and GenerationabstractVILA-U is a Unified foundation model that integrates Video, Image, Language understanding and generation. Traditional visual language models (VLMs) use separate modules for understanding and generating visual content, which can lead to misalignment and increased complexity. In contrast, VILA-U employs a single autoregressive next-token prediction framework for both tasks, eliminating the need for additional components like diffusion models. This approach not only simplifies the model but also achieves near state-of-the-art performance in visual language understanding and generation. The success of VILA-U is attributed to two main factors: the unified vision tower that aligns discrete visual tokens with textual inputs during pretraining, which enhances visual perception, and autoregressive image generation can achieve similar quality as diffusion models with high-quality dataset. This allows VILA-U to perform comparably to more complex models using a fully token-based autoregressive framework. Yecheng Wu, Zhuoyang Zhang, Junyu Chen 0003, Haotian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi 0001, Song Han 0003, Yao Lu 0006 |
ICLR | 5 |
| 2025 | SORRY-Bench: Systematically Evaluating Large Language Model Safety RefusalabstractEvaluating aligned large language models' (LLMs) ability to recognize and reject unsafe user requests is crucial for safe, policy-compliant deployments. Existing evaluation efforts, however, face three limitations that we address with **SORRY-Bench**, our proposed benchmark. **First**, existing methods often use coarse-grained taxonomies of unsafe topics, and are over-representing some fine-grained topics. For example, among the ten existing datasets that we evaluated, tests for refusals of self-harm instructions are over 3x less represented than tests for fraudulent activities. SORRY-Bench improves on this by using a fine-grained taxonomy of 44 potentially unsafe topics, and 440 class-balanced unsafe instructions, compiled through human-in-the-loop methods. **Second**, evaluations often overlook the linguistic formatting of prompts, like different languages, dialects, and more --- which are only implicitly considered in many evaluations. We supplement SORRY-bench with 20 diverse linguistic augmentations to systematically examine these effects. **Third**, existing evaluations rely on large LLMs (e.g., GPT-4) for evaluation, which can be computationally expensive. We investigate design choices for creating a fast, accurate automated safety evaluator. By collecting 7K+ human annotations and conducting a meta-evaluation of diverse LLM-as-a-judge designs, we show that fine-tuned 7B LLMs can achieve accuracy comparable to GPT-4 scale LLMs, with lower computational cost. Putting these together, we evaluate over 50 proprietary and open-weight LLMs on SORRY-Bench, analyzing their distinctive safety refusal behaviors. We hope our effort provides a building block for systematic evaluations of LLMs' safety refusal capabilities, in a balanced, granular, and efficient manner. Benchmark demo, data, code, and models are available through [https://sorry-bench.github.io](https://sorry-bench.github.io). Tinghao Xie, Xiangyu Qi, Yi Zeng 0005, Yangsibo Huang, T. W. U. Madhushani, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng 0007, Ruoxi Jia 0001, Bo Li 0026, Kai Li 0001, Danqi Chen 0001, Peter Henderson 0002, Prateek Mittal |
ICLR | 9 |
| 2025 | Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal SparsityabstractDiffusion Transformers (DiTs) dominate video generation but their high computational cost severely limits real-world applicability, usually requiring tens of minutes to generate a few seconds of video even on high-performance GPUs. This inefficiency primarily arises from the quadratic computational complexity of 3D full attention with respect to the context length. In this paper, we propose a training-free framework termed Sparse VideoGen (SVG) that leverages the inherent sparsity in 3D full attention to boost inference efficiency. We reveal that the attention heads can be dynamically classified into two groups depending on distinct sparse patterns: (1) Spatial Head, where only spatially-related tokens within each frame dominate the attention output, and (2) Temporal Head, where only temporally-related tokens across different frames dominate. Based on this insight, SVG proposes an online profiling strategy to capture the dynamic sparse patterns and predicts the type of attention head. Combined with a novel hardware-efficient tensor layout transformation and customized kernel implementations, SVG achieves up to 2.28$\times$ and 2.33$\times$ end-to-end speedup on CogVideoX-v1.5 and HunyuanVideo, respectively, while preserving generation quality. Our code will be open-sourced upon publication. Haocheng Xi, Shuo Yang 0011, Yilong Zhao 0002, Chenfeng Xu, Xiuyu Li, Yujun Lin 0001, Han Cai, Dacheng Li, Jianfei Chen 0001, Ion Stoica, Kurt Keutzer, Song Han 0003 |
ICML | 10 |
| 2025 | WorldModelBench: Judging Video Generation Models As World ModelsabstractVideo generation models have rapidly progressed, positioning themselves as video world models capable of supporting decision-making applications like robotics and autonomous driving. However, current benchmarks fail to rigorously evaluate these claims, focusing only on general video quality, ignoring important factors to world models such as physics adherence.To bridge this gap, we propose WorldModelBench, a benchmark designed to evaluate the world modeling capabilities of video generation models in application-driven domains. WorldModelBench offers two key advantages: (1) Against to nuanced world modeling violations: By incorporating instruction-following and physics-adherence dimensions, WorldModelBench detects subtle violations, such as irregular changes in object size that breach the mass conservation law—issues overlooked by prior benchmarks. (2) Aligned with large-scale human preferences: We crowd-source 67K human labels to accurately measure 14 frontier models. Using our high-quality human labels, we further fine-tune an accurate judger to automate the evaluation procedure, achieving 9.9% lower error in predicting world modeling violations than GPT-4o with 2B parameters. In addition, we demonstrate that training to align human annotations by maximizing the rewards from the judger noticeably improve the world modeling capability. The dataset is hosted in HuggingFace at https://huggingface.co/datasets/Efficient-Large-Model/worldmodelbench. The code to run evaluation is available at https://github.com/WorldModelBench-Team/WorldModelBench. Dacheng Li, Yunhao Fang, Yukang Chen, Shuo Yang 0011, Shiyi Cao, Justin Wong, Michael Luo, Xiaolong Wang 0004, Hongxu Yin, Joseph Gonzalez 0001, Ion Stoica, Song Han 0003, Yao Lu 0006 |
NeurIPS | 1 |
| 2025 | Land Surface Emissivity Retrieval From Landsat 9 Data in Combination With Land Cover Data and Spectral LibraryabstractLand surface emissivity (LSE) is crucial for retrieving land surface temperature (LST) from Landsat 9 TIRS-2 thermal infrared data. However, the single-band LSE product (band 10) provided officially is insufficient for split-window (SW) algorithm requiring dual-band emissivity inputs. This letter proposes a land cover and channel transformed (LCCT-LSE) method to estimate band 11 LSE and enables LST retrieval using SW algorithm on Google Earth Engine. Cross-validation with MOD21 LSE products showed that the LCCT-LSE method achieved a mean absolute error (MAE) of 0.004 and a root mean square error (RMSE) of 0.005, outperforming classification-based method, NDVI threshold method, and vegetation cover (VCM) methods. In situ validation showed SW-retrieved LST attains MAE/RMSE of 1.27 K/ 2.13 K, with consistent accuracy across diverse land covers (water: 0.86 K, soil: 1.58 K, desert: 1.71 K, sand: 1.80 K, vegetation: 0.87 K). A comparison with the official Landsat 9 LST product indicated that the bias of retrieved LST is within 1K for all land cover classes (cropland, forest, grassland, shrubland, water, barren and impervious) in Beijing. These results demonstrated that the LCCT-LSE method is capable of estimating the LSE in Landsat 9 band 11 with a reliable and accurate result. This study provides a new insight for LST retrieval from Landsat 9 data. Qi Zhang 0084, Yonggang Qian, Kun Li 0019, Qiyao Li, Dacheng Li |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceabstractLarge Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To address this issue, we introduce Chatbot Arena, an open platform for evaluating LLMs based on human preferences. Our methodology employs a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing. The platform has been operational for several months, amassing over 240K votes. This paper describes the platform, analyzes the data we have collected so far, and explains the tried-and-true statistical methods we are using for efficient and accurate evaluation and ranking of models. We confirm that the crowdsourced questions are sufficiently diverse and discriminating and that the crowd-sourced human votes are in good agreement with those of expert raters. These analyses collectively establish a robust foundation for the credibility of Chatbot Arena. Because of its unique value and openness, Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies. The platform is publicly available at https://chat.lmsys.org. Wei-Lin Chiang, Lianmin Zheng, Ying Sheng 0007, Anastasios Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang 0025, Michael I. Jordan, Joseph Gonzalez 0001, Ion Stoica |
ICML | 6 |
| 2024 | Fairness in Serving Large Language Models
Ying Sheng 0007, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li 0001, Danyang Zhuo, Joseph Gonzalez 0001, Ion Stoica |
OSDI | 3 |
| 2024 | 3D Detection Under Foggy Weather: The Effective Multi-Sensor Fusion of LiDAR and RadarabstractAdverse weather conditions are a significant challenge for fully automated driving systems in autonomous vehicles. While radar sensor integration can alleviate some of the adverse effects of weather, the inherent sparsity and altitude discrepancies in radar data remain significant concerns. In response to these challenges, we introduce an innovative sensor fusion method for LiDAR and radar data under foggy weather conditions. This method effectively augments radar data in voxel space. It facilitates the synchronization and complementation of Li-DAR and radar features in the bird's eye view space, thereby achieving high-precision 3D object detection. The effectiveness of this method has been extensively validated through extensive experimentation on the Oxford Radar RobotCar dataset under foggy weather conditions. Jiake Tian, Dacheng Li, Yi Zou 0001, Jiale Lai, Zufeng Liang |
SECON | 2 |
| 2024 | mmHPE: Human Pose Estimation Based on Point Cloud from Millimeter-wave RadarabstractRehabilitation therapy involving repetitive exer-cises targeting specific human joints under the supervision of a doctor is crucial for patients with movement disorders. Yet the cost of commuting and the demand for medical resources are inconvenient for patients. Human-computer interaction can provide remote rehabilitation guidance for patients at home through human pose estimation technology, but privacy concerns with optical sensors and the cost and discomfort of wearable sensors have hindered progress in this field. To address the above challenges, we propose mmHPE, an innovative 3D human pose estimation framework that uses millimeter-wave (mmWave) radar sensors. It initially manipulates the raw data captured from radar sensors to generate a spatiotemporal sequence point cloud dataset. Afterward, we create a Convolutional Neural Network (CNN) that is linked to a Bidirectional Long Short-Term Memory (Bi-LSTM) network. Moreover, a multi-head attention mechanism is employed to boost the network's performance and to accurately estimate the locations of human skeletons. Ultimately, the 21 points with the corre-sponding human pose position are successfully reconstructed. We investigate the mmHPE framework's feasibility and cross-domain stability in different home environments in real-world scenarios. This innovation proffers patients a convenient and privacy-conscious solution for their rehabilitation training req-uisites. Jiale Lai, Jiake Tian, Yi Zou 0001, Xianfeng Song, Fangming Liu, Dacheng Li |
SMC | 6 |
| 2024 | A novel reinforcement learning based Heap-based optimizer
Xuesen Ma, Zhineng Zhong, Yangyu Li, Dacheng Li, Yan Qiao 0001 |
Knowl. Based Syst. | 4 |
| 2024 | Vessel Target Detection Method Based on an Improved CFAR Method in Nighttime Remote Sensing ImagesabstractThe Visible Infrared Imaging Radiometer Suite (VIIRS) day/night band (DNB) data are very sensitive to low radiation and capable of detecting faint light sources emitted by vessels at night. Existing vessel detection methods are mainly based on pixel statistical feature estimation or manual experience to determine the threshold for detection. In complex scenes, problems such as interference targets and inaccurate background modeling result in inaccurate detection accuracy and poor robustness. In this article, a vessel detection method based on an improved constant false alarm rate (CFAR) detector is proposed, and more accurate threshold selection is obtained through background sample truncation and distribution adjustment. Then, the local peak detection algorithm removes the bright spots in the nonvessel location and realizes the vessel detection. We selected data from two research areas, the waters around the East China Sea and the waters around the Gulf of Mexico of the United States, to build a vessel dataset and applied the algorithm to the dataset. The experimental results show that the total detection accuracy of the algorithm is 94.48%, and the recall rate is 93.40%, which is higher than the other two comparison methods. Especially in complex scenarios with high target density, such as ports, the recall rate is significantly improved, which proves the applicability of the algorithm. Weiyuan Yao, Zhaoyan Liu, Xinhong Wang, Feihong Wang, Hongjia Cheng, Xiqing Guo, Dacheng Li |
IEEE Trans. Geosci. Remote. Sens. | 12 |
| 2023 | MPCFORMER: Fast, Performant and Provate Transformer Inference with MPC
Dacheng Li, Hongyi Wang 0001, Rulin Shao, Eric P. Xing, Hao Zhang 0025 |
ICLR | 1 |
| 2023 | DeFusion: Aerial Image Matching Based on Fusion of Handcrafted and Deep Features
Xianfeng Song, Dacheng Li |
ICONIP (14) | 5 |
| 2023 | Comparisons of Three Single-Channel Algorithms for Retrieving Land Surface Temperature from HJ-1B Satellite DataabstractHJ-1B Infrared Scanner (IRS) has accumulated long-term data, but there are no comparative studies on land surface temperature (LST) inversion algorithms for IRS data. This study compared the radiative transfer equation (RTE) and two generalized single-channel algorithms (GSCwand GSCwT) for retrieving LST from IRS data. Firstly, the coefficients of the GSCwand the GSCwTalgorithms were simulated using MODTRAN 5.2 and the TIGR atmospheric profiles. ERA5 atmospheric profiles were used for the RTE algorithm. Second, land surface emissivity was calculated using the ASTER global emissivity dataset and vegetation/snow cover products based on the vegetation cover method. Finally, the LST retrievals were evaluated using ground measurements of twelve sites during the Heihe Watershed Allied Telemetry Experimental Research (HiWATER) experiment from June 2012 to June 2014. The results showed that the accuracy of the RTE with bias (RMSE) of 0.74 K (2.47 K) is superior to the accuracy of the GSCwT with bias (RMSE) of -1.18 K (2.50 K), followed by the GSCwwith bias (RMSE) of 1.60 K (2.77 K). The results can support the production of HJ-1B/IRS LST products. Guoqin Zhang, Dacheng Li, Hua Li 0005, Hui Jia, Yue Lyu |
IGARSS | 2 |
| 2023 | Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaabstractEvaluating large language model (LLM) based chat assistants is challenging due to their broad capabilities and the inadequacy of existing benchmarks in measuring human preferences.To address this, we explore using strong LLMs as judges to evaluate these models on more open-ended questions.We examine the usage and limitations of LLM-as-a-judge, including position, verbosity, and self-enhancement biases, as well as limited reasoning ability, and propose solutions to mitigate some of them.We then verify the agreement between LLM judges and human preferences by introducing two benchmarks: MT-bench, a multi-turn question set; and Chatbot Arena, a crowdsourced battle platform.Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80\% agreement, the same level of agreement between humans.Hence, LLM-as-a-judge is a scalable and explainable way to approximate human preferences, which are otherwise very expensive to obtain.Additionally, we show our benchmark and traditional benchmarks complement each other by evaluating several variants of LLaMA and Vicuna.The MT-bench questions, 3K expert votes, and 30K conversations with human preferences are publicly available at https://github.com/lm-sys/FastChat/tree/main/fastchat/llm_judge. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng 0007, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang 0001, Zi Lin, Zhuohan Li 0001, Dacheng Li, Eric P. Xing, Hao Zhang 0025, Joseph Gonzalez 0001, Ion Stoica |
NeurIPS | 9 |
| 2022 | AMP: Automatically Finding Model Parallel Strategies with Heterogeneity AwarenessabstractScaling up model sizes can lead to fundamentally new capabilities in many machine learning (ML) tasks. However, training big models requires strong distributed system expertise to carefully design model-parallel execution strategies that suit the model architectures and cluster setups. In this paper, we develop AMP, a framework that automatically derives such strategies. AMP identifies a valid space of model parallelism strategies and efficiently searches the space for high-performed strategies, by leveraging a cost model designed to capture the heterogeneity of the model and cluster specifications. Unlike existing methods, AMP is specifically tailored to support complex models composed of uneven layers and cluster setups with more heterogeneous accelerators and bandwidth. We evaluate AMP on popular modelsand cluster setups from public clouds and show that AMP returns parallel strategies that match the expert-tuned strategies on typical cluster setups. On heterogeneous clusters or models with heterogeneous architectures, AMP finds strategies with 1.54$\times$ and 1.77$\times$ higher throughput than state-of-the-art model-parallel systems, respectively. Dacheng Li, Hongyi Wang 0001, Eric P. Xing, Hao Zhang 0025 |
NeurIPS | 1 |
| 2021 | Dual Contradistinctive Generative AutoencoderabstractWe present a new generative autoencoder model with dual contradistinctive losses to improve generative autoencoder that performs simultaneous inference (reconstruction) and synthesis (sampling). Our model, named dual contradistinctive generative autoencoder (DC-VAE), integrates an instance-level discriminative loss (maintaining the instance-level fidelity for the reconstruction/synthesis) with a set-level adversarial loss (encouraging the set-level fidelity for the reconstruction/synthesis), both being contradistinctive. Extensive experimental results by DC-VAE across different resolutions including 32×32, 64×64, 128×128, and 512×512 are reported. The two contradistinctive losses in VAE work harmoniously in DC-VAE leading to a significant qualitative and quantitative performance enhancement over the baseline VAEs without architectural changes. State-of-the-art or competitive results among generative autoencoders for image reconstruction, image synthesis, image interpolation, and representation learning are observed. DC-VAE is a general-purpose VAE model, applicable to a wide variety of downstream tasks in computer vision and machine learning. Gaurav Parmar, Dacheng Li, Kwonjoon Lee, Zhuowen Tu |
CVPR | 2 |
| 2020 | Adaptive Detection Algorithm for Hazardous Clouds Based on Infrared Remote Sensing Spectroscopy and the LASSO MethodabstractLongwave infrared (LWIR) spectroscopy is useful for detecting and identifying hazardous clouds by passive remote sensing technology. Gaseous constituents are usually assumed to be thin plumes in a three-layer model, from which the spectral signatures are linearly superimposed on the brightness temperature spectrum. However, the thin-plume model performs poorly in cases of thick clouds. A modification to this method is made using synthetic references as target spectra, which allow linear models to be used for thick clouds. The prior background, which is generally unknown in most applications, is reconstructed through a regression method using predefined references. However, large residuals caused by fitting errors may distort the extracted spectral signatures and identification results if the predefined references are not consistent with the real spectral shapes. A group of references are generated to represent the possible spectral shapes, and the least absolute shrinkage and selection operator (LASSO) method is used to select the most appropriate reference for spectral fitting. Small residuals and adaptive identification are achieved by automatically selecting the reference spectrum. Two experiments are performed to verify the algorithm proposed in this article. Ethylene is adaptively detected during an indoor release process, and the spectral shape varies with the amount released. In addition, ammonia is measured under different humidity conditions, and the background is adaptively removed using the LASSO method. Based on this research, LWIR remote sensing technology can be applied in various target-detection scenarios, and adaptive identification is achieved to promote hazardous cloud detection. Dacheng Li, Fangxiao Cui, Anjing Wang, Yangyu Li, Yanli Qiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | A sensor-based scheme for assessing cloud coverage in HJ-1 CCD dataabstractCurrently, most of the effective cloud detection algorithms utilize either temperature properties or reference image(s). Typically, the automatic cloud cover assessment (ACCA) algorithm for Landsat 7 ETM+ images detects cloudy pixels relying on the thermal channel (three times used in two passes). In another study, Fernando flags clouds through analyzing the reflectance distribution between the high resolution cloudy image and a reference cloud-free MODIS image. However, for those sensors whose thermal channels are absent, less achievement has been obtained when no reference image can be acquired. In this paper, we develop a sensor-based cloud detection scheme for HJ-1 CCD data. HJ-1 (A/B) satellite (the environment and disaster monitoring and forecasting small satellites) is an optical satellite which covers a 720 square km area and consists of four visible and near-infrared spectral bands with 30 m resolution, and has been applied in land use and disaster detection. Dacheng Li |
IGARSS | 1 |
| 2012 | An adaptive and automated method for masking cloud on Landsat dataabstractIn this paper, an automated cloud masking approach for Landsat images has been proposed. The methodology couples linear constraint extracted from the Tasseled Cap Transformation (TCT) and the Normalized Difference Vegetation Index (NDVI) for masking cloud. The results are assessed in comparison with the Automatic Cloud Cover Assessment (ACCA) algorithm and show a fine agreement for Landsat ETM+ images. Since the algorithm proposed here doesn't rely on the thermal band and other accessorial data, this method could be applied to other high resolution sensors, like SPOT, HJ, IRS, et al. Dacheng Li, Yanqin Ge |
IGARSS | 1 |