Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dechen Zhan

dblp:74/4048 · also De-chen Zhan · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-5973-9542ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Computer networks · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Planning, search and constraint satisfaction · 31% Information extraction and text analysis · 29% Vision and language · 20%
Network and information security
1 paper
Security and privacy of machine learning · 67% Privacy and data protection · 33%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 50% Reconfigurable computing and FPGAs · 50%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
vision-language model
1.722025
SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment · ACM Multimedia 2025
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning · ACM Multimedia 2025
Natural language and speech › Information extraction and text analysis
semantic parsing
1.222023
MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic Parsing · AAAI 2023
Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic Knowledge · EMNLP 2022
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
1.222023
MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic Parsing · AAAI 2023
Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic Knowledge · EMNLP 2022
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning
0.912025
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning · ACM Multimedia 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › procedural reasoning
multimodal procedural planning
0.912025
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning · ACM Multimedia 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › task planning
procedure planning
0.912025
LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning · ACM Multimedia 2025
Visual content generation and editing › layout generation
graphic layout generation
0.812024
Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe Locator · AAAI 2024
Security and privacy of machine learning
adversarial attack
0.812024
Adversarial Perturbation Prediction for Real-Time Protection of Speech Privacy · IEEE Trans. Inf. Forensics Secur. 2024
Security and privacy of machine learning › adversarial attack
adversarial perturbation
0.812024
Adversarial Perturbation Prediction for Real-Time Protection of Speech Privacy · IEEE Trans. Inf. Forensics Secur. 2024
Privacy and data protection › data confidentiality › content privacy › multimedia privacy
speech privacy
0.812024
Adversarial Perturbation Prediction for Real-Time Protection of Speech Privacy · IEEE Trans. Inf. Forensics Secur. 2024
Machine learning › Efficient and distributed learning
model compression
0.412019
Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity · FPGA 2019
Machine learning › Efficient and distributed learning › model compression › sparsity
structured sparsity
0.412019
Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity · FPGA 2019
Reconfigurable computing and FPGAs
FPGA accelerator
0.412019
Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity · FPGA 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.412019
Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity · FPGA 2019
Machine learning › Trustworthy machine learning
interpretability
0.312025
SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment · ACM Multimedia 2025
Machine learning › Generative modeling
autoregressive model
0.212024
Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe Locator · AAAI 2024
Natural language and speech › Speech recognition and synthesis
speaker recognition
0.212024
Adversarial Perturbation Prediction for Real-Time Protection of Speech Privacy · IEEE Trans. Inf. Forensics Secur. 2024
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
0.212023
MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic Parsing · AAAI 2023
Data models and query languages
SQL
0.212022
Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic Knowledge · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

wireframe locator · 1.5non-autoregressive decoding · 1.5iterative refinement · 1.5adversarial perturbation prediction · 1.5vision-language model · 0.9task-oriented segmentation · 0.9self-guided fact enhancement · 0.9reranking · 0.9entropy-aware direct preference optimization · 0.9counterfactual reasoning · 0.9streaming audio processing · 0.8formulaic knowledge · 0.6weight pruning · 0.4compressed sparse banks · 0.4bank-balanced sparsity · 0.4
YearPublicationVenuePosition
2025 LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
abstract
While large language models (LLMs) have advanced procedural planning for embodied AI systems through strong reasoning abilities, the integration of multimodal inputs and counterfactual reasoning remains underexplored. To tackle these challenges, we introduce LLaPa, a vision-language model framework designed for multimodal procedural planning. LLaPa generates executable action sequences from textual task descriptions and visual environmental images using vision-language models (VLMs). Furthermore, we enhance LLaPa with two auxiliary modules to improve procedural planning. The first module, the Task-Environment Reranker (TER), leverages task-oriented segmentation to create a task-sensitive feature space, aligning textual descriptions with visual environments and emphasizing critical regions for procedural execution. The second module, the Counterfactual Activities Retriever (CAR), identifies and emphasizes potential counterfactual conditions, enhancing the model's reasoning capability in counterfactual scenarios. Extensive experiments on ActPlan-1K and ALFRED benchmarks demonstrate that LLaPa generates higher-quality plans with superior LCS and correctness, outperforming advanced models. The code and models are available https://github.com/sunshibo1234/LLaPa.
Shibo Sun, Xue Li 0011, Donglin Di, Lanshun Nie, Weinan Zhang 0003, Dechen Zhan, Yang Song 0001, Lei Fan 0007
ACM Multimedia7
2025 SAGE: A Visual Language Model for Anomaly Detection via Fact Enhancement and Entropy-aware Alignment
abstract
While Vision-Language Models (VLMs) have shown promising progress in general multimodal tasks, they often struggle with industrial anomaly detection and reasoning, particularly in delivering interpretable explanations and generalizing to unseen categories. This limitation stems from the inherently domain-specific nature of anomaly detection, which hinders the applicability of existing VLMs in industrial scenarios that require precise, structured, and context-aware analysis. To address these challenges, we propose SAGE, a VLM-based framework that enhances anomaly reasoning through Self-Guided Fact Enhancement (SFE) and Entropy-aware Direct Preference Optimization (E-DPO). SFE integrates domain-specific knowledge into visual reasoning via fact extraction and fusion, while E-DPO aligns model outputs with expert preferences using entropy-aware optimization. Additionally, we introduce AD-PL, a preference-optimized dataset tailored for industrial anomaly reasoning, consisting of 28,415 question-answering instances with expert-ranked responses. To evaluate anomaly reasoning models, we develop Multiscale Logical Evaluation (MLE), a quantitative framework analyzing model logic and consistency. SAGE demonstrates superior performance on industrial anomaly datasets under zero-shot and one-shot settings. The code, model, and dataset are available at https://github.com/amoreZgx1n/SAGE.
Guoxin Zang, Xue Li 0011, Donglin Di, Lanshun Nie, Dechen Zhan, Yang Song 0001, Lei Fan 0007
ACM Multimedia5
2025 S2C-HAR: A Semi-Supervised Human Activity Recognition Framework Based on Contrastive Learning
abstract
ABSTRACT Human activity recognition (HAR) has emerged as a critical element in various domains, such as smart healthcare, smart homes, and intelligent transportation, owing to the rapid advancements in wearable sensing technology and mobile computing. Nevertheless, existing HAR methods predominantly rely on deep supervised learning algorithms, necessitating a substantial supply of high‐quality labeled data, which significantly impacts their accuracy and reliability. Considering the diversity of mobile devices and usage environments, the quest for optimizing recognition performance in deep models while minimizing labeled data usage has become a prominent research area. In this paper, we propose a novel semi‐supervised HAR framework based on contrastive learning named S2C‐HAR, which is capable of generating accurate pseudo‐labels for unlabeled data, thus achieving comparable performance with supervised learning with only a few labels applied. First, a contrastive learning model for HAR (CLHAR) is designed for more general feature representations, which contains a contrastive augmentation transformer pre‐trained exclusively on unlabeled data and fine‐tuned in conjunction with a model‐agnostic classification network. Furthermore, based on the FixMatch technique, unlabeled data with two different perturbations imposed are fed into the CLHAR to produce pseudo‐labels and prediction results, which effectively provides a robust self‐training strategy and improves the quality of pseudo‐labels. To validate the efficacy of our proposed model, we conducted extensive experiments, yielding compelling results. Remarkably, even with only 1% labeled data, our model achieves satisfactory recognition performance, outperforming state‐of‐the‐art methods by approximately 5%.
Xue Li 0011, Mingxing Liu, Lanshun Nie, Wenxiao Cheng, Xiaohe Wu, Dechen Zhan
Concurr. Comput. Pract. Exp.6
2024 Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe Locator
abstract
Layout generation is a critical step in graphic design to achieve meaningful compositions of elements. Most previous works view it as a sequence generation problem by concatenating element attribute tokens (i.e., category, size, position). So far the autoregressive approach (AR) has achieved promising results, but is still limited in global context modeling and suffers from error propagation since it can only attend to the previously generated tokens. Recent non-autoregressive attempts (NAR) have shown competitive results, which provides a wider context range and the flexibility to refine with iterative decoding. However, current works only use simple heuristics to recognize erroneous tokens for refinement which is inaccurate. This paper first conducts an in-depth analysis to better understand the difference between the AR and NAR framework. Furthermore, based on our observation that pixel space is more sensitive in capturing spatial patterns of graphic layouts (e.g., overlap, alignment), we propose a learning-based locator to detect erroneous tokens which takes the wireframe image rendered from the generated layout sequence as input. We show that it serves as a complementary modality to the element sequence in object space and contributes greatly to the overall performance. Experiments on two public datasets show that our approach outperforms both AR and NAR baselines. Extensive studies further prove the effectiveness of different modules with interesting findings. Our code will be available at https://github.com/ffffatgoose/SpotError.
Jieru Lin, Danqing Huang, Tiejun Zhao, Dechen Zhan, Chin-Yew Lin
AAAI4
2024 Adversarial Perturbation Prediction for Real-Time Protection of Speech Privacy
abstract
The widespread collection and analysis of private speech signals have become increasingly prevalent, raising significant privacy concerns. To protect speech signals from unauthorized analysis, adversarial attack methods for deceiving speaker recognition models have been proposed. While a few of these methods are specifically designed for real-time protection of speech signals, they introduce significant delays that can severely impact speech communication when applied to streaming speech data. In this paper, we present a novel approach that aims to offer real-time protection for speech signals without delays. By utilizing observed data only, we generate initial adversarial seed perturbations and refine them to obtain the necessary adversarial perturbations predicted for adjacent unobserved signals. This refinement process is conducted via a proposed model called PAPG. On the basis of perturbation prediction, we develop a streaming audio processing framework that generates perturbations in synchronization with the playback of the original signal, effectively eliminating delays. The experimental results demonstrate that under the proposed attack, the average Top-1 accuracy of various advanced speaker recognition methods is reduced by 89%, and the average equal error rate (EER) increases to 36%. Remarkably, these results are achieved without delays while maintaining superior perceptual quality.
Zhaoyang Zhang 0002, Shen Wang 0004, Guopu Zhu, Dechen Zhan, Jiwu Huang
IEEE Trans. Inf. Forensics Secur.4
2023 MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic Parsing
abstract
Text-to-SQL semantic parsing is an important NLP task, which facilitates the interaction between users and the database. Much recent progress in text-to-SQL has been driven by large-scale datasets, but most of them are centered on English. In this work, we present MultiSpider, the largest multilingual text-to-SQL semantic parsing dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese). Upon MultiSpider we further identify the lexical and structural challenges of text-to-SQL (caused by specific language properties and dialect sayings) and their intensity across different languages. Experimental results under various settings (zero-shot, monolingual and multilingual) reveal a 6.1% absolute drop in accuracy in non-English languages. Qualitative and quantitative analyses are conducted to understand the reason for the performance drop of each language. Besides the dataset, we also propose a simple schema augmentation framework SAVe (Schema-Augmentation-with-Verification), which significantly boosts the overall performance by about 1.8% and closes the 29.5% performance gap across languages.
Longxu Dou, Yan Gao 0002, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Jian-Guang Lou
AAAI6
2023 Cross-domain network attack detection enabled by heterogeneous transfer learning
Chunrui Zhang 0002, Gang Wang 0027, Shen Wang 0004, Dechen Zhan, Mingyong Yin
Comput. Networks4
2022 Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic Knowledge
abstract
Longxu Dou, Yan Gao, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Longxu Dou, Yan Gao 0002, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou
EMNLP7
2022 Few shot learning-based fast adaptation for human activity recognition
Lanshun Nie, Xue Li 0011, Tianying Gong, Dechen Zhan
Pattern Recognit. Lett.4
2021 Enhancing Representation of Deep Features for Sensor-Based Activity Recognition
Xue Li 0011, Lanshun Nie, Xiandong Si, Renjie Ding, Dechen Zhan
Mob. Networks Appl.5
2020 A Class Incremental Temporal-Spatial Model Based on Wireless Sensor Networks for Activity Recognition
Xue Li 0011, Lanshun Nie, Xiandong Si, Dechen Zhan
WASA (1)4
2019 Efficient and Effective Sparse LSTM on FPGA with Bank-Balanced Sparsity
abstract
Neural networks based on Long Short-Term Memory (LSTM) are widely deployed in latency-sensitive language and speech applications. To speed up LSTM inference, previous research proposes weight pruning techniques to reduce computational cost. Unfortunately, irregular computation and memory accesses in unrestricted sparse LSTM limit the realizable parallelism, especially when implemented on FPGA. To address this issue, some researchers propose block-based sparsity patterns to increase the regularity of sparse weight matrices, but these approaches suffer from deteriorated prediction accuracy. This work presents Bank-Balanced Sparsity (BBS), a novel sparsity pattern that can maintain model accuracy at a high sparsity level while still enable an efficient FPGA implementation. BBS partitions each weight matrix row into banks for parallel computing, while adopts fine-grained pruning inside each bank to maintain model accuracy. We develop a 3-step software-hardware co-optimization approach to apply BBS in real FPGA hardware. First, we propose a bank-balanced pruning method to induce the BBS pattern on weight matrices. Then we introduce a decoding-free sparse matrix format, Compressed Sparse Banks (CSB), that transparently exposes inter-bank parallelism in BBS to hardware. Finally, we design an FPGA accelerator that takes advantage of BBS to eliminate irregular computation and memory accesses. Implemented on Intel Arria-10 FPGA, the BBS accelerator can achieve 750.9 GOPs on sparse LSTM networks with a batch size of 1. Compared to state-of-the-art FPGA accelerators for LSTM with different compression techniques, the BBS accelerator achieves 2.3 ~ 3.7x improvement on energy efficiency and 7.0 ~ 34.4x reduction on latency with negligible loss of model accuracy.
Shijie Cao, Chen Zhang 0001, Zhuliang Yao, Wencong Xiao, Lanshun Nie, Dechen Zhan, Yunxin Liu 0001, Ming Wu 0007
FPGA6
2019 An Automatic Method to Estimate the Calibration Quality of the Aeromagnetic Compensation
abstract
Aeromagnetic compensation is an effective technique to reduce magnetic interference of the aircraft during an aero-magnetic survey. As an important indicator to evaluate the compensation effect, the figure of merit (FOM) is widely applied. However, calculating FOM needs to label the maneuvers manually. In order to solve this problem, we propose an automatic method to calculate the FOM. This method labels the maneuvers through classifying the flight direction and recognizing the attitude by combining two clustering algorithms, k-means and Gaussian mixture model (GMM). The required magnetic fields are collected by a vector-magnetometer, which is commonly used during the magnetic surveys. Results of the tests demonstrate the availability of the proposed method.
Qi Han 0002, Dechen Zhan
IGARSS4
2019 Automatic determination of types number of mixed binary protocols
abstract
In the absence of prior knowledge, it is a challenge to determine the number of protocol types in frames, which are completely unknown. These frames might be mixed with multiple protocols from the data link layer to the application layer. In this study, the authors combine the spectral clustering algorithm with a method of determining the number of protocol frame types, and further more design the refinement clustering of different hierarchical protocols based on the eigenvectors of the Laplace matrix. They use three clustering validity indices, which are Calinski–Harabasz index, Davies–Bouldinn index and Silhouette index, to quantify the clustering effect in order to calculate the number of protocol types. Extensive experiments on several open datasets obtain relatively satisfying results without prior knowledge and demonstrate the significant advantages of their methods clearly.
Chunrui Zhang 0002, Shen Wang 0004, Dechen Zhan
IET Commun.3
2019 FlexSaaS: A Reconfigurable Accelerator for Web Search Selection
abstract
Web search engines deploy large-scale selection services on CPUs to identify a set of web pages that match user queries. An FPGA-based accelerator can exploit various levels of parallelism and provide a lower latency, higher throughput, more energy-efficient solution than commodity CPUs. However, maintaining such a customized accelerator in a commercial search engine is challenging because selection services are changed often. This article presents our design for FlexSaaS (Flexible Selection as a Service), an FPGA-based accelerator for web search selection. To address efficiency and flexibility challenges, FlexSaaS abstracts computing models and separates memory access from computation. Specifically, FlexSaaS (i) contains a reconfigurable number of matching processors that can handle various possible query plans, (ii) decouples index stream reading from matching computation to fetch and decode index files, and (iii) includes a universal memory accessor that hides the complex memory hierarchy and reduces host data access latency. Evaluated on FPGAs in the selection service of a commercial web search--the Bing web search engine—FlexSaaS can be evolved quickly to adapt to new updates. Compared to the software baseline, FlexSaaS on Arria 10 reduces average latency by 30% and increases throughput by 1.5×.
Shijie Cao, Lanshun Nie, Dechen Zhan, Ningyi Xu, Ramashis Das, Ming Wu 0007, Derek Chiou
ACM Trans. Reconfigurable Technol. Syst.3
2006 Service-Oriented Infrastructure for Collaborative Product Design in ETO Enterprises
abstract
Engineer-to-order (ETO) products have some unique characteristics on their production process and supply chain models. These characteristics require that their design process should not only export product structure and production procedure information, but also consider other objectives, e.g., feasibility of production planning, cost, quality, service, etc. Different roles, including customers, suppliers, and related departments in ETO enterprises, should be integrated closely for frequent collaborations by exchanging mass of data during design process. To address this issue, in this paper we introduce a Web service based infrastructure for ETO product collaborative design, in which related product design systems, e.g., CAD, CAPP and PDM, and related business information systems, e.g., ERP, are all partially encapsulated as basic Web services, and according to several primitive service interoperability patterns, these services are collaborated together to realize data and process integration with the aid of enterprise service bus (ESB), consequently forms an integrated design ecosystem, to improve design efficiency for ETO products
Zhongjie Wang 0003, Dechen Zhan, Xiaofei Xu 0001
CSCWD2
2006 Model for Negotiating Prices and Due Dates with Suppliers in Make-to-Order Supply Chains
Lanshun Nie, Xiaofei Xu 0001, Dechen Zhan
PRIMA3
2005 A Component Optimization Design Method based on Variation Point Decomposition
abstract
Traditional component design methods pay much attention to component's usefulness, while usually ignore the optimization on component's usability, such as reuse cost and reuse efficiency. During the whole lifecycle of component reuse, it is necessary for component to be continuously re-designed so as to reduce reuse cost under the guarantee of high reusability. In this paper, a feature-based component model is firstly introduced, with emphasis on reuse mechanism based on variation point. By analysis of constituents of component reuse cost, the optimization goal, i.e., increasing the proportion of fixed part of a component, is put forward. Then a component optimization method based on variation point decomposition is presented, i.e., decomposing every feature item of a variable feature into a set of sub-items and picking up those common sub-items as much as possible to extend fixed part of the component. Finally the effectiveness of this method is proved theoretically and practically.
Zhongjie Wang 0003, Xiaofei Xu 0001, Dechen Zhan
SERA3
2004 Viewing the Web as a Cube: The Vision and Approach
Xiaofei Xu 0001, Dechen Zhan
APWeb3
2004 A Reuse-Oriented Business Component Acquisition Method Based on Greatest Common Sub-Process
Zhongjie Wang 0003, Xiaofei Xu 0001, Dechen Zhan
SNPD3
2000 Dynamic Organization and Methodology for Agile Virtual Enterprises
Xiaofei Xu 0001, Quanlong Li, Dechen Zhan
J. Comput. Sci. Technol.4