EDBT 2026 Demo / reviewers in the wild / expert
Zhi Han
dblp:14/3563
· DBLP profile ↗
65ranked-venue papers
9as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2Theory of computation · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bring Your Dreams to Life: Continual Text-to-Video CustomizationabstractCustomized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do not expand incrementally over time. Additionally, they struggle with forgetting and concept neglect when continuously learning new concepts, including subjects and motions. To resolve the above challenges, we develop a novel Continual Customized Video Diffusion (CCVD) model, which can continuously learn new concepts to generate videos across various text-to-video generation tasks by tackling forgetting and concept neglect. To address catastrophic forgetting, we introduce a concept-specific attribute retention module and a task-aware concept aggregation strategy. They can capture the unique characteristics and identities of old concepts during training, while combining all subject and motion adapters of old concepts based on their relevance during testing. Besides, to tackle concept neglect, we develop a controllable conditional synthesis to enhance regional features and align video contexts with user conditions, by incorporating layer-specific region attention-guided noise estimation. Extensive experimental comparisons demonstrate that our CCVD outperforms existing CTVG models. Jiahua Dong 0001, Wenqi Liang, Zongyan Han, Meng Cao 0002, Duzhen Zhang, Hanbin Zhao, Zhi Han, Salman Khan 0001, Fahad Shahbaz Khan |
AAAI | 8 |
| 2026 | SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningabstractSequential-Horizon Vision-and-Language Navigation (SH-VLN) presents a challenging scenario where agents should sequentially execute multi-task trajectory navigation guided by complex, long-horizon natural language instructions. Current vision-and-language navigation models exhibit significant performance degradation with such instructions, as information overload impairs the agent's ability to attend to observationally relevant details. To address this problem, we propose SeqWalker, a novel navigation model built on a hierarchical planning framework. Our SeqWalker features: (1) A High-Level Planner that dynamically selects global instructions into contextually relevant sub-instructions based on the agent's current visual observations, thus reducing cognitive load; (2) A Low-Level Planner incorporating an Exploration-Verification strategy that leverages the inherent logical structure of instructions for trajectory error correction. To evaluate SH-VLN performance, we also extend the IVLN dataset and establish a new benchmark. Extensive experiments are performed to demonstrate the effectiveness and superiority of SeqWalker. Zebin Han, Baichen Liu, Qi Lyu, Zhenduo Shang, Jiahua Dong 0001, Lianqing Liu, Zhi Han |
AAAI | 8 |
| 2026 | Lifelong Language-Conditioned Robotic Manipulation Learning
Zebin Han, Gan Li, Jiahua Dong 0001, Baichen Liu, Lianqing Liu, Zhi Han |
AAAI | 8 |
| 2026 | ObjecTok: Learning Holistic and Robust Object Tokens for MLLMsabstractMainstream multimodal large language models (MLLMs) rely on patch-based tokenization methods, which compromise the integrity of objects and thereby limit the model's perception capabilities while triggering object-related hallucinations. To address this issue, we propose ObjecTok, an innovative object tokenization framework. ObjecTok generates a single, holistic object token for each object in an image. This token is produced by a specially trained object encoder that embeds the object's semantic, positional, and shape information into a single compact representation, thereby preserving the object's integrity. To mitigate the imperfections of upstream object proposer models, we introduce learnable confidence embeddings. These embeddings enable the MLLM to learn the reliability of each object's information, significantly enhancing the model's robustness. Additionally, ObjecTok employs a hybrid input strategy, combining object tokens with traditional image patch tokens, allowing the model to leverage both object-level information and global scene context. By integrating ObjecTok into the LLaVA architecture, we achieve notable performance improvements on multiple object-centric benchmarks, effectively reducing object hallucinations and enhancing perception capabilities. Experimental results robustly demonstrate that the object tokens generated by our ObjecTok framework hold great potential for building more powerful and reliable MLLMs. Xiyao Liu 0002, Lianqing Liu, Zhi Han |
AAAI | 4 |
| 2026 | Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation AlignmentabstractWe present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require complex architectural redesigns, ARRA aligns LLM's hidden states with visual representations from external visual foundational models via a global visual alignment loss and a hybrid token, . This token enforces dual constraints: local next-token prediction and global semantic distillation, enabling LLMs to implicitly learn spatial and contextual coherence while retaining their original autoregressive paradigm. Extensive experiments validate ARRA's plug-and-play versatility. When training T2I LLMs from scratch, ARRA reduces FID by 16.6% (ImageNet), 12.0% (LAION-COCO) for autoregressive LLMs like LlamaGen, without modifying original architecture and inference mechanism. For training from text-generation-only LLMs, ARRA reduces FID by 25.5% (MIMIC-CXR), 8.8% (DeepEyeNet) for advanced LLMs like Chameleon. For domain adaptation, ARRA aligns general-purpose LLMs with specialized models (e.g., BioMedCLIP), achieving an 18.6% FID reduction over direct fine-tuning on medical imaging (MIMIC-CXR). These results demonstrate that training objective redesign, rather than architectural modifications, can resolve cross-modal global coherence challenges. ARRA offers a complementary paradigm for advancing autoregressive models. Jiawei Liu 0003, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu |
AAAI | 5 |
| 2026 | Zero-shot single-image 3D generation via multi-grained semantic guidance
Xiyao Liu 0002, Xiai Chen, Lianqing Liu, Zhi Han |
Neurocomputing | 5 |
| 2026 | Rail image harmonization dataset: a seed to generate evaluation resources for track vision inspection systems
Zishen Zhao, Zhi Han, Chunlei Chen, Jinfei Hao, Shengzhe Tao |
Multim. Syst. | 3 |
| 2026 | TacMan-Turbo: Proactive Tactile Control for Robust and Efficient Articulated Object ManipulationabstractAdept manipulation of articulated objects is essential for robots to operate successfully in human environments. Such manipulation requires both effectiveness—reliable operation despite uncertain object structures—and efficiency—swift execution with minimal redundant steps and smooth trajectories. Existing approaches struggle to achieve both objectives simultaneously: methods relying on predefined kinematic models lack robustness when encountering structural variations, while the tactile-informed approach achieves robust manipulation but sacrifices efficiency through reactive, step-by-step execute-and-recover cycles. To address this challenge, this paper introduces TacMan-Turbo, a proactive tactile control framework that unifies the cycles into a continuous control loop. Our key insight is to interpret tactile signals temporally in addition to spatially: instead of treating contact deviations merely as instantaneous error signals requiring immediate compensation, we analyze sequential tactile observations to reveal local kinematic information. This temporal perspective enables our controller to predict future object states and proactively modulate manipulation velocities, eliminating the previous execute-and-recover cycle. In evaluations across 200 diverse simulated articulated objects and real-world experiments, our approach maintains a near-perfect performance while significantly enhancing time efficiency, action efficiency, and trajectory smoothness (all p-values < 0.0001). These results demonstrate that tactile feedback serves a dual purpose: providing not only spatial error signals for reactive control but also critical temporal structural information that enables proactive prediction and control. By enabling robots to manipulate articulated objects both reliably and efficiently, this work advances robot capabilities for seamless operation in dynamic, human-centric environments. Zihang Zhao, Zhenghao Qi, Leiyao Cui, Zhi Han, Lecheng Ruan, Yixin Zhu 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video Generation
Siyuan Wang 0010, Jiawei Liu 0003, Wei Wang 0474, Yeying Jin, Jinsong Du, Zhi Han |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance SegmentationabstractContinual video instance segmentation (CVIS) requires the plasticity to absorb new categories while maintaining the stability to retain previously learned knowledge. Crucially, the model must also preserve temporal consistency of instances across video frames. In this work, we introduce Contrastive Residual Injection and Semantic Prompting (CRISP), a framework tailored to address instance-wise, category-wise, and task-wise confusion in CVIS. For instance-wise learning, we model instance tracking and construct instance correlation loss, which emphasizes the correlation with the prior query space while strengthening the specificity of the current task query. For category-wise learning, we build an adaptive residual semantic prompt (ARSP) learning framework, which constructs a learnable semantic residual prompt pool generated by category text and uses an adjustive query-prompt matching mechanism to build a mapping relationship between the query of the current task and the semantic residual prompt. Meanwhile, a semantic consistency loss based on the contrastive learning is introduced to maintain semantic coherence between object queries and residual prompts during incremental training. For task-wise learning, to ensure the correlation at the inter-task level within the query space, we introduce a concise yet powerful initialization strategy for incremental prompts. Extensive experiments on YouTube-VIS-2019 and YouTube-VIS-2021 datasets demonstrate that CRISP significantly outperforms existing continual segmentation methods in the long-term continual video instance segmentation task, avoiding catastrophic forgetting and effectively improving segmentation and classification performance. The code is available at https://github.com/LyuQi127/CRISP. Baichen Liu, Qi Lyu, Jiahua Dong 0001, Lianqing Liu, Zhi Han |
IEEE Trans. Image Process. | 6 |
| 2026 | DVG-Diffusion: Dual-View-Guided Diffusion Model for CT Reconstruction From X-RaysabstractDirectly reconstructing 3D CT volume from few-view 2D X-rays using an end-to-end deep learning network is a challenging task, as X-ray images are merely projection views of the 3D CT volume. In this work, we facilitate complex 2D X-ray image to 3D CT mapping by incorporating new view synthesis, and reduce the learning difficulty through view-guided feature alignment. Specifically, we propose a dual-view guided diffusion model (DVG-Diffusion), which couples a real input X-ray view and a synthesized new X-ray view to jointly guide CT reconstruction. First, a novel view parameter-guided encoder captures features from X-rays that are spatially aligned with CT. Next, we concatenate the extracted dual-view features as conditions for the latent diffusion model to learn and refine the CT latent representation. Finally, the CT latent representation is decoded into a CT volume in pixel space. By incorporating view parameter guided encoding and dual-view guided CT reconstruction, our DVG-Diffusion can achieve an effective balance between high fidelity and perceptual quality for CT reconstruction. Experimental results demonstrate our method outperforms state-of-the-art methods. Based on experiments, the comprehensive analysis and discussions for views and reconstruction are also presented. The model and code are available at https://github.com/xiexing0916/DVG-Diffusion. Jiawei Liu 0003, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu |
IEEE Trans. Image Process. | 4 |
| 2025 | Label smoothing regularization-based no hyperparameter domain generalization
Xiyao Liu 0002, Fupeng Chu, Zhi Han |
Knowl. Based Syst. | 6 |
| 2025 | Rope-net: deep convolutional neural network via robust principal component analysis
Baichen Liu, Zhi Han, Xi'ai Chen, Yandong Tang |
Mach. Learn. | 2 |
| 2025 | Gradient Projection for Continual Parameter-Efficient TuningabstractParameter-efficient tunings (PETs) have demonstrated impressive performance and promising perspectives in training large models, while they are still confronted with a common problem: the trade-off between learning new content and protecting old knowledge, leading to zero-shot generalization collapse, and cross-modal hallucination. In this paper, we reformulate Adapter, LoRA, Prefix-tuning, and Prompt-tuning from the perspective of gradient projection, and first propose a unified framework called Parameter Efficient Gradient Projection (PEGP). We introduce orthogonal gradient projection into different PET paradigms and theoretically demonstrate that the orthogonal condition for the gradient can effectively resist forgetting even for large-scale models. It therefore modifies the gradient towards the direction that has less impact on the old feature space, with less extra memory space and training time. We extensively evaluate our method with different backbones, including ViT and CLIP, on diverse datasets, and experiments comprehensively demonstrate its efficiency in reducing forgetting in class, online class, domain, task, and multi-modality continual settings. Jingyang Qiao, Zhizhong Zhang 0001, Xin Tan 0002, Yanyun Qu, Wensheng Zhang 0002, Zhi Han, Yuan Xie 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | DUAL-GDFQ: A Dual-Generator, Dual-Phase Learning Approach for Data-Free QuantizationabstractData-free quantization (DFQ) seeks to maximize the performance of quantized networks without requiring original training data. Conventional methods, which use synthetic samples from generators for network fine-tuning, often yield inferior results compared to training conducted with real data. To mitigate this problem, we introduce a dual-generator, dual-phase learning generative data-free quantization (DUAL-GDFQ) method, which utilizes two generators: a knowledge-matching generator and a knowledge-promoting generator for replicating the original data distribution as well as keeping samples informative. Additionally, inspired by meta-learning, the proposed novel dual-phase learning scheme can effectively utilize the capabilities of both generators by aligning their gradient descent directions. Theoretical analysis and extensive experiments demonstrate that our method successfully minimizes performance degradation in quantized networks and can achieve performance levels comparable to training with real data. Zhi Han, Xiyao Liu 0002 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Facial Expression Monitoring via Fine-Grained Vision-Language AlignmentabstractIn the fields of health care and clinical monitoring, vision-based Facial Expression Recognition (FER) has achieved significantly progress, but it still faces the challenge of poor generalization ability under unconstrained conditions of occlusions and pose variation. Recently, Vision-Language Model (VLM) has greatly advanced the FER task. However, the existing VLM-based FER methods typically leverage a hard-crafted prompt (e.g., “a photo of [class]”) and only focus on the holistic semantic alignment, which may suffer from modal heterogeneity. In this work, we propose a fine-grained vision-language model via Prompt Masking for FER (PMFER). Specifically, for each expression, we first create fine-grained prompts using facial action units to guide the image encoder to learn discriminative representations. Further, to finely align text prompts and visual action units, we randomly drop a phrase description in the prompts and then predict the dropped phrase by conducting modal cross attention, implicitly promoting fine-grained vision-language alignment. In addition, we also design a modal-adversarial strategy to holistically eliminate the modal difference between visual and textual embeddings in a common latent space. Experimental results demonstrate that our PMFER model outperforms the state-of-the-art methods on several FER benchmarks, especially under the conditions of occlusions and pose variations. Note to Practitioners—Facial expression recognition is very important in health care and clinical monitoring, which provides an useful tool to assess the psychological and physiological conditions of patients. Although FER has made significant progress with the development of deep learning technologies, it still faces problems in the complex environments (e.g., occlusions and pose variations). To address the above issues, we propose a novel FER method in this work based on the recent vision-language model. It takes RGB image and text prompts as input and finally predicts the expression classification. Different from the existing methods, the proposed PMFER can enable fine-grained modal alignment for facial key units. Compared with the state-of-the-art methods on the public datasets, it can achieve better results, especially under the conditions of occlusions and pose variations. Also, we evaluate the proposed method on a real-world pain dataset, and the results demonstrate that PMFER has a good generalization and can be applied to health care. Weihong Ren, Yu Gao 0010, Xi'ai Chen, Zhi Han, Zhiyong Wang 0009, Jiaole Wang, Honghai Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | A Subspace-Based Method for Facial Image EditingabstractIn the realm of computational social systems, the ability to edit facial attributes accurately plays a crucial role in enhancing user experience on social media platforms and virtual environments. However, we face significant challenges in isolated attribute manipulation and balancing the tradeoff between editing fidelity and facial identity preservation. Here, this article presents a novel approach to constructing an orthogonal decomposition subspace, enabling precise editing control over individual attributes with minimal impact on others and maintaining identity consistency. We introduce an adaptive weight modulation (AWM) method and a maximum slope truncation (MST) formula. The AWM method, founded on a sufficient convergent criterion, performs singular value decomposition to yield subspace parameters that preserve rich facial knowledge within the generative model, facilitating high-quality facial generation with reduced parameterization. This empowers meaningful semantic interpretation of attributes, supporting diverse editing tasks such as pose, age, and eyewear adjustments. The MST formula rigorously defines the editing bounds to effectively navigate the tradeoff between editing depth and identity retention. We also propose a guideline for deciphering the specific meanings of unsupervised semantics, potentially advancing interpretability in social behavioral studies. An accompanying web application, available athttps://github.com/mickoluan/GreenLimeSia, has been developed, granting users the freedom to perform tailored facial edits. Extensive experimental results show we pave the way for more personalized and authentic interactions within computational social platforms. MengChu Zhou, Xin Luan, Liang Qi 0001, Yandong Tang, Zhi Han |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Synergistic Prompting Learning for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection, as a foundational task in human-centric understanding, aims to detect interactive triplets in real-world scenarios. To better distinguish diverse HOIs within an open-world context, current HOI detectors utilize pre-trained Visual-Language Models (VLMs) to extract prior knowledge through textual prompts (i.e., descriptive texts for each HOI instance). However, relying on predetermined descriptive texts, such approaches only acquire a fixed set of textual knowledge for HOI prediction, consequently resulting in inferior performance and limited generalization. To remedy this, we propose a novel VLM-based method, which jointly performs prompting learning from both visual and textual perspectives and synergizes visual-textual prompting for HOI detection. Initially, we design a hierarchical adaptation architecture to perform progressive prompting: visual prompting is facilitated through gradual token migration from VLM's image encoder, while textual prompting is initialized with progressively leveled interaction descriptions. In addition, to synergize the visual-textual prompting learning, a text-supervising and image-tuning loop is introduced, in which the text-supervising stage guides visual prompting learning through contrastive learning and the image-tuning stage refines textual prompting by modal matching. Finally, we employ an interaction-aware knowledge merging mechanism to effectively transfer visual-textual knowledge encapsulated within synergistic prompting for HOI detection. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings. Jinguo Luo, Weihong Ren, Zhiyong Wang 0009, Xi'ai Chen, Huijie Fan, Zhi Han, Honghai Liu 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Diverse Representations Embedding for Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) aims to continuously learn from sequential data streams, enabling cross-camera matching of individuals over time. A critical challenge in LReID lies in balancing the preservation of previously acquired knowledge with the incremental acquisition of new information, due to task-level gaps and limited representation capacity. Conventional methods relying on CNN backbones struggle to fully capture the diverse perspectives of each instance, leading to suboptimal model performance. To tackle these limitations, we propose a diverse representation embedding (DRE) framework that balances preserving old knowledge with adapting to new information. Specifically, our DRE incorporates a robust Transformer-based backbone that utilizes maximum embedding (ME) and multiple class tokens to generate overlapping representations for each instance. To further enhance the model's representation capacity, we design an adaptive constraint module (ACM), which performs integration and discrimination operations on overlapping representations to yield diverse yet diverse representations. Furthermore, we propose two strategies: knowledge update (KU) and knowledge preservation (KP), implemented within the adjustment and learner models, respectively. The KU strategy enhances the learner model's ability to adapt to new information by leveraging prior knowledge from the adjustment model. The KP strategy ensures the retention of historical knowledge while maintaining the model's adaptability. Extensive experiments validate that our DRE surpasses state-of-the-art approaches across large-scale, occluded, and holistic datasets, demonstrating significant performance gains. Our code is available at https://github.com/LiuShiBen/DRE. Shiben Liu, Huijie Fan, Qiang Wang 0015, Xi'ai Chen, Zhi Han, Yandong Tang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Discovering Syntactic Interaction Clues for Human-Object Interaction DetectionabstractRecently, Vision-Language Model (VLM) has greatly ad-vanced the Human-Object Interaction (HOI) detection. The existing VLM-based HOI detectors typically adopt a hand-crafted template (e.g., a photo of a person [action] a/an [object]) to acquire text knowledge through the VLM text encoder. However, such approaches, only encoding the action-specific text prompts in vocabulary level, may suffer from learning ambiguity without exploring the fine-grained clues from the perspective of interaction context. In this paper, we propose a novel method to discover Syntactic Interaction Clues for HOI detection (SICHOI) by using VLM. Specifically, we first investigate what are the essen-tial elements for an interaction context, and then establish a syntactic interaction bank from three levels: spatial relationship, action-oriented posture and situational condition. Further, to align visual features with the syntactic interaction bank, we adopt a multi-view extractor to jointly aggre-gate visual features from instance, interaction, and image levels accordingly. In addition, we also introduce a dual cross-attention decoder to perform context propagation be-tween text knowledge and visual features, thereby enhancing the HOI detection. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on HICO-DET and V-COCO. Jinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen, Qiang Wang 0015, Zhi Han, Honghai Liu 0001 |
CVPR | 6 |
| 2024 | Uni-YOLO: Vision-Language Model-Guided YOLO for Robust and Fast Universal Detection in the Open WorldabstractUniversal object detectors aim to detect any object in any scene without human annotation, exhibiting superior generalization. However, the current universal object detectors show degraded performance in harsh weather, and their insufficient real-time capabilities limit their application. In this paper, we present Uni-YOLO, a universal detector designed for complex scenes with real-time performance. Uni-YOLO is a one-stage object detector that uses general object confidence to distinguish between objects and backgrounds, and employs a grid cell regression method for real-time detection. To improve its robustness in harsh weather conditions, the input of Uni-YOLO is adaptively enhanced with a physical model-based enhancement module. During training and inference, Uni-YOLO is guided by the extensive knowledge of the vision-language model CLIP. An object augmentation method is proposed to improve generalization in training by utilizing multiple source datasets with heterogeneous annotations. Furthermore, an online self-enhancement method is proposed to allow Uni-YOLO to further focus on specific objects through self-supervised fine-tuning in a given scene. Extensive experiments on public benchmarks and a UAV deployment are conducted to validate its superiority and practical value. Weihong Ren, Xi'ai Chen, Huijie Fan, Yandong Tang, Zhi Han |
ACM Multimedia | 6 |
| 2024 | Feature distillation and guide network for unsupervised underwater image enhancement
Xin Luan, Qiang Wang 0015, Huijie Fan, Xiai Chen, Zhi Han, Yandong Tang |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | CCR: Facial Image Editing with Continuity, Consistency and Reversibility
Xin Luan, Huidi Jia, Zhi Han, Xiaofeng Li 0001, Yandong Tang |
Int. J. Comput. Vis. | 4 |
| 2024 | Wavelet-pixel domain progressive fusion network for underwater image enhancement
Shiben Liu, Huijie Fan, Qiang Wang 0015, Zhi Han, Yandong Tang |
Knowl. Based Syst. | 4 |
| 2024 | Learning Self- and Cross-Triplet Context Clues for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection aims to infer interactions between humans and objects, and it is very important for scene analysis and understanding. The existing methods usually focus on exploring instance-level (e.g., object appearance) or interaction-level (e.g., action semantic) features to conduct interaction prediction. However, most of these methods only consider the self-triplet feature aggregation, which may lead to learning ambiguity without exploring the cross-triplet context exchange. In this paper, from both visual and textual perspectives, we propose a novel method to jointly explore self-and cross-triplet interaction context clues for HOI detection. First, we employ a graph neural network to perform self-triplet aggregation, where human and object features represent graph nodes and visual interaction feature and textual prior knowledge are acted as two different edges. Furthermore, we also attempt to explore cross-triplet context exchange by incorporating symbiotic and layout relationships among different HOI triplets. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones and achieves the impressive performance of 40.32 mAP on HICO-DET and 69.1 mAP on V-COCO datasets, respectively. Weihong Ren, Jinguo Luo, Weibo Jiang, Liangqiong Qu, Zhi Han, Jiandong Tian, Honghai Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Online Video Sparse Noise Removing via Nonlocal Robust PCAabstractOnline schemes and nonlocal similarity are two effective approaches for strengthening robust principal component analysis (RPCA) techniques in video denoising. However, their limitations are also evident. The online scheme is usually highly efficient but lacks consideration of regional appearance information, thus it cannot effectively handle videos with complex dynamics such as object movements. On the other hand, nonlocal similarity is used to better utilize regional information but incurs a heavy computational cost. Moreover, these two techniques are incompatible and challenging to work together. To overcome this barrier and harness the advantages of both approaches, this paper proposes a novel online nonlocal RPCA method. 1) A clustering based nonlocal strategy (ClusNonlocal) is adopted, which not only greatly reduces the computation cost, but also forms low-dimensional subspaces for online processing; 2) a new weighted RPCA model is proposed, which regards samples with different importances and improves the performance of subspace pursuit and video recovery; 3) a multi-level subspace updating scheme and weighted projection method is proposed, which keeps the performance of online video data processing at a high level at all time. A series of video denoising experiments are carried out to demonstrate the overall advantages of our procedure over several other ones, in terms of both visual quality and running speed. Zhi Han, Huijie Fan, Yandong Tang, Yao Wang 0003 |
IEEE Trans. Multim. | 1 |
| 2023 | Neural network equivalent model for highly efficient massive data classification
Siquan Yu, Zhi Han, Yandong Tang, Chengdong Wu 0001 |
Sci. China Inf. Sci. | 2 |
| 2023 | Self-taught cross-domain few-shot learning with weakly supervised object localization and task-decomposition
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Zhi Han |
Knowl. Based Syst. | 4 |
| 2023 | Dual Distillation Discriminator Networks for Domain Adaptive Few-Shot Learning
Xiyao Liu 0002, Zhong Ji, Yanwei Pang, Zhi Han |
Neural Networks | 4 |
| 2023 | Tensor Robust Principal Component Analysis With Side Information: Models and ApplicationsabstractAs a domain-dependent prior knowledge, side information has been introduced into Robust Principal Component Analysis (RPCA) to alleviate its degenerate or suboptimal performance in some real applications. It has recently realized that the natural structural information can be better retained if the observed data is kept in the original tensor form rather than matricizing it or other order reduction means. Hence, studies on RPCA of tensor version have attracted more and more attentions. To share the merits from both direct tensor modeling and side information, we propose three models to deal with the problem of Tensor RPCA with side information based on tensor Singular Value Decomposition (t-SVD). To solve these models, we develop an efficient algorithm with convergence guarantee using the well-known alternating direction method of multiplier. Extensive experimental studies on both synthetic and real-world tensor data have been carried out to demonstrate the superiority of the proposed models over several other state-of-the-arts. Our code is released athttps://github.com/zsj9509/TPCPSF. Zhi Han, Junping Yao, Yao Wang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | TSAFinder: exhaustive tumor-specific antigen detection with RNAseqabstractMOTIVATION: Tumor-specific antigen (TSA) identification in human cancer predicts response to immunotherapy and provides targets for cancer vaccine and adoptive T-cell therapies with curative potential, and TSAs that are highly expressed at the RNA level are more likely to be presented on major histocompatibility complex (MHC)-I. Direct measurements of the RNA expression of peptides would allow for generalized prediction of TSAs. Human leukocyte antigen (HLA)-I genotypes were predicted with seq2HLA. RNA sequencing (RNAseq) fastq files were translated into all possible peptides of length 8-11, and peptides with high and low expressions in the tumor and control samples, respectively, were tested for their MHC-I binding potential with netMHCpan-4.0. RESULTS: A novel pipeline for TSA prediction from RNAseq was used to predict all possible unique peptides size 8-11 on previously published murine and human lung and lymphoma tumors and validated on matched tumor and control lung adenocarcinoma (LUAD) samples. We show that neoantigens predicted by exomeSeq are typically poorly expressed at the RNA level, and a fraction is expressed in matched normal samples. TSAs presented in the proteomics data have higher RNA abundance and lower MHC-I binding percentile, and these attributes are used to discover high confidence TSAs within the validation cohort. Finally, a subset of these high confidence TSAs is expressed in a majority of LUAD tumors and represents attractive vaccine targets. AVAILABILITY AND IMPLEMENTATION: The datasets were derived from sources in the public domain as follows: TSAFinder is open-source software written in python and R. It is licensed under CC-BY-NC-SA and can be downloaded at https://github.com/RNAseqTSA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael F. Sharpnack, Travis S. Johnson, Robert Chalkley, Zhi Han, David Carbone, Kun Huang 0001 |
Bioinform. | 4 |
| 2022 | Application of unsupervised deep learning algorithms for identification of specific clusters of chronic cough patients from EMR dataabstractBACKGROUND: Chronic cough affects approximately 10% of adults. The lack of ICD codes for chronic cough makes it challenging to apply supervised learning methods to predict the characteristics of chronic cough patients, thereby requiring the identification of chronic cough patients by other mechanisms. We developed a deep clustering algorithm with auto-encoder embedding (DCAE) to identify clusters of chronic cough patients based on data from a large cohort of 264,146 patients from the Electronic Medical Records (EMR) system. We constructed features using the diagnosis within the EMR, then built a clustering-oriented loss function directly on embedded features of the deep autoencoder to jointly perform feature refinement and cluster assignment. Lastly, we performed statistical analysis on the identified clusters to characterize the chronic cough patients compared to the non-chronic cough patients. RESULTS: The experimental results show that the DCAE model generated three chronic cough clusters and one non-chronic cough patient cluster. We found various diagnoses, medications, and lab tests highly associated with chronic cough patients by comparing the chronic cough cluster with the non-chronic cough cluster. Comparison of chronic cough clusters demonstrated that certain combinations of medications and diagnoses characterize some chronic cough clusters. CONCLUSIONS: To the best of our knowledge, this study is the first to test the potential of unsupervised deep learning methods for chronic cough investigation, which also shows a great advantage over existing algorithms for patient data clustering. Wei Shao 0005, Xiao Luo 0002, Zuoyi Zhang, Zhi Han, Vasu Chandrasekaran, Vladimir Turzhitsky, Vishal Bali, Anna R. Roberts, Megan Metzger, Jarod Baker, Carmen La Rosa, Jessica Weaver, Paul Richard Dexter, Kun Huang 0001 |
BMC Bioinform. | 4 |
| 2022 | A novel compact design of convolutional layers with spatial transformation towards lower-rank representation for image classificationabstractConvolutional neural networks (CNNs) usually come with numerous parameters and thus are not convenient for some situations, such as when the storage space is limited. Low-rank decomposition is one effective way for network compression or compaction. However, the current methods are far from theoretical optimal compression performance because the low-rankness of the commonly trained convolution filter sets is limited because of the versatility of convolution filters. We propose a novel compact design for convolutional layers with spatial transformation for achieving a much lower-rank form. The convolution filters in our design are generated using a predefined Tucker product form, followed by learnable individual spatial transformations on each filter. The low-rank (Tucker) part lowers the parameter capacity while the transformation part enhances the feature representation capacity. We validate our proposed approach on an image classification task. Our approach focuses on compressing parameters while also improving accuracy. We perform experiments on the MNIST, CIFAR10, CIFAR100, and ImageNet datasets. On the ImageNet dataset, our approach outperforms low-rank based state-of-the-arts by 2% to 6% in top-1 validation accuracy. Furthermore, our approach outperforms a series of low-rank-based state-of-the-arts on various datasets. The experiments validate the efficacy of our proposed method. Our code is available at https://github.com/liubc17/low_rank_compact_transformed. Baichen Liu, Zhi Han, Xiai Chen, Wenming Shao, Huidi Jia, Yandong Tang |
Knowl. Based Syst. | 2 |
| 2022 | Depth Selection for Deep ReLU Nets in Feature Extraction and GeneralizationabstractDeep learning is recognized to be capable of discovering deep features for representation learning and pattern recognition without requiring elegant feature engineering techniques by taking advantages of human ingenuity and prior knowledge. Thus it has triggered enormous research activities in machine learning and pattern recognition. One of the most important challenges of deep learning is to figure out relations between a feature and the depth of deep neural networks (deep nets for short) to reflect the necessity of depth. Our purpose is to quantify this feature-depth correspondence in feature extraction and generalization. We present the adaptivity of features to depths and vice-verse via showing a depth-parameter trade-off in extracting both single feature and composite features. Based on these results, we prove that implementing the classical empirical risk minimization on deep nets can achieve the optimal generalization performance for numerous learning tasks. Our theoretical results are verified by a series of numerical experiments including toy simulations and a real application of earthquake seismic intensity prediction. Zhi Han, Siquan Yu, Shaobo Lin, Ding-Xuan Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Effective Tensor Completion via Element-Wise Weighted Low-Rank Tensor Train With Overlapping Ket AugmentationabstractTensor completion methods based on the tensor train (TT) have the issues of inaccurate weight assignment and ineffective tensor augmentation pre-processing. In this work, we propose a novel tensor completion approach via the element-wise weighted technique. Accordingly, a novel formulation for tensor completion and an effective optimization algorithm, called tensor completion by parallel weighted matrix factorization via tensor train (TWMac-TT), is proposed. In addition, we specifically consider the recovery quality of edge elements from adjacent blocks. Different from traditional reshaping and ket augmentation, we utilize a new tensor augmentation technique called overlapping ket augmentation, which can further avoid blocking artifacts. We then conduct extensive performance evaluations on synthetic data and several real image data sets. Our experimental results demonstrate that the proposed algorithm TWMac-TT outperforms several other competing tensor completion methods. The code is available athttps://github.com/yzcv/TWMac-TT-OKA Yang Zhang 0073, Yao Wang 0003, Zhi Han, Xiai Chen, Yandong Tang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | A Computational Framework to Analyze the Associations Between Symptoms and Cancer Patient Attributes Post Chemotherapy Using EHR DataabstractPatients with cancer, such as breast and colorectal cancer, often experience different symptoms post-chemotherapy. The symptoms could be fatigue, gastrointestinal (nausea, vomiting, lack of appetite), psychoneurological symptoms (depressive symptoms, anxiety), or other types. Previous research focused on understanding the symptoms using survey data. In this research, we propose to utilize the data within the Electronic Health Record (EHR). A computational framework is developed to use a natural language processing (NLP) pipeline to extract the clinician-documented symptoms from clinical notes. Then, a patient clustering method is based on the symptom severity levels to group the patient in clusters. The association rule mining is used to analyze the associations between symptoms and patient attributes (smoking history, number of comorbidities, diabetes status, age at diagnosis) in the patient clusters. The results show that the various symptom types and severity levels have different associations between breast and colorectal cancers and different timeframes post-chemotherapy. The results also show that patients with breast or colorectal cancers, who smoke and have severe fatigue, likely have severe gastrointestinal symptoms six months after the chemotherapy. Our framework can be generalized to analyze symptoms or symptom clusters of other chronic diseases where symptom management is critical. Xiao Luo 0002, Priyanka Gandhi, Susan Storey, Zuoyi Zhang, Zhi Han, Kun Huang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Weakly Supervised Deep Ordinal Cox Model for Survival Prediction From Whole-Slide Pathological ImagesabstractWhole-Slide Histopathology Image (WSI) is generally considered the gold standard for cancer diagnosis and prognosis. Given the large inter-operator variation among pathologists, there is an imperative need to develop machine learning models based on WSIs for consistently predicting patient prognosis. The existing WSI-based prediction methods do not utilize the ordinal ranking loss to train the prognosis model, and thus cannot model the strong ordinal information among different patients in an efficient way. Another challenge is that a WSI is of large size (e.g., 100,000-by-100,000 pixels) with heterogeneous patterns but often only annotated with a single WSI-level label, which further complicates the training process. To address these challenges, we consider the ordinal characteristic of the survival process by adding a ranking-based regularization term on the Cox model and propose a weakly supervised deep ordinal Cox model (BDOCOX) for survival prediction from WSIs. Here, we generate amounts of bags from WSIs, and each bag is comprised of the image patches representing the heterogeneous patterns of WSIs, which is assumed to match the WSI-level labels for training the proposed model. The effectiveness of the proposed method is well validated by theoretical analysis as well as the prognosis and patient stratification results on three cancer datasets from The Cancer Genome Atlas (TCGA). Wei Shao 0005, Tongxin Wang, Zhi Han, Jie Zhang 0010, Kun Huang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Low-rank decomposition on transformed feature maps domain for image denoising
Qiong Luo 0003, Baichen Liu, Yang Zhang 0073, Zhi Han, Yandong Tang |
Vis. Comput. | 4 |
| 2020 | Multi-task multi-modal learning for joint diagnosis and prognosis of human cancers
Wei Shao 0005, Tongxin Wang, Liang Sun 0009, Tianhan Dong, Zhi Han, Jie Zhang 0010, Daoqiang Zhang, Kun Huang 0001 |
Medical Image Anal. | 5 |
| 2020 | Integrative Analysis of Pathological Images and Multi-Dimensional Genomic Data for Early-Stage Cancer PrognosisabstractThe integrative analysis of histopathological images and genomic data has received increasing attention for studying the complex mechanisms of driving cancers. However, most image-genomic studies have been restricted to combining histopathological images with the single modality of genomic data (e.g., mRNA transcription or genetic mutation), and thus neglect the fact that the molecular architecture of cancer is manifested at multiple levels, including genetic, epigenetic, transcriptional, and post-transcriptional events. To address this issue, we propose a novel ordinal multi-modal feature selection (OMMFS) framework that can simultaneously identify important features from both pathological images and multi-modal genomic data (i.e., mRNA transcription, copy number variation, and DNA methylation data) for the prognosis of cancer patients. Our model is based on a generalized sparse canonical correlation analysis framework, by which we also take advantage of the ordinal survival information among different patients for survival outcome prediction. We evaluate our method on three early-stage cancer datasets derived from The Cancer Genome Atlas (TCGA) project, and the experimental results demonstrated that both the selected image and multi-modal genomic markers are strongly correlated with survival enabling effective stratification of patients with distinct survival than the comparing methods, which is often difficult for early-stage cancer patients. Wei Shao 0005, Kun Huang 0001, Zhi Han, Jun Cheng 0006, Tongxin Wang, Liang Sun 0009, Zixiao Lu, Jie Zhang 0010, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Multilinear Multitask Learning by Rank-Product RegularizationabstractMultilinear multitask learning (MLMTL) considers an MTL problem in which tasks are arranged by multiple indices. By exploiting the higher order correlations among the tasks, MLMTL is expected to improve the performance of traditional MTL, which only considers the first-order correlation across all tasks, e.g., low-rank structure of the coefficient matrix. The key to MLMTL is designing a rational regularization term to represent the latent correlation structure underlying the coefficient tensor instead of matrix. In this paper, we propose a new MLMTL model by employing the rank-product regularization term in the objective, which on one hand can automatically rectify the weights along all its tensor modes and on the other hand have an explicit physical meaning. By using this regularization, the intrinsic high-order correlations among tasks can be more precisely described, and thus, the overall performance of all tasks can be improved. To solve the resulted optimization model, we design an efficient algorithm by applying the alternating direction method of multipliers (ADMM). We also analyze the convergence and show that the proposed algorithm, with certain restriction, is asymptotically regular. Experiments on both synthetic and real data sets substantiate the superiority of the proposed method beyond the existing MLMTL methods in terms of accuracy and efficiency. Qian Zhao 0002, Xiangyu Rui, Zhi Han, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Diagnosis-Guided Multi-modal Feature Selection for Prognosis Prediction of Lung Squamous Cell Carcinoma
Wei Shao 0005, Tongxin Wang, Jun Cheng 0006, Zhi Han, Daoqiang Zhang, Kun Huang 0001 |
MICCAI (4) | 5 |
| 2018 | Genetic Mutations Associated with Histopathology Changes in Kidney Cancer
Jun Cheng 0006, Zhi Han, Qianjin Feng 0002, Jie Zhang 0010, Kun Huang 0001 |
AMIA | 2 |
| 2018 | Ordinal Multi-modal Feature Selection for Survival Analysis of Early-Stage Renal Cancer
Wei Shao 0005, Jun Cheng 0006, Liang Sun 0009, Zhi Han, Qianjin Feng 0002, Daoqiang Zhang, Kun Huang 0001 |
MICCAI (2) | 4 |
| 2018 | Robust video denoising with sparse and dense noise modelings
Guiping Shen, Zhi Han, Xiai Chen, Yandong Tang |
Sci. China Inf. Sci. | 2 |
| 2018 | On Convergence Properties of Implicit Self-paced Objective
Zilu Ma, Shiqi Liu 0001, Deyu Meng, Sio-Long Lo, Zhi Han |
Inf. Sci. | 6 |
| 2018 | Snowflake Removal for Videos via Global and Local Low-Rank DecompositionabstractFalling snow not only blocks human vision, but also significantly degrades the effectiveness of computer vision systems in outdoor environment. In this paper, we aim to remove snowflakes in videos by using the global and local low-rank property of snowflake-removed scenes. The stationary background and the mixture of moving foreground as well as falling snowflake are extracted via the global low-rank matrix decomposition. Some snowflake features, such as its color and size, are used to separate out the snowflakes from other moving objects. Then, the mean absolute difference based patch matching is applied to align every same moving object over frames to grab its low-rank structure. As such, the falling snowflake in front of moving objects can be removed via the local low-rank decomposition. Finally, the snowflake removed videos are generated by pasting moving foreground to stationary backgrounds. Experiments show that our method can remove snowflakes effectively and outperforms the comparison methods. Jiandong Tian, Zhi Han, Weihong Ren, Xiai Chen, Yandong Tang |
IEEE Trans. Multim. | 2 |
| 2018 | A Generalized Model for Robust Tensor Factorization With Noise Modeling by Mixture of GaussiansabstractThe low-rank tensor factorization (LRTF) technique has received increasing attention in many computer vision applications. Compared with the traditional matrix factorization technique, it can better preserve the intrinsic structure information and thus has a better low-dimensional subspace recovery performance. Basically, the desired low-rank tensor is recovered by minimizing the least square loss between the input data and its factorized representation. Since the least square loss is most optimal when the noise follows a Gaussian distribution, -norm-based methods are designed to deal with outliers. Unfortunately, they may lose their effectiveness when dealing with real data, which are often contaminated by complex noise. In this paper, we consider integrating the noise modeling technique into a generalized weighted LRTF (GWLRTF) procedure. This procedure treats the original issue as an LRTF problem and models the noise using a mixture of Gaussians (MoG), a procedure called MoG GWLRTF. To extend the applicability of the model, two typical tensor factorization operations, i.e., CANDECOMP/PARAFAC factorization and Tucker factorization, are incorporated into the LRTF procedure. Its parameters are updated under the expectation-maximization framework. Extensive experiments indicate the respective advantages of these two versions of MoG GWLRTF in various applications and also demonstrate their effectiveness compared with other competing methods. Xiai Chen, Zhi Han, Yao Wang 0003, Qian Zhao 0002, Deyu Meng, Lin Lin 0007, Yandong Tang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Video Desnowing and Deraining Based on Matrix DecompositionabstractThe existing snow/rain removal methods often fail for heavy snow/rain and dynamic scene. One reason for the failure is due to the assumption that all the snowflakes/rain streaks are sparse in snow/rain scenes. The other is that the existing methods often can not differentiate moving objects and snowflakes/rain streaks. In this paper, we propose a model based on matrix decomposition for video desnowing and deraining to solve the problems mentioned above. We divide snowflakes/rain streaks into two categories: sparse ones and dense ones. With background fluctuations and optical flow information, the detection of moving objects and sparse snowflakes/rain streaks is formulated as a multi-label Markov Random Fields (MRFs). As for dense snowflakes/rain streaks, they are considered to obey Gaussian distribution. The snowflakes/rain streaks, including sparse ones and dense ones, in scene backgrounds are removed by low-rank representation of the backgrounds. Meanwhile, a group sparsity term in our model is designed to filter snow/rain pixels within the moving objects. Experimental results show that our proposed model performs better than the state-of-the-art methods for snow and rain removal. Weihong Ren, Jiandong Tian, Zhi Han, Antoni B. Chan, Yandong Tang |
CVPR | 3 |
| 2017 | Tensor RPCA by Bayesian CP Factorization with Complex NoiseabstractThe RPCA model has achieved good performances in various applications. However, two defects limit its effectiveness. Firstly, it is designed for dealing with data in matrix form, which fails to exploit the structure information of higher order tensor data in some pratical situations. Secondly, it adopts L1-norm to tackle noise part which makes it only valid for sparse noise. In this paper, we propose a tensor RPCA model based on CP decomposition and model data noise by Mixture of Gaussians (MoG). The use of tensor structure to raw data allows us to make full use of the inherent structure priors, and MoG is a general approximator to any blends of consecutive distributions, which makes our approach capable of regaining the low dimensional linear subspace from a wide range of noises or their mixture. The model is solved by a new proposed algorithm inferred under a variational Bayesian framework. The superiority of our approach over the existing state-of-the-art approaches is demonstrated by extensive experiments on both of synthetic and real data. Qiong Luo 0003, Zhi Han, Xiai Chen, Yao Wang 0003, Deyu Meng, Yandong Tang |
ICCV | 2 |
| 2017 | A New Intrinsic-Lighting Color Space for Daytime Outdoor ImagesabstractExtracting or separating intrinsic information and illumination from natural images is crucial for better solving computer vision tasks. In this paper, we present a new illumination-based color space, the IL (intrinsic information and lighting level) space. Its first two channels represent 2D intrinsic information, and the third channel is for lighting levels. The IL color space has a one-to-one correspondence with the RGB color space. One valuable benefit of the IL color space is that illumination-related processing can be realized by directly operating on the lighting channel. As an example, based on the extracted lighting channel, we propose a new algorithm to estimate the intrinsic lighting level of an image such that the shadow-free color image and relighting series are obtained. In contrast to the existing color spaces for display or printing, the IL color space intuitively shows the information of reflectance and lighting levels for colors separately. Zhi Han, Jiandong Tian, Liangqiong Qu, Yandong Tang |
IEEE Trans. Image Process. | 1 |
| 2016 | Robust Tensor Factorization with Unknown NoiseabstractBecause of the limitations of matrix factorization, such as losing spatial structure information, the concept of tensor factorization has been applied for the recovery of a low dimensional subspace from high dimensional visual data. Generally, the recovery is achieved by minimizing the loss function between the observed data and the factorization representation. Under different assumptions of the noise distribution, the loss functions are in various forms, like L1 and L2 norms. However, real data are often corrupted by noise with an unknown distribution. Then any specific form of loss function for one specific kind of noise often fails to tackle such real data with unknown noise. In this paper, we propose a tensor factorization algorithm to model the noise as a Mixture of Gaussians (MoG). As MoG has the ability of universally approximating any hybrids of continuous distributions, our algorithm can effectively recover the low dimensional subspace from various forms of noisy observations. The parameters of MoG are estimated under the EM framework and through a new developed algorithm of weighted low-rank tensor factorization (WLRTF). The effectiveness of our algorithm are substantiated by extensive experiments on both of synthetic data and real image data. Xiai Chen, Zhi Han, Yao Wang 0003, Qian Zhao 0002, Deyu Meng, Yandong Tang |
CVPR | 2 |
| 2016 | Super-resolution reconstruction of hyperspectral images via low rank tensor modeling and total variation regularizationabstractIn this paper, we propose a novel approach to hyperspectral image super-resolution by modeling the global spatial-and-spectral correlation and local smoothness properties over hyperspectral images. Specifically, we utilize the tensor nuclear norm and tensor folded-concave penalty functions to describe the global spatial-and-spectral correlation hidden in hyperspectral images, and 3D total variation (TV) to characterize the local spatial-and-spectral smoothness across all hyperspectral bands. Then, we develop an efficient algorithm for solving the resulting optimization problem by combing the local linear approximation (LLA) strategy and alternative direction method of multipliers (ADMM). Experimental results on one hyperspectral image dataset illustrate the merits of the proposed approach. Shiying He, Haiwei Zhou, Yao Wang 0003, Wenfei Cao, Zhi Han |
IGARSS | 5 |
| 2016 | Nonconvex plus quadratic penalized low-rank and sparse decomposition for noisy image alignment
Xiai Chen, Zhi Han, Yao Wang 0003, Yandong Tang |
Sci. China Inf. Sci. | 2 |
| 2015 | Folded-concave penalization approaches to tensor completion
Wenfei Cao, Yao Wang 0003, Can Yang 0002, Xiangyu Chang, Zhi Han, Zongben Xu |
Neurocomputing | 5 |
| 2014 | A hyperdense semantic domain for hybrid dynamic systems to model different classes of discontinuitiesabstractThe physics of technical systems, such as embedded and cyber-physical systems, is frequently modeled using the notion of continuous time. The underlying continuous phenomena may, however, occur at a time scale much faster than the system behavior of interest. In such situations, it is desirable to approximate the detailed continuous-time behavior by discontinuous change. Two classes of discontinuous change can be identified: pinnacles and mythical modes. This work shows how pinnacles are well modeled using a hyperreal notion of time while a superdense notion of time applies well to mythical modes. Thus, the combination, called hyperdense time, is proposed to allow for the expression of the semantics of both pinnacles and mythical modes. Further, the hyperdense semantic domain is translated into a computational representation as a three-dimensional model of time. In particular, continuous-time behavior is mapped onto floating point numbers, while the mythical mode and pinnacle event iterations each map onto an integer dimension. A modified Newton's cradle is used as a case study and to illustrate the computational implementation. Pieter J. Mosterman, Gabor Simko, Justyna Zander, Zhi Han |
HSCC | 4 |
| 2013 | Towards sensitivity analysis of hybrid systems using simulinkabstractIn the design of engineered systems two types of models are used: (i) analysis models and (ii) system models. The system models are primary deliverables between design stages whereas analysis models are employed within a design stage. Sensitivity analysis studies the behavior of the system under small parameter variations which proves to be useful in design. To enable sensitivity analysis in verification of hybrid dynamic systems that model industry-size problems, support for simulation-based methods is desired. The computational semantics for simulation of corresponding analysis models must then be consistent with the computational semantics of the system models. A method is presented that enables direct sensitivity analysis on system models via an implementation in the Simulink(R) software. The approach relies on the existing ordinary differential equation solver of Simulink and the block-by-block analytic Jacobian computation to provide the analytic Jacobian for solving the sensitivity equations. Results of a prototype implementation show that sensitivity analysis can be applied to moderate size Simulink models of continuous and hybrid systems. Zhi Han, Pieter J. Mosterman |
HSCC | 1 |
| 2012 | A signal processing approach for enriched region detection in RNA polymerase II ChIP-seq dataabstractBACKGROUND: RNA polymerase II (PolII) is essential in gene transcription and ChIP-seq experiments have been used to study PolII binding patterns over the entire genome. However, since PolII enriched regions in the genome can be very long, existing peak finding algorithms for ChIP-seq data are not adequate for identifying such long regions. METHODS: Here we propose an enriched region detection method for ChIP-seq data to identify long enriched regions by combining a signal denoising algorithm with a false discovery rate (FDR) approach. The binned ChIP-seq data for PolII are first processed using a non-local means (NL-means) algorithm for purposes of denoising. Then, a FDR approach is developed to determine the threshold for marking enriched regions in the binned histogram. RESULTS: We first test our method using a public PolII ChIP-seq dataset and compare our results with published results obtained using the published algorithm HPeak. Our results show a high consistency with the published results (80-100%). Then, we apply our proposed method on PolII ChIP-seq data generated in our own study on the effects of hormone on the breast cancer cell line MCF7. The results demonstrate that our method can effectively identify long enriched regions in ChIP-seq datasets. Specifically, pertaining to MCF7 control samples we identified 5,911 segments with length of at least 4 Kbp (maximum 233,000 bp); and in MCF7 treated with E2 samples, we identified 6,200 such segments (maximum 325,000 bp). CONCLUSIONS: We demonstrated the effectiveness of this method in studying binding patterns of PolII in cancer cells which enables further deep analysis in transcription regulation and epigenetics. Our method complements existing peak detection algorithms for ChIP-seq experiments. Zhi Han, Thierry Pécot, Tim Hui-Ming Huang, Raghu Machiraju, Kun Huang 0001 |
BMC Bioinform. | 1 |
| 2011 | Video Primal Sketch: A generic middle-level representation of videoabstractThis paper presents a middle-level video representation named Video Primal Sketch (VPS), which integrates two regimes of models: i) sparse coding model using static or moving primitives to explicitly represent moving corners, lines, feature points, etc., ii) FRAME/MRF model with spatio-temporal filters to implicitly represent textured motion, such as water and fire, by matching feature statistics, i.e. histograms. This paper makes three contributions: i) learning a dictionary of video primitives as parametric generative model; ii) studying the Spatio-Temporal FRAME (ST-FRAME) model for modeling and synthesizing textured motion; and iii) developing a parsimonious hybrid model for generic video representation. VPS selects the proper representation automatically and is compatible with high-level action representations. In the experiments, we synthesize a series of dynamic textures, reconstruct real videos and show varying VPS over the change of densities causing by the scale transition in videos. Zhi Han, Zongben Xu, Song-Chun Zhu |
ICCV | 1 |
| 2011 | Incremental Alignment Manifold Learning
Zhi Han, Deyu Meng, Zongben Xu, Nannan Gu |
J. Comput. Sci. Technol. | 1 |
| 2005 | A two-stage handwritten character segmentation approach in mail address recognitionabstractCharacter segmentation has become a crucial step for mail address recognition in the automatic post mail sorting system. In this paper, a two-stage character segmentation algorithm according to the characteristics of handwritten mail address characters is proposed. In the simple segmentation stage, the block sequence is extracted from the mail address image using the structure-based methods, including projection profile analysis, connected components analysis and stroke cross number analysis. In the precise segmentation stage, all candidate segmentation paths are created by combining the neighboring blocks and represented with a candidate segmentation graph first. Then several optimal candidate paths are selected from the graph by dynamic programming searching based on recognition confidence. Finally the best segmentation path is determined by matching these paths with the known post address database. In the experiment on more than 500 real envelop images with the this approach, the correct sorting rate of address recognition is up to 79.46% and that of address-postcode integrated recognition is up to 96.26%. Zhi Han, Chang-Ping Liu |
ICDAR | 1 |
| 2005 | Financial Document Image Coding with Regions of Interest Using JPEG2000abstractDocument image coding is a very important issue in document analysis and recognition systems provided with vast samples. An image compression algorithm with regions of interest (ROIs) using JPEG2000 is proposed for financial document images which have various categories, complex layouts, and irregular noises. Three types of ROIs: filled information ROIs, seal ROIs, and handwriting ROIs, are detected and extracted through document knowledge analysis and handwriting identification. The first ROIs are detected by document classification, the second are extracted by connected component analysis based on color and shape information, and the third are located by handwriting identification using an incremental Fisher linear discriminant classifier. A ROI mask with a random shape is constructed by thresholding and merging these ROIs. Finally, a financial document image is encoded using JPEG2000 Part I with this ROI mask. Compared to JPEG and DjVu, the method improves visual quality while decreasing storing space. Xu-Cheng Yin, Chang-Ping Liu, Zhi Han |
ICDAR | 3 |
| 2005 | Feature combination using boosting
Xu-Cheng Yin, Chang-Ping Liu, Zhi Han |
Pattern Recognit. Lett. | 3 |
| 2004 | Managing Verification Activities Using SVM
Bill Aldrich, Ansgar Fehnker, Peter H. Feiler, Zhi Han, Bruce H. Krogh, Eric Lim, Shiva Sivashankar |
ICFEM | 4 |
| 2003 | Verification of Hybrid Systems Based on Counterexample-Guided Abstraction Refinement
Edmund M. Clarke, Ansgar Fehnker, Zhi Han, Bruce H. Krogh, Olaf Stursberg, Michael Theobald |
TACAS | 3 |