VLDB 2026 Research / reviewers in the wild / expert
Ao Luo
dblp:63/9431
· DBLP profile ↗
55ranked-venue papers
21as first author
44since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 14 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 11 first-author · 30 since 2021Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-DisentangleabstractCompositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives, have not achieved effective state-object decoupling and causal interventional invariance, limiting their performance on unseen compositions. To tackle this challenge, this study introduces I2CD (Invertible Causal framework via Disentangle-Compose-Disentangle), a novel framework that integrates invertible neural networks with causal intervention techniques to achieve state-object disentanglement. The framework employs a disentangle-compose-disentangle mechanism for counterfactual generation within the disentangled representation space, ensuring that modifications to one primitive (attribute or object) maintain independence from the other, thus enabling robust causal disentanglement. Representational consistency is maintained through semantic alignment between initial disentangled representations and their recomposed-then-disentangled counterparts with corresponding textual concepts. Comprehensive evaluations on three benchmark datasets—MIT-States, UT-Zappos, and C-GQA—demonstrate the framework's effectiveness in achieving both disentanglement and compositional generalization in CZSL tasks. Zhaoquan Yuan, Yuankang Pan, Ao Luo, Wei Li 0110, Xiao Wu 0001, Changsheng Xu |
AAAI | 4 |
| 2026 | A 4.7 μW Dual-Phase Front End with Dual-Mode Buffer for Dry-Electrode ECG Acquisition
Ao Luo, Zhechang Hu, Pujia Xing, Liang Qi 0002, Yan Liu 0016 |
ISCAS | 1 |
| 2026 | Learning Efficient Meshflow and Optical Flow From Event CamerasabstractIn this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we review the state-of-the-art in event-based flow estimation, highlighting two key areas for further research: i) the lack of meshflow-specific event datasets and methods, and ii) the underexplored challenge of event data density. First, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280 × 720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (30×faster) of our EEMFlow model compared to the recent state-of-the-art flow method. As an extension, we expand HREM into HREM+, a multi-density event dataset contributing to a thorough study of the robustness of existing methods across data with varying densities, and propose an Adaptive Density Module (ADM) to adjust the density of input event data to a more optimal range, enhancing the model's generalization ability. We empirically demonstrate that ADM helps to significantly improve the performance of EEMFlow and EEMFlow+ by 8% and 10%, respectively. Xinglong Luo, Ao Luo, Kunming Luo, Zhengning Wang, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Learning Unknowns Without Forgetting Knowns: Compositional and Bidirectional Low-Rank Adaptive Open-World Detection TransformerabstractOpen-World Object Detection (OWOD) aims to detect unseen objects as “unknown” while incrementally learning them without catastrophic forgetting. This problem presents two major challenges: (1) the lack of annotations for unknown objects during training, and (2) the risk of catastrophic forgetting during model updates. To address these issues, we propose the COmpositional and Bidirectional low-Rank Adaptive open-world detection transformer (COBRA)-a novel framework built upon a pre-trained Deformable DETR model. Specifically, COBRA first employs an attentional filtering mechanism that prunes previously known (P-Known) and currently known (C-Known) objects, yielding a purified set of candidateunknowns. To system-atically pseudo-label theseunknowns, we introduce a Primitive Composition Recognition (PCR) module, which evaluates set-level similarity between candidate objects and learned primitives, enabling accurate labeling ofpseudo-unknowns. To mitigate catastrophic forgetting during incremental updates, COBRA leverages Bidirectional Low-Rank Adaptation (Bi-LoRA)-a parameter-efficient mechanism that supports forward knowledge transfer and stable backward integration. Together, these components form a synergistic pipeline for continual object discovery and knowledge consolidation. Extensive experiments on MS COCO and PASCAL VOC demonstrate that our rehearsal-free COBRA framework outperforms SAM-powered methods in unknown recall while achieving lower forgetting compared to rehearsal-based competitors. Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Wei Li 0110, Ao Luo, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Voxel-Based Point Cloud Geometry Compression With Space-to-Channel ContextabstractVoxel-based methods are among the most efficient for point cloud geometry compression, particularly with dense point clouds. However, they face limitations due to a restricted receptive field from the upsampling operation, especially when handling high-bit-depth point clouds. To overcome this issue, we introduce a stage-wise Space-to-Channel (S2C) context model for both dense point clouds and the low-level part of sparse point clouds. This model utilizes a channel-wise autoregressive strategy to effectively integrate neighborhood information at a coarse resolution. For the high-level part of sparse point clouds, we further propose a level-wise S2C context model that addresses resolution limitations by incorporating Geometry Residual Coding (GRC) for consistent-resolution cross-level prediction. Additionally, we use the spherical coordinate system for its compact representation and enhance our GRC approach with a Residual Probability Approximation (RPA) module, which features a large kernel size. Experimental results show that our S2C context model not only achieves bit savings while maintaining or improving reconstruction quality but also reduces computational complexity compared to state-of-the-art voxel-based compression methods. Yangzhi Ma, Ao Luo, Li Li 0040, Dong Liu 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | Forensics Adapter: Adapting CLIP for Generalizable Face Forgery DetectionabstractWe describe the Forensics Adapter, an adapter network designed to transform CLIP into an effective and generalizable face forgery detector. Although CLIP is highly versatile, adapting it for face forgery detection is nontrivial as forgery-related knowledge is entangled with a wide range of unrelated knowledge. Existing methods treat CLIP merely as a feature extractor, lacking task-specific adaptation, which limits their effectiveness. To address this, we introduce an adapter to learn face forgery traces – the blending boundaries unique to forged faces, guided by task-specific objectives. Then we enhance the CLIP visual tokens with a dedicated interaction strategy that communicates knowledge across CLIP and the adapter. Since the adapter is alongside CLIP, its versatility is highly retained, naturally ensuring strong generalizability in face forgery detection. With only 5.7M trainable parameters, our method achieves a significant performance boost, improving by approximately 7% on average across five standard datasets. We believe the proposed method can serve as a baseline for future CLIP-based face forgery detection methods. The code is available at https://github.com/OUCVAS/ForensicsAdapter. Xinjie Cui, Yuezun Li, Ao Luo, Jiaran Zhou, Junyu Dong |
CVPR | 3 |
| 2025 | MExD: An Expert-Infused Diffusion Model for Whole-Slide Image ClassificationabstractWhole Slide Image (WSI) classification poses unique challenges due to the vast image size and numerous non-informative regions, which introduce noise and cause data imbalance during feature aggregation. To address these issues, we propose MExD, an Expert-Infused Diffusion Model that combines the strengths of a Mixture-of-Experts (MoE) mechanism with a diffusion model for enhanced classification. MExD balances patch feature distribution through a novel MoE-based aggregator that selectively emphasizes relevant information, effectively filtering noise, addressing data imbalance, and extracting essential features. These features are then integrated via a diffusion-based generative process to directly yield the class distribution for the WSI. Moving beyond conventional discriminative approaches, MExD represents the first generative strategy in WSI classification, capturing fine-grained details for robust and precise results. Our MExD is validated on three widely-used benchmarks—Camelyon16, TCGANSCLC, and BRACS—consistently achieving state-of-the-art performance in both binary and multi-class tasks. Our code and model are available at https://github.com/JWZhao-uestc/MExD. Xin Li 0079, Fan Yang 0054, Qiang Zhai, Ao Luo, Yang Zhao 0024, Hong Cheng 0002, Huazhu Fu |
CVPR | 5 |
| 2025 | Latent Interactiveness Field for Non-Contact Human Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection serves a broad spectrum of applications. Despite significant progress, current approaches encounter difficulties in effectively handling Non-Contact Human-Object Interaction (NCHOI) scenarios, where humans and objects remain physically apart. To address these challenges, this paper proposes a novel approach, named Latent Interactiveness Field Modeling (LIFM), which enhances HOI detection by capturing long-range contextual dependencies. Specifically, the Latent Interactiveness Field (LIF) is introduced to define potential interactive relationships between humans and objects. To complement this, the LIF Fusion Encoder is designed to adaptively fuse visual features with LIF, resulting in more informative and discriminative feature representations. The Mobile Scanning HOI Dataset (MSHD) is introduced as a comprehensive benchmark to systematically assess the robustness of existing methods on both common HOI and NCHOI in real-world applications. Extensive experimentation indicates that the proposed approach outperforms existing state-of-the-art techniques. It offers substantial improvements, particularly in NCHOI scenarios, which highlight its effectiveness in resolving issues related to long-range interactions. Xiang Huang 0004, Ao Luo, Xiao Wu 0001, Zhaoquan Yuan |
ACM Multimedia | 2 |
| 2025 | Storage-and-Memory-Efficient Learned Image Compression With Quality-Aware Hyperprior PruningabstractABSTRACT Learned image compression (LIC) has become more and more important in recent years. The hyperprior‐module‐based LIC models, which use hyperprior module to predict the distribution of image features and improve entropy coder performance, have achieved remarkable rate‐distortion (RD) performance. However, the storage and memory costs of these LIC models are too high, resulting in higher difficulty to be applied to various devices, especially portable or edge devices. The storage and memory cost are directly linked to the parameter number. As a preliminary experiment, we manually assigned half channels for the hyperprior module in LIC models, reducing about 30% parameters in the model. The pruned models still kept similar RD performance to the original ones. This reveals that the hyperprior module in LIC models is highly redundant. In the meanwhile, LIC models with different reconstruction qualities require different amounts of parameters for the hyperprior module. Based on these phenomena, we propose a quality‐aware hyperprior pruning method that efficiently reduces the storage and memory cost of the hyperprior module and various context models. It consists of two parts. The first part is the pruning method itself, called enhanced ResRep on hyper path (ERHP). The second part is a quality‐aware threshold searching method, called pruning threshold searching (PTS), which prunes the hyperprior module based on the reconstruction qualities of LIC models. The experiments on various LIC models show that our methods reduce large volumes of storage cost (up to 74.6%) and memory cost (up to 41.5%), while keeping the performance the same before pruning. Ao Luo, Diego Fujii, Keisuke Nonaka, Heming Sun, Jiro Katto |
IET Image Process. | 1 |
| 2025 | MDLPCC: Misalignment-aware dynamic LiDAR point cloud compressionabstractLiDAR point cloud plays an important role in various real-world areas. It is usually generated as sequences by LiDAR on moving vehicles. Regarding the large data size of LiDAR point clouds, Dynamic Point Cloud Compression (DPCC) methods are developed to reduce transmission and storage data costs. However, most existing DPCC methods neglect the intrinsic misalignment in LiDAR point cloud sequences, limiting the rate–distortion (RD) performance. This paper proposes a Misalignment-aware Dynamic LiDAR Point Cloud Compression method (MDLPCC), which alleviates the misalignment problem in both macroscope and microscope. MDLPCC exploits a global transformation (GlobTrans) method to eliminate the macroscopic misalignment problem, which is the obvious gap between two continuous point cloud frames. MDLPCC also uses a spatial–temporal mixed structure to alleviate the microscopic misalignment, which still exists in the detailed parts of two point clouds after GlobTrans. The experiments on our MDLPCC show superior performance over existing point cloud compression methods. Ao Luo, Linxin Song, Keisuke Nonaka, Jinming Liu 0001, Kyohei Unno, Kohei Matsuzaki, Heming Sun, Jiro Katto |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | Event-Triggered Optimal Consensus Control for MASs With Multiple Constraints: A Flexible Performance ApproachabstractThis paper investigates the challenge of achieving event-triggered optimal consensus control for multiagent systems (MASs) with multiple constraints, encompassing saturation constraint at the input and performance constraint at the output. To achieve performance constraint while satisfying input saturation, a flexible prescribed performance method (FPPM) is designed. Utilizing non-negative signals generated by the improved auxiliary system to design the performance functions, the FPPM can adaptively adjust the performance constraint boundaries to ensure safe operation of the MASs with multiple constraints. Meanwhile, the proposed FPPM can achieve different performance behaviors by changing core parameters without the need to alter the control structure. Subsequently, a simplified reinforcement learning algorithm with actor-critic structure is integrated into the FPPM. By designing actor-critic neural networks and dynamic event-triggered mechanism, optimal consensus control for MASs under multiple constraint conditions is achieved cleverly while avoiding unnecessary communication transmissions. Finally, a simulation example verifies the effectiveness of the proposed method. Note to Practitioners—Considering the limitations of physical devices and the practical requirements for control performance, the input saturation constraint and performance constraint often coexist during the operation of practical systems, such as robotic systems, manipulator systems and aerospace systems. Therefore, this paper aims to design an event-triggered reinforcement learning algorithm for MASs with multiple constraints. To resolve the conflict problem caused by input saturation and performance constraint, a FPPM with adjustable performance functions is proposed. By flexibly adjusting the performance constraint boundaries, the coexistence problem of multiple constraints can be solved effectively. Meanwhile, the constructed FPPM framework can achieve various performance behaviors by adjusting parameters according to the practical application scenario without changing the controller structure. Additionally, the proposed event-triggered reinforcement learning algorithm can optimize the designed cost function and promote the utilization of communication resources. Ao Luo, Qi Zhou 0002, Hui Ma 0010, Hongyi Li 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Estimator-Based Reinforcement Learning Consensus Control for Multiagent Systems With Discontinuous ConstraintsabstractThis article focuses on the optimal consensus control problem for multiagent systems (MASs) with discontinuous constraints. The case of discontinuous constraints is a particular instance of state constraints, which has been studied less but occurs in many practical situations. Due to the discontinuous constraint boundaries, the traditional barrier function-based backstepping methods cannot be used directly. In response to this thorny problem, a novel constraint boundary reconstruction technique is proposed by designing a class of switch-like functions. The technique can convert discontinuous constraint boundaries into continuous ones, and it strictly proves that when the states satisfy the transformed constraint boundaries, the original constraints are also absolutely fulfilled. Meanwhile, with the aid of the barrier function and distributed event-triggered estimator, an improved coordinate transformation is constructed, which can remove the "feasibility condition" and simplify the controller design. In addition, by introducing prediction error and revised term into the learning process of neural networks (NNs), the optimal consensus problem is resolved by constructing a modified reinforcement learning strategy. Finally, the stability of the MASs is testified through the Lyapunov stability theory, and a simulation example verifies the effectiveness of the proposed method. Ao Luo, Hui Ma 0010, Hongru Ren, Hongyi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | SCP: Spherical-Coordinate-Based Learned Point Cloud CompressionabstractIn recent years, the task of learned point cloud compression has gained prominence. An important type of point cloud, LiDAR point cloud, is generated by spinning LiDAR on vehicles. This process results in numerous circular shapes and azimuthal angle invariance features within the point clouds. However, these two features have been largely overlooked by previous methodologies. In this paper, we introduce a model-agnostic method called Spherical-Coordinate-based learned Point cloud compression (SCP), designed to fully leverage the features of circular shapes and azimuthal angle invariance. Additionally, we propose a multi-level Octree for SCP to mitigate the reconstruction error for distant areas within the Spherical-coordinate-based Octree. SCP exhibits excellent universality, making it applicable to various learned point cloud compression techniques. Experimental results demonstrate that SCP surpasses previous state-of-the-art methods by up to 29.14% in point-to-point PSNR BD-Rate. Ao Luo, Linxin Song, Keisuke Nonaka, Kyohei Unno, Heming Sun, Masayuki Goto, Jiro Katto |
AAAI | 1 |
| 2024 | FastForensics: Efficient Two-Stream Design for Real-Time Image Manipulation Detection
Yangxiang Zhang, Yuezun Li, Ao Luo, Jiaran Zhou, Junyu Dong |
BMVC | 3 |
| 2024 | Efficient Meshflow and Optical Flow Estimation from Event CamerasabstractIn this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280×720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (39× faster) of our EEMFlow model compared to recent state-of-the-art flow methods. Our code is available at https://github.com/boomluo02/EEMFlow. Xinglong Luo, Ao Luo, Zhengning Wang, Chunyu Lin, Bing Zeng 0001, Shuaicheng Liu |
CVPR | 2 |
| 2024 | FlowDiffuser: Advancing Optical Flow Estimation with Diffusion ModelsabstractOptical flow estimation, a process of predicting pixel-wise displacement between consecutive frames, has commonly been approached as a regression task in the age of deep learning. Despite notable advancements, this de facto paradigm unfortunately falls short in generalization performance when trained on synthetic or constrained data. Pioneering a paradigm shift, we reformulate optical flow estimation as a conditional flow generation challenge, unveiling FlowDiffuser — a new family of optical flow models that could have stronger learning and generalization capabilities. FlowDiffuser estimates optical flow through a ‘noise-to-flow’ strategy, progressively eliminating noise from randomly generated flows conditioned on the provided pairs. To optimize accuracy and efficiency, our FlowDiffuser incorporates a novel Conditional Recurrent Denoising Decoder (Conditional-RDD), streamlining the flow estimation process. It incorporates a unique Hidden State Denoising (HSD) paradigm, effectively leveraging the information from previous time steps. Moreover, FlowDiffuser can be easily integrated into existing flow networks, leading to significant improvements in performance metrics compared to conventional implementations. Experiments on challenging benchmarks, including Sintel and KITTI, demonstrate the effectiveness of our FlowDiffuser with superior performance to existing state-of-the-art models. Code is available at https://github.com/LA30/FlowDiffuser. Ao Luo, Fan Yang 0054, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu |
CVPR | 1 |
| 2024 | RecDiffusion: Rectangling for Image Stitching with Diffusion ModelsabstractImage stitching from different captures often results in non-rectangular boundaries, which is often considered un-appealing. To solve non-rectangular boundaries, current solutions involve cropping, which discards image content, inpainting, which can introduce unrelated content, or warping, which can distort non-linear features and introduce artifacts. To overcome these issues, we introduce a novel diffusion-based learning framework, RecDiffusion, for image stitching rectangling. This framework combines Motion Diffusion Models (MDM) to generate motion fields, ef-fectively transitioning from the stitched image's irregular borders to a geometrically corrected intermediary. Fol-lowed by Content Diffusion Models (CDM) for image de-tail refinement. Notably, our sampling process utilizes a weighted map to identify regions needing correction during each iteration of CDM. Our RecDiffusion ensures geomet-ric accuracy and overall visual appeal, surpassing all pre-vious methods in both quantitative and qualitative measures when evaluated on public benchmarks. Code is released at https://github.com/haippp/RecDiffusion. Tianhao Zhou, Haipeng Li 0001, Ao Luo, Chen-Lin Zhang, Bing Zeng 0001, Shuaicheng Liu |
CVPR | 4 |
| 2024 | LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models
Hai Jiang 0006, Ao Luo, Xiaohong Liu 0001, Songchen Han, Shuaicheng Liu |
ECCV (48) | 2 |
| 2024 | FocusDiffuser: Perceiving Local Disparities for Camouflaged Object Detection
Xin Li 0079, Fan Yang 0054, Qiang Zhai, Ao Luo, Zicheng Jiao, Hong Cheng 0002 |
ECCV (53) | 5 |
| 2024 | MagicCartoon: 3D Pose and Shape Estimation for Bipedal Cartoon CharactersabstractThe 3D model can be estimated by regressing the pose and shape parameters from the image data of the digital model. The reconstruction of 3D cartoon characters poses a challenging task due to diverse visual representations and postural variations. This paper proposes a dual-branch structure named MagicCartoon for 3D bipedal cartoon character estimation, which models pose and shape independently through feature decoupling. Considering the correlation between category difference and shape parameters, a hybrid feature fusion technique is introduced, which integrates the global features of the original image with the corresponding local features expressed by the puzzle image, reducing the abstractness of understanding shape parameter differences. To semantically align image and geometric between feature space, a geometric-guided feedback loop is proposed in an iterative way, so that the pose of modeling results can be expressed consistently with the image. Moreover, a feature consistency loss is designed to augment the training data by incorporating the same character with different postures and the same posture of different characters. It enhances the correlation between the features extracted by the backbone network and the specific task. Experiments conducted on the 3DBiCar dataset demonstrate that MagicCartoon outperforms the state-of-the-art methods. Yu-Pei Song, Yuantong Liu, Xiao Wu 0001, Qi He 0007, Zhaoquan Yuan, Ao Luo |
ACM Multimedia | 6 |
| 2024 | Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain SchedulerabstractIn Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and accurately quantify category novelty, which is critical for applications in dynamic environments. Recently, meta-learning techniques have demonstrated superior results in OSDG, effectively orchestrating the meta-train and -test tasks by employing varied random categories and predefined domain partition strategies. These approaches prioritize a well-designed training schedule over traditional methods that focus primarily on data augmentation and the enhancement of discriminative feature learning.
The prevailing meta-learning models in OSDG typically utilize a predefined sequential domain scheduler to structure data partitions. However, a crucial aspect that remains inadequately explored is the influence brought by strategies of domain schedulers during training.
In this paper, we observe that an adaptive domain scheduler benefits more in OSDG compared with prefixed sequential and random domain schedulers. We propose the Evidential Bi-Level Hardest Domain Scheduler (EBiL-HaDS) to achieve an adaptive domain scheduler. This method strategically sequences domains by assessing their reliabilities in utilizing a follower network, trained with confidence scores learned in an evidential manner, regularized by max rebiasing discrepancy, and optimized in a bilevel manner. We verify our approach on three OSDG benchmarks, i.e., PACS, DigitsDG, and OfficeHome. The results show that our method substantially improves OSDG performance and achieves more discriminative embeddings for both the seen and unseen categories, underscoring the advantage of a judicious domain scheduler for the generalizability to unseen domains and unseen categories. The source code is publicly available at https://github.com/KPeng9510/EBiL-HaDS. Kunyu Peng, Di Wen 0006, Kailun Yang 0001, Ao Luo, Yufan Chen 0001, Jia Fu 0001, M. Saquib Sarfraz, Alina Roitberg, Rainer Stiefelhagen |
NeurIPS | 4 |
| 2024 | Reinforcement learning-based consensus control for MASs with intermittent constraints
Ao Luo, Qi Zhou 0002, Hongru Ren, Hui Ma 0010, Renquan Lu |
Neural Networks | 1 |
| 2024 | Observer-Based Consensus Control for MASs With Prescribed Constraints via Reinforcement Learning AlgorithmabstractIn this article, an adaptive optimal consensus control problem is studied for multiagent systems (MASs) with external disturbances, unmeasurable states, and prescribed constraints. First, by using neural networks (NNs), a composite observer is constructed to estimate the unmeasurable states and disturbances simultaneously. Then, the consensus error is guaranteed within a prescribed boundary by presenting an improved prescribed performance control (PPC) technique, and the initial conditions for the error are eliminated. In addition, the updating laws of actor-critic NNs are established by using a simplified reinforcement learning (RL) algorithm based on the uniqueness of optimal solution, and the asymmetric input saturation is resolved by designing auxiliary system instead of using nonquadratic cost functions in other optimal control methods. Finally, the boundedness of all signals in the closed-loop system is proved by using Lyapunov stability theory. The effectiveness of the proposed control method is verified by a simulation example. Ao Luo, Qi Zhou 0002, Hui Ma 0010, Hongyi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | DMHomo: Learning Homography with Diffusion ModelsabstractSupervised homography estimation methods face a challenge due to the lack of adequate labeled training data. To address this issue, we propose DMHomo , a diffusion model-based framework for supervised homography learning. This framework generates image pairs with accurate labels, realistic image content, and realistic interval motion, ensuring that they satisfy adequate pairs. We utilize unlabeled image pairs with pseudo labels such as homography and dominant plane masks, computed from existing methods, to train a diffusion model that generates a supervised training dataset. To further enhance performance, we introduce a new probabilistic mask loss, which identifies outlier regions through supervised training, and an iterative mechanism to optimize the generative and homography models successively. Our experimental results demonstrate that DMHomo effectively overcomes the scarcity of qualified datasets in supervised homography learning and improves generalization to real-world scenes. The code and dataset are available at GitHub ( https://github.com/lhaippp/DMHomo ). Haipeng Li 0001, Hai Jiang 0006, Ao Luo, Ping Tan 0002, Haoqiang Fan, Bing Zeng 0001, Shuaicheng Liu |
ACM Trans. Graph. | 3 |
| 2023 | Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-trainingabstractIn this paper, we introduce Cross-View Language Modeling, a simple and effective pretraining framework that unifies cross-lingual and cross-modal pre-training with shared architectures and objectives.Our approach is motivated by a key observation that cross-lingual and cross-modal pre-training share the same goal of aligning two different views of the same object into a common semantic space.To this end, the cross-view language modeling framework considers both multi-modal data (i.e., image-caption pairs) and multi-lingual data (i.e., parallel sentence pairs) as two different views of the same object, and trains the model to align the two views by maximizing the mutual information between them with conditional masked language modeling and contrastive learning.We pre-train CCLM, a Crosslingual Cross-modal Language Model, with the cross-view language modeling framework.Empirical results on IGLUE, a multi-lingual multi-modal benchmark, and two multi-lingual image-text retrieval datasets show that while conceptually simpler, CCLM significantly outperforms the prior state-of-the-art with an average absolute improvement of over 10%.Moreover, CCLM is the first multi-lingual multimodal pre-trained model that surpasses the translate-test performance of representative English vision-language models by zero-shot cross-lingual transfer.1 Yan Zeng 0003, Wangchunshu Zhou, Ao Luo, Ziming Cheng, Xinsong Zhang |
ACL (1) | 3 |
| 2023 | Explicit Motion Disentangling for Efficient Optical Flow EstimationabstractIn this paper, we propose a novel framework for optical flow estimation that achieves a good balance between performance and efficiency. Our approach involves disentangling global motion learning from local flow estimation, treating global matching and local refinement as separate stages. We offer two key insights: First, the multi-scale 4D cost-volume based recurrent flow decoder is computationally expensive and unnecessary for handling small displacement. With the separation, we can utilize lightweight methods for both parts and maintain similar performance. Second, a dense and robust global matching is essential for both flow initialization as well as stable and fast convergence for the refinement stage. Towards this end, we introduce EMD-Flow, a framework that explicitly separates global motion estimation from the recurrent refinement stage. We propose two novel modules: Multi-scale Motion Aggregation (MMA) and Confidence-induced Flow Propagation (CFP). These modules leverage cross-scale matching prior and self-contained confidence maps to handle the ambiguities of dense matching in a global manner, generating a dense initial flow. Additionally, a lightweight decoding module is followed to handle small displacements, resulting in an efficient yet robust flow estimation framework. We further conduct comprehensive experiments on standard optical flow benchmarks with the proposed framework, and the experimental results demonstrate its superior balance between performance and runtime. Code is available at https://github.com/gddcx/EMD-Flow. Changxing Deng, Ao Luo, Shaodan Ma, Jiangyu Liu, Shuaicheng Liu |
ICCV | 2 |
| 2023 | Learning Optical Flow from Event Camera with Rendered DatasetabstractWe study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real scenes by event cameras or synthesizing from images with pasted foreground objects. The former case can produce real event values but with calculated flow labels, which are sparse and inaccurate. The latter case can generate dense flow labels but the interpolated events are prone to errors. In this work, we propose to render a physically correct event-flow dataset using computer graphics models. In particular, we first create indoor and outdoor 3D scenes by Blender with rich scene content variations. Second, diverse camera motions are included for the virtual capturing, producing images and accurate flow labels. Third, we render high-framerate videos between images for accurate events. The rendered dataset can adjust the density of events, based on which we further introduce an adaptive density module (ADM). Experiments show that our proposed dataset can facilitate event-flow learning, whereas previous approaches when trained on our dataset can improve their performances constantly by a relatively large margin. In addition, event-flow pipelines when equipped with our ADM can further improve performances. Our code is available at https://github.com/boomluo02/ADMFlow. Xinglong Luo, Kunming Luo, Ao Luo, Zhengning Wang, Ping Tan 0002, Shuaicheng Liu |
ICCV | 3 |
| 2023 | GAFlow: Incorporating Gaussian Attention into Optical FlowabstractOptical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving consistent representations of the same category, optical flow raises extra demands for obtaining local discrimination and smoothness, which yet is not fully explored by existing approaches. In this paper, we push Gaussian Attention (GA) into the optical flow models to accentuate local properties during representation learning and enforce the motion affinity during matching. Specifically, we introduce a novel Gaussian-Constrained Layer (GCL) which can be easily plugged into existing Transformer blocks to highlight the local neighborhood that contains fine-grained structural information. Moreover, for reliable motion analysis, we provide a new Gaussian-Guided Attention Module (GGAM) which not only inherits properties from Gaussian distribution to instinctively revolve around the neighbor fields of each point but also is empowered to put the emphasis on contextually related regions during matching. Our fully-equipped model, namely Gaussian Attention Flow network (GAFlow), naturally incorporates a series of novel Gaussian-based modules into the conventional optical flow framework for reliable motion analysis. Extensive experiments on standard optical flow datasets consistently demonstrate the exceptional performance of the proposed approach in terms of both generalization ability evaluation and online benchmark testing. Code is available at https://github.com/LA30/GAFlow. Ao Luo, Fan Yang 0054, Xin Li 0005, Lang Nie, Chunyu Lin, Haoqiang Fan, Shuaicheng Liu |
ICCV | 1 |
| 2023 | PTS-LIC: Pruning Threshold Searching for Lightweight Learned Image CompressionabstractLearned Image Compression (LIC), which uses neural networks to compress images, has experienced significant growth in recent years. The hyperprior-module-based LIC model has achieved higher performance than classical codecs. However, the LIC models are too heavy (in calculation and parameter amounts) to apply to edge devices. To solve this problem, some former papers focus on structural pruning for LIC models. However, they either cause noticeable performance decrement or neglect the appropriate pruning threshold for each LIC model. These problems keep their pruning results sub-optimal. This paper proposes a Pruning Threshold Searching on the hyperprior module for different-quality LIC models. Our method removes most parameters and calculations while keeping the performance the same as the models before pruning. We removed at least 49.8% of parameters and 28.5% of calculations for the Channel-Wise-Context-Model-based models and 29.1% of parameters for the Cheng-2020 models. Ao Luo, Heming Sun, Jinming Liu 0001, Fangzheng Lin, Jiro Katto |
VCIP | 1 |
| 2023 | Robust Scene Parsing by Mining Supportive Knowledge From DatasetabstractScene parsing, or semantic segmentation, aims at labeling all pixels in an image with the predefined categories of things and stuff. Learning a robust representation for each pixel is crucial for this task. Existing state-of-the-art (SOTA) algorithms employ deep neural networks to learn (discover) the representations needed for parsing from raw data. Nevertheless, these networks discover desired features or representations only from the given image (content), ignoring more generic knowledge contained in the dataset. To overcome this deficiency, we make the first attempt to explore the meaningful supportive knowledge, including general visual concepts (i.e., the generic representations for objects and stuff) and their relations from the whole dataset to enhance the underlying representations of a specific scene for better scene parsing. Specifically, we propose a novel supportive knowledge mining module (SKMM) and a knowledge augmentation operator (KAO), which can be easily plugged into modern scene parsing networks. By taking image-specific content and dataset-level supportive knowledge into full consideration, the resulting model, called knowledge augmented neural network (KANN), can better understand the given scene and provide greater representational power. Experiments are conducted on three challenging scene parsing and semantic segmentation datasets: Cityscapes, Pascal-Context, and ADE20K. The results show that our KANN is effective and achieves better results than all existing SOTA methods. Ao Luo, Fan Yang 0054, Xin Li 0079, Yuezun Li, Zhicheng Jiao, Hong Cheng 0002, Siwei Lyu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Low-Light Image Enhancement with Wavelet-Based Diffusion ModelsabstractDiffusion models have achieved promising results in image restoration tasks, yet suffer from time-consuming, excessive computational resource consumption, and unstable restoration. To address these issues, we propose a robust and efficient Diffusion-based Low-Light image enhancement approach, dubbed DiffLL. Specifically, we present a wavelet-based conditional diffusion model (WCDM) that leverages the generative power of diffusion models to produce results with satisfactory perceptual fidelity. Additionally, it also takes advantage of the strengths of wavelet transformation to greatly accelerate inference and reduce computational resource usage without sacrificing information. To avoid chaotic content and diversity, we perform both forward diffusion and denoising in the training phase of WCDM, enabling the model to achieve stable denoising and reduce randomness during inference. Moreover, we further design a high-frequency restoration module (HFRM) that utilizes the vertical and horizontal details of the image to complement the diagonal information for better fine-grained restoration. Extensive experiments on publicly available real-world benchmarks demonstrate that our method outperforms the existing state-of-the-art methods both quantitatively and visually, and it achieves remarkable improvements in efficiency compared to previous diffusion-based methods. In addition, we empirically show that the application for low-light face detection also reveals the latent practical values of our method. Code is available at https://github.com/JianghaiSCU/Diffusion-Low-Light. Hai Jiang 0006, Ao Luo, Haoqiang Fan, Songchen Han, Shuaicheng Liu |
ACM Trans. Graph. | 2 |
| 2022 | Learning Optical Flow with Adaptive Graph ReasoningabstractEstimating per-pixel motion between video frames, known as optical flow, is a long-standing problem in video understanding and analysis. Most contemporary optical flow techniques largely focus on addressing the cross-image matching with feature similarity, with few methods considering how to explicitly reason over the given scene for achieving a holistic motion understanding. In this work, taking a fresh perspective, we introduce a novel graph-based approach, called adaptive graph reasoning for optical flow (AGFlow), to emphasize the value of scene/context information in optical flow. Our key idea is to decouple the context reasoning from the matching procedure, and exploit scene information to effectively assist motion estimation by learning to reason over the adaptive graph. The proposed AGFlow can effectively exploit the context information and incorporate it within the matching procedure, producing more robust and accurate results. On both Sintel clean and final passes, our AGFlow achieves the best accuracy with EPE of 1.43 and 2.47 pixels, outperforming state-of-the-art approaches by 11.2% and 13.6%, respectively. Code is publicly available at https://github.com/megvii-research/AGFlow. Ao Luo, Fan Yang 0054, Kunming Luo, Xin Li 0079, Haoqiang Fan, Shuaicheng Liu |
AAAI | 1 |
| 2022 | Learning Optical Flow with Kernel Patch AttentionabstractOptical flow is a fundamental method used for quantitative motion estimation on the image plane. In the deep learning era, most works treat it as a task of ‘matching of features’, learning to pull matched pixels as close as possible in feature space and vice versa. However, spatial affinity (smoothness constraint), another important component for motion understanding, has been largely overlooked. In this paper, we introduce a novel approach, called kernel patch attention (KPA), to better resolve the ambiguity in dense matching by explicitly taking the local context relations into consideration. Our KPA operates on each local patch, and learns to mine the context affinities for better inferring the flow fields. It can be plugged into contemporary optical flow architecture and empower the model to conduct comprehensive motion analysis with both feature similarities and spatial relations. On Sintel dataset, the proposed KPA-Flow achieves the best performance with EPE of 1.35 on clean pass and 2.36 on final pass, and it sets a new record of 4.60% in F1-all on KITTI-15 benchmark. Code is publicly available at https://github.com/megvii-research/KPAFlow. Ao Luo, Fan Yang 0054, Xin Li 0079, Shuaicheng Liu |
CVPR | 1 |
| 2022 | RealFlow: EM-Based Realistic Optical Flow Dataset Generation from Videos
Yunhui Han, Kunming Luo, Ao Luo, Jiangyu Liu, Haoqiang Fan, Guiming Luo, Shuaicheng Liu |
ECCV (19) | 3 |
| 2022 | Memory-Efficient Learned Image Compression with Pruned Hyperprior ModuleabstractLearned Image Compression (LIC) gradually became more and more famous in these years. The hyperprior-module-based LIC models have achieved remarkable rate-distortion performance. However, the memory cost of these LIC models is too large to actually apply them to various devices, especially to portable or edge devices. The parameter scale is directly linked with memory cost. In our research, we found the hyperprior module is not only highly over-parameterized, but also its latent representation contains redundant information. Therefore, we propose a novel pruning method named ERHP in this paper to efficiently reduce the memory cost of hyperprior module, while improving the network performance. The experiments show our method is effective, reducing at least 22.6% parameters in the whole model while achieving better rate-distortion performance. Ao Luo, Heming Sun, Jinming Liu 0001, Jiro Katto |
ICIP | 1 |
| 2022 | Attention-Based Deep Driving Model for Autonomous Vehicles with Surround-View CamerasabstractExperienced human drivers always make safe driving decisions by selectively observing the front, rear and side- view mirrors. Several end - to-end methods have been pro-posed to learn driving models with multi-view visual infor-mation. However, these benchmark methods lack semantic understanding of multi-view image contents, where human drivers usually reason these information for decision making with different visual region of interests. In this paper, we propose an attention-based deep learning method to learn a driving model with input of surround-view visual information and the route planner, in which a multi-view attention module is designed for obtaining region of interests from human drivers. We evaluate our model on the Drive360 dataset with comparison of benchmarking deep driving models. Results demonstrate that our model achieves a competitive accuracy in both steering angle and speed prediction than benchmarking methods. Code is available at https://githuh.com/jet-uestc/MVA-Net. Yang Zhao 0024, Rui Huang 0008, Boqi Li 0001, Ao Luo, Yaochen Li, Hong Cheng 0002 |
IROS | 5 |
| 2022 | ASFlow: Unsupervised Optical Flow Learning With Adaptive Pyramid SamplingabstractWe present an unsupervised optical flow estimation method by proposing an adaptive pyramid sampling in the deep pyramid network. Specifically, in the pyramid downsampling, we propose a Content-Aware Pooling (CAP) module, which promotes local feature gathering by avoiding cross region pooling, so that the learned features become more representative. In the pyramid upsampling, we propose an Adaptive Flow Upsampling (AFU) module, where cross edge interpolation can be avoided, producing sharp motion boundaries. Equipped with these two modules, our method achieves the best performance for unsupervised optical flow estimation on multiple leading benchmarks, including MPI-Sintel, KITTI 2012 and KITTI 2015. Particularly, we achieve EPE=1.5 on KITTI 2012 and F1=9.67% KITTI 2015, which outperform the previous state-of-the-art methods by 16.7% and 13.1%, respectively. Shuaicheng Liu, Kunming Luo, Ao Luo, Chuan Wang 0001, Fanman Meng, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | EFRNet: Efficient Feature Reconstructing Network for Real-Time Scene ParsingabstractIn this paper, we introduce a light-weight and powerful convolutional neural network, termed asefficient feature reconstructing network(EFRNet), for real-time scene parsing. Our key idea is to decompose the process of learning high-resolution representations into two stages: i) bottom-up codebook/coding matrix learning and ii) top-down feature reconstructing. Specifically, the bottom-up process focuses on learningimage-specificcodewords (codebook) using deep-layer features and generating a coding matrix with the shallow-layer feature map. In the top-down process, the learned codebook and coding matrix are used to rebuild high-resolution features via a lightweightfeature reconstructing operator(FRO). In addition, our EFRNet is constructed on a new building block, named efficient adaptive abstraction (EAA) block, to further reduce the overall network parameters and achieve a significant speed up. Extensive experiments are conducted on challenging benchmarks, such as CamVid and Cityscapes. The results show that EFRNet demonstrates state-of-the-art performance with an optimal balance between accuracy and speed. Xin Li 0079, Fan Yang 0054, Ao Luo, Zhicheng Jiao, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Probabilistic Model Distillation for Semantic CorrespondenceabstractSemantic correspondence is a fundamental problem in computer vision, which aims at establishing dense correspondences across images depicting different instances under the same category. This task is challenging due to large intra-class variations and a severe lack of ground truth. A popular solution is to learn correspondences from synthetic data. However, because of the limited intra-class appearance and background variations within synthetically generated training data, the model’s capability for handling “real” image pairs using such strategy is intrinsically constrained. We address this problem with the use of a novel Probabilistic Model Distillation (PMD) approach which transfers knowledge learned by a probabilistic teacher model on synthetic data to a static student model with the use of unlabeled real image pairs. A probabilistic supervision reweighting (PSR) module together with a confidence-aware loss (CAL) is used to mine the useful knowledge and alleviate the impact of errors. Experimental results on a variety of benchmarks show that our PMD achieves state-of-the-art performance. To demonstrate the generalizability of our approach, we extend PMD to incorporate stronger supervision for better accuracy – the probabilistic teacher is trained with stronger key-point supervision. Again, we observe the superiority of our PMD. The extensive experiments verify that PMD is able to infer more reliable supervision signals from the probabilistic teacher for representation learning and largely alleviate the influence of errors in pseudo labels. Code is available at https://github.com/fanyang587/PMD. Xin Li 0079, Deng-Ping Fan, Fan Yang 0054, Ao Luo, Hong Cheng 0002, Zicheng Liu 0001 |
CVPR | 4 |
| 2021 | WebSRC: A Dataset for Web-Based Structural Reading ComprehensionabstractWeb search is an essential way for humans to obtain information, but it's still a great challenge for machines to understand the contents of web pages.In this paper, we introduce the task of structural reading comprehension (SRC) on web.Given a web page and a question about it, the task is to find the answer from the web page.This task requires a system not only to understand the semantics of texts but also the structure of the web page.Moreover, we proposed Web-SRC, a novel Web-based Structural Reading Comprehension dataset.WebSRC consists of 400K question-answer pairs, which are collected from 6.4K web pages.Along with the QA pairs, corresponding HTML source code, screenshots, and metadata are also provided in our dataset.Each question in WebSRC requires a certain structural understanding of a web page to answer, and the answer is either a text span on the web page or yes/no.We evaluate various baselines on our dataset to show the difficulty of our task.We also investigate the usefulness of structural information and visual features.Our dataset and baselines have been publicly available 1 . Zihan Zhao 0001, Lu Chen 0002, Jiabao Ji, Ao Luo, Yuxuan Xiong, Kai Yu 0004 |
EMNLP (1) | 6 |
| 2021 | Uncertainty-Guided Transformer Reasoning for Camouflaged Object DetectionabstractSpotting objects that are visually adapted to their surroundings is challenging for both humans and AI. Conventional generic / salient object detection techniques are suboptimal for this task because they tend to only discover easy and clear objects, while overlooking the difficult-to-detect ones with inherent uncertainties derived from indistinguishable textures. In this work, we contribute a novel approach using a probabilistic representational model in combination with transformers to explicitly reason under uncertainties, namely uncertainty-guided transformer reasoning (UGTR), for camouflaged object detection. The core idea is to first learn a conditional distribution over the backbone's output to obtain initial estimates and associated uncertainties, and then reason over these uncertain regions with attention mechanism to produce final predictions. Our approach combines the benefits of both Bayesian learning and Transformer-based reasoning, allowing the model to handle camouflaged object detection by leveraging both deterministic and probabilistic information. We empirically demonstrate that our proposed approach can achieve higher accuracy than existing state-of-the-art models on CHAMELEON, CAMO and COD10K datasets. Code is available at https://github.com/fanyang587/UGTR. Fan Yang 0054, Qiang Zhai, Xin Li 0079, Rui Huang 0008, Ao Luo, Hong Cheng 0002, Deng-Ping Fan |
ICCV | 5 |
| 2021 | TemporalFusion: Temporal Motion Reasoning with Multi-Frame Fusion for 6D Object Pose Estimationabstract6D object pose estimation is an essential task in vision-based robotic grasping and manipulation. Prior works extract spatial features by fusing the RGB image and depth without considering the temporal motion information, limiting their performance in heavy occlusion robotic grasping scenarios. In this paper, we present an end-to-end model named TemporalFusion, which integrates the temporal motion information from RGB-D images for 6D object pose estimation. The core of proposed TemporalFusion model is to embed and fuse the temporal motion information from multi-frame RGB-D sequences, which could handle heavy occlusion in robotic grasping tasks. Furthermore, the proposed deep model can also obtain stable pose sequences, which is essential for real-time robotic grasping tasks. We evaluated the proposed method in the YCB-Video dataset, and experimental results show our model outperforms state-of-the-art approaches. Our code is available at https://github.com/mufengjun260/TemporalFusion21. Fengjun Mu, Rui Huang 0008, Ao Luo, Xin Li 0079, Jing Qiu 0004, Hong Cheng 0002 |
IROS | 3 |
| 2021 | Student Class Behavior Dataset: a video dataset for recognizing, detecting, and captioning students' behaviors in classroom scenes
Bo Sun 0006, Kaijie Zhao, Jun He 0009, Lejun Yu, Huanqing Yan, Ao Luo |
Neural Comput. Appl. | 7 |
| 2021 | EKENet: Efficient knowledge enhanced network for real-time scene parsing
Ao Luo, Fan Yang 0054, Xin Li 0079, Rui Huang 0008, Hong Cheng 0002 |
Pattern Recognit. | 1 |
| 2020 | Hybrid Graph Neural Networks for Crowd CountingabstractCrowd counting is an important yet challenging task due to the large scale and density variation. Recent investigations have shown that distilling rich relations among multi-scale features and exploiting useful information from the auxiliary task, i.e., localization, are vital for this task. Nevertheless, how to comprehensively leverage these relations within a unified network architecture is still a challenging problem. In this paper, we present a novel network structure called Hybrid Graph Neural Network (HyGnn) which targets to relieve the problem by interweaving the multi-scale features for crowd density as well as its auxiliary task (localization) together and performing joint reasoning over a graph. Specifically, HyGnn integrates a hybrid graph to jointly represent the task-specific feature maps of different scales as nodes, and two types of relations as edges: (i) multi-scale relations capturing the feature dependencies across scales and (ii) mutual beneficial relations building bridges for the cooperation between counting and localization. Thus, through message passing, HyGnn can capture and distill richer relations between nodes to obtain more powerful representations, providing robust and accurate results. Our HyGnn performs significantly well on four challenging datasets: ShanghaiTech Part A, ShanghaiTech Part B, UCF_CC_50 and UCF_QNRF, outperforming the state-of-the-art algorithms by a large margin. Ao Luo, Fan Yang 0054, Xin Li 0079, Dong Nie, Zhicheng Jiao, Shangchen Zhou, Hong Cheng 0002 |
AAAI | 1 |
| 2020 | Cascade Graph Neural Networks for RGB-D Salient Object Detection
Ao Luo, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Hong Cheng 0002, Siwei Lyu |
ECCV (12) | 1 |
| 2020 | Relative Floor Estimation for Indoor Co-navigation: A Machine Learning Approach
Chanxin Zhou, Ao Luo, Bang Wang 0001 |
GPC | 2 |
| 2020 | Fast Portrait Segmentation With Highly Light-Weight NetworkabstractIn this paper, we describe a fast and light-weight portrait segmentation method based on a new highly light-weight backbone (HLB) architecture. The core element of HLB is a bottleneck-based factorized block (BFB) that has much fewer parameters than existing alternatives while keeping good learning capacity. Consequently, the HLB-based portrait segmentation method can run faster than the existing methods yet retaining the competitive accuracy performance with state-of-the-arts. Experiments conducted on two benchmark datasets demonstrate the effectiveness and efficiency of our method. Yuezun Li, Ao Luo, Siwei Lyu |
ICIP | 2 |
| 2020 | Webly-supervised learning for salient object detection
Ao Luo, Xin Li 0079, Fan Yang 0054, Zhicheng Jiao, Hong Cheng 0002 |
Pattern Recognit. | 1 |
| 2019 | A CNN-based Approach to the Detection of SQL Injection AttacksabstractSQL injection has always been a major threat in the field of web application security. Traditional methods such as the rule-matching-based SQL injection detection solutions, which are inefficient to cope with the ever-changing SQL injection techniques and there is always a risk of bypassing variants. In this paper, we extract SQL injection attack related payloads from network flow and propose a SQL injection detection model based on Convolutional Neural Network (CNN), which can take the advantages of high-dimensional features of SQL injection behavior to deal with this issue. The proposed approach was tested in a real-traffic case study along with ModSecurity, which is the representative rule-matching-based method. The experimental results show that the CNN based model has higher accuracy, precision and recall rate, which validate its detection effectiveness and robustness against obfuscation of attacks. Ao Luo, Wenqing Fan |
ICIS | 1 |
| 2019 | Jintide®: A Hardware Security Enhanced Server CPU with Xeon® Cores under Runtime Surveillance by an In-Package Dynamically Reconfigurable ProcessorabstractThis article consists of a collection of slides from the author's conference presentation. Leibo Liu, Ao Luo, Guanhua Li, Jianfeng Zhu 0001, Gang Shan, Jianfeng Pan, Shouyi Yin, Shaojun Wei |
Hot Chips Symposium | 2 |
| 2019 | End-to-End Driving Model for Steering Control of Autonomous Vehicles with Future Spatiotemporal FeaturesabstractEnd-to-end deep learning has gained considerable interests in autonomous driving vehicles in both academic and industrial fields, especially in decision making process. One critical issue in decision making process of autonomous driving vehicles is steering control. Researchers has already trained different artificial neural networks to predict steering angle with front-facing camera data stream. However, existing end-to-end methods only consider the spatiotemporal relation on a single layer and lack the ability of extracting future spatiotemporal information. In this paper, we propose an end-to-end driving model based on Convolutional Long Short-Term Memory (Conv-LSTM) neural network with a Multi-scale Spatiotemporal Integration (MSI) module, which aiming to encode the spatiotemporal information from different scales for steering angle prediction. Moreover, we employ future sequential information to enhance spatiotemporal features of the end-to-end driving model. We demonstrate the efficiency of proposed end-to-end driving model on the public Udacity dataset with comparison of some existing methods. Experimental results show that the proposed model has better performances than other existing methods, especially in some complex scenarios. Furthermore, we evaluate the proposed driving model on a real-time autonomous vehicle, and results show that the proposed driving model is able to predict the steering angle with high accuracy compared to skilled human driver. Ao Luo, Rui Huang 0008, Hong Cheng 0002, Yang Zhao 0024 |
IROS | 2 |
| 2019 | Goal-Oriented Knowledge-Driven Neural Dialog Generation System
Ao Luo, Shengfeng Pan |
NLPCC (2) | 1 |
| 2018 | Robust spatial-temporal Bayesian view synthesis for video stitching with occlusion handling
Jianmei Su, Hong Cheng 0002, Lu Yang 0002, Ao Luo |
Mach. Vis. Appl. | 4 |
| 2010 | A VLSI design of sensor node for wireless image sensor networkabstractThis paper presents a single chip VLSI architecture of wireless image sensor node, which is constituted by an enhanced embedded 8051 microcontroller, a CMOS camera interface and hardware accelerators. The algorithms and control flows of the IEEE 802.15.4 MAC layer are accelerated by hardware, results in 45% less code size compared with the conventional software stack. An innovated CFA preprocessing algorithm and JPEG-LS compressing method is adopted and implemented by hardware, which has a minimal 46.3dB PSNR, an average compression ratio of about 3.0bit/pixel and an approximately 5fps at 16MHz system clock. Furthermore, low power design and techniques are employed to extend battery life, resulting in 60mW max system power consumption when the SoC is in full working mode (i.e. processor, image processing and wireless communication are active simultaneously) in 0.18μm CMOS process. Renyan Zhou, Leibo Liu, Shouyi Yin, Ao Luo, Xinkai Chen, Shaojun Wei |
ISCAS | 4 |