EDBT 2026 Demo / reviewers in the wild / expert
Tiejian Luo
dblp:00/6108
· DBLP profile ↗
54ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0002-9793-5878ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 since 2021Databases, data management, data science and information retrieval · 11Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structured Context Learning for Generic Event Boundary DetectionabstractGeneric Event Boundary Detection (GEBD) aims to identify moments in videos that humans perceive as event boundaries. This paper proposes a novel method for addressing this task, called Structured Context Learning, which introduces the Structured Partition of Sequence (SPoS) to provide a structured context for learning temporal information. Our approach is end-to-end trainable and flexible, not restricted to specific temporal models like GRU, LSTM, and Transformers. This flexibility enables our method to achieve a better speed-accuracy trade-off. Specifically, we apply SPoS to partition the input frame sequence and provide a structured context for the subsequent temporal model. Notably, SPoS’s overall computational complexity is linear with respect to the video length. We next calculate group similarities to capture differences between frames, and a lightweight fully convolutional network is utilized to determine the event boundaries based on the grouped similarity maps. To remedy the ambiguities of boundary annotations, we adapt the Gaussian kernel to preprocess the ground-truth event boundaries. Our proposed method has been extensively evaluated on the challenging Kinetics-GEBD, TAPOS, and shot transition detection datasets, demonstrating its superiority over existing state-of-the-art methods. Xin Gu 0003, Dexiang Hong, Libo Zhang 0001, Tiejian Luo, Longyin Wen, Heng Fan 0001 |
WACV | 6 |
| 2025 | Multi-Reward as Condition for Instruction-based Image EditingabstractHigh-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable Diffusion, DALL-E) which are not trained for image editing. Accordingly, these datasets suffer from inaccurate instruction following, poor detail preserving, and generation artifacts. In this paper, we propose to address the training data quality issue with multi-perspective reward data instead of refining the ground-truth image quality. 1) we first design a quantitative metric system based on best-in-class LVLM (Large Vision Language Model), i.e., GPT-4o in our case, to evaluate the generation quality from 3 perspectives, namely, instruction following, detail preserving, and generation quality. For each perspective, we collected quantitative score in $0\sim 5$ and text descriptive feedback on the specific failure points in ground-truth edited images, resulting in a high-quality editing reward dataset, i.e., RewardEdit20K. 2) We further proposed a novel training framework to seamlessly integrate the metric output, regarded as multi-reward, into editing models to learn from the imperfect training triplets. During training, the reward scores and text descriptions are encoded as embeddings and fed into both the latent space and the U-Net of the editing models as auxiliary conditions. During inference, we set these additional conditions to the highest score with no text description for failure points, to aim at the best generation outcome. 3) We also build a challenging evaluation benchmark with real-world images/photos and diverse editing instructions, named as Real-Edit. Experiments indicate that our multi-reward conditioned model outperforms its no-reward counterpart on two popular editing pipelines, i.e., InsPix2Pix and SmartEdit. Code is released at https://github.com/bytedance/Multi-Reward-Editing. Xin Gu 0003, Libo Zhang 0001, Longyin Wen, Tiejian Luo, Sijie Zhu |
ICLR | 6 |
| 2025 | Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video GroundingabstractTransformer has attracted increasing interest in spatio-temporal video grounding, or STVG, owing to its end-to-end pipeline and promising result. Existing Transformer-based STVG approaches often leverage a set of object queries, which are initialized simply using zeros and then gradually learn target position information via iterative interactions with multimodal features, for spatial and temporal localization. Despite simplicity, these zero object queries, due to lacking target-specific cues, are hard to learn discriminative target information from interactions with multimodal features in complicated scenarios (e.g., with distractors or occlusion), resulting in degradation. Addressing this, we introduce a novel $\textbf{T}$arget-$\textbf{A}$ware Transformer for $\textbf{STVG}$ ($\textbf{TA-STVG}$), which seeks to adaptively generate object queries via exploring target-specific cues from the given video-text pair, for improving STVG. The key lies in two simple yet effective modules, comprising text-guided temporal sampling (TTS) and attribute-aware spatial activation (ASA), working in a cascade. The former focuses on selecting target-relevant temporal cues from a video utilizing holistic text information, while the latter aims at further exploiting the fine-grained visual attribute information of the object from previous target-aware temporal cues, which is applied for object query initialization. Compared to existing methods leveraging zero-initialized queries, object queries in our TA-STVG, directly generated from a given video-text pair, naturally carry target-specific cues, making them adaptive and better interact with multimodal features for learning more discriminative information to improve STVG. In our experiments on three benchmarks, including HCSTVG-v1/-v2 and VidSTG, TA-STVG achieves state-of-the-art performance and significantly outperforms the baseline, validating its efficacy. Moreover, TTS and ASA are designed for general purpose. When applied to existing methods such as TubeDETR and STCAT, we show substantial performance gains, verifying its generality. Code is released at https://github.com/HengLan/TA-STVG. Xin Gu 0003, Yaojie Shen, Chenxi Luo, Tiejian Luo, Yan Huang 0002, Yuewei Lin, Heng Fan 0001, Libo Zhang 0001 |
ICLR | 4 |
| 2024 | Context-Guided Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding (or STVG) task aims at locating a spatio-temporal tube for a specific instance given a text query. Despite advancements, current methods easily suffer the distractors or heavy object appearance variations in videos due to insufficient object information from the text, leading to degradation. Addressing this, we propose a novel framework, context-guided STVG (CG-STVG), which mines discriminative instance context for object in videos and applies it as a supplementary guidance for target localization. The key of CG-STVG lies in two specially designed modules, including instance context generation (ICG), which focuses on discovering visual context information (in both appearance and motion) of the instance, and instance context refinement (ICR), which aims to improve the instance context from ICG by eliminating irrelevant or even harmful information from the context. During grounding, ICG, together with ICR, are deployed at each decoding stage of a transformer architecture for instance context learning. Particularly, instance context learned from one decoding stage is fed to the next stage, and leveraged as a guidance containing rich and discriminative object feature to enhance the target-awareness in decoding feature, which conversely benefits generating better new instance context to improve localization finally. Compared to existing methods, CG-STVG enjoys object information in text query and guidance from mined instance visual context for more accurate target localization. In experiments on HCSTVG-v1/-v2 and VidSTG, CG-STVG sets new state-of-the-arts in m_tIoU and m_vIoU on all of them, showing efficacy. Code is released at https://github.com/HengLan/CGSTVG. Xin Gu 0003, Heng Fan 0001, Yan Huang 0002, Tiejian Luo, Libo Zhang 0001 |
CVPR | 4 |
| 2024 | MaGIC: Multi-modality Guided Image CompletionabstractVanilla image completion approaches exhibit sensitivity to large missing regions, attributed to the limited availability of reference information for plausible generation. To mitigate this, existing methods incorporate the extra cue as guidance for image completion. Despite improvements, these approaches are often restricted to employing a *single modality* (e.g., *segmentation* or *sketch* maps), which lacks scalability in leveraging multi-modality for more plausible completion.
In this paper, we propose a novel, simple yet effective method for **M**ulti-mod**a**l **G**uided **I**mage **C**ompletion, dubbed **MaGIC**, which not only supports a wide range of single modality as the guidance (e.g., *text*, *canny edge*, *sketch*, *segmentation*, *depth*, and *pose*), but also adapts to arbitrarily customized combinations of these modalities (i.e., *arbitrary multi-modality*) for image completion.
For building MaGIC, we first introduce a modality-specific conditional U-Net (MCU-Net) that injects single-modal signal into a U-Net denoiser for single-modal guided image completion. Then, we devise a consistent modality blending (CMB) method to leverage modality signals encoded in multiple learned MCU-Nets through gradient guidance in latent space. Our CMB is *training-free*, thereby avoiding the cumbersome joint re-training of different modalities, which is the secret of MaGIC to achieve exceptional flexibility in accommodating new modalities for completion.
Experiments show the superiority of MaGIC over state-of-the-art methods and its generalization to various completion tasks. Hao Wang 0093, Tiejian Luo, Heng Fan 0001, Libo Zhang 0001 |
ICLR | 3 |
| 2024 | An Improved Consistent Hashing-Based Data Indexing Method for Distributed Photovoltaic Stations on HighwaysabstractEstablishing photovoltaic stations on highways is considered a promising way for fulfilling future carbon neutrality ambition in China. Data management of those operation data generated face challenge of reliability, time efficiency, balance and scalability. This paper proposes an improved consistent hashing-based indexing technique to deal with such problems. We allow multiple copies for one data item to improve system reliability. To deal with scalability issue, we calculate several candidate storage nodes and choose a fixed number of nodes with minimal storage load. The system performance and comparison result shows our proposal can deal with the challenges quite well. Zhu Wang 0004, Tiejian Luo |
INDIN | 2 |
| 2024 | CSFIR: Leveraging Code-Specific Features to Augment Information Retrieval in Low-Resource Code DatasetsabstractFrom search engines like Google to advanced applications such as Retrieval Augmented Generation integrating Large Language Model (LLM), Information Retrieval (IR) serves a crucial role. To facilitate the development of increasingly large and complex code program projects, researchers introduce IR systems into the code domain. Unfortunately, although IR systems achieve significant success in retrieving natural language corpus and query, they face challenges when tasked with retrieving corpus consisting of code sequences. Primarily, in practical applications, most code sequences lack corresponding natural language annotations, which are known as low-resource scenarios, hindering the training of neural network-based retriever in IR systems. Furthermore, modern IR systems often overlook the structural features of code sequences, which may be beneficial for understanding these sequences. Additionally, the length of most code sequences exceeds that of equivalent natural language expressions, complicating the processing of relationships between code and natural language sequences. To address these challenges, we propose a novel IR system CSFIR, which leverages Code-Speclflc Features to augment IR. For the prevalent issue of unlabeled code sequences in low-resource scenarios, we employ a supervised fine-tuned LLM as a generator to generate natural language queries for unlabeled code sequence. Subsequently, we extract structural features from the abstract syntax tree of code sequence using graph convolution networks and integrate these features to enhance the original retriever. Finally, given the adverse effects of lengthy code sequences on generators, we propose a subtree segmentation algorithm, which reduces the length of code sequences without compromising their original meaning, thereby enhancing the quality of queries generated by the generator. We conduct comparative experiments to ascertain the efficacy of our method. Regarding Recall@100, our CSFIR system improves from 90.12 in a traditional IR system to 96.18. Our code is available at https://github.com/tzy3141S/CSFIR.git. Zhenyu Tong, Chenxi Luo, Tiejian Luo |
SMC | 3 |
| 2024 | Local Compressed Video Stream Learning for Generic Event Boundary Detection
Libo Zhang 0001, Xin Gu 0003, Tiejian Luo, Heng Fan 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Robust Domain Adaptive Object Detection With Unified Multi-Granularity AlignmentabstractDomain adaptive detection aims to improve the generalization of detectors on target domain. To reduce discrepancy in feature distributions between two domains, recent approaches achieve domain adaption through feature alignment in different granularities via adversarial learning. However, they neglect the relationship between multiple granularities and different features in alignment, degrading detection. Addressing this, we introduce a unified multi-granularity alignment (MGA)-based detection framework for domain-invariant feature learning. The key is to encode the dependencies across different granularities including pixel-, instance-, and category-levels simultaneously to align two domains. Specifically, based on pixel-level features, we first develop an omni-scale gated fusion (OSGF) module to aggregate discriminative representations of instances with scale-aware convolutions, leading to robust multi-scale detection. Besides, we introduce multi-granularity discriminators to identify where, either source or target domains, different granularities of samples come from. Note that, MGA not only leverages instance discriminability in different categories but also exploits category consistency between two domains for detection. Furthermore, we present an adaptive exponential moving average (AEMA) strategy that explores model assessments for model update to improve pseudo labels and alleviate local misalignment problem, boosting detection robustness. Extensive experiments on multiple domain adaption scenarios validate the superiority of MGA over other approaches on FCOS and Faster R-CNN detectors. Libo Zhang 0001, Wenzhang Zhou, Heng Fan 0001, Tiejian Luo, Haibin Ling |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Text with Knowledge Graph Augmented Transformer for Video CaptioningabstractVideo captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for real-world applications, mainly due to the long-tail words challenge. In this paper, we propose a text with knowledge graph augmented transformer (TextKG)for video captioning. Notably, TextKG is a two-stream transformer, formed by the external stream and internal stream. The external stream is designed to absorb additional knowledge, which models the interactions between the additional knowledge, e.g., pre-built knowledge graph, and the built-in information of videos, e.g., the salient object regions, speech transcripts, and video captions, to mitigate the long-tail words challenge. Meanwhile, the internal stream is designed to exploit the multi-modality information in videos (e.g., the appearance of video frames, speech transcripts, and video captions) to ensure the quality of caption results. In addition, the cross attention mechanism is also used in between the two streams for sharing information. In this way, the two streams can help each other for more accurate results. Extensive experiments conducted on four challenging video captioning datasets, i.e., YouCookII, ActivityNet Captions, MSR-VTT, and MSVD, demonstrate that the proposed method performs favorably against the state-of-the-art methods. Specifically, the proposed TextKG method out-performs the best published results by improving 18.7% absolute CIDEr scores on the YouCookII dataset. Xin Gu 0003, Libo Zhang 0001, Tiejian Luo, Longyin Wen |
CVPR | 5 |
| 2023 | Unsupervised Domain Adaptive Detection with Network Stability AnalysisabstractDomain adaptive detection aims to improve the generality of a detector, learned from the labeled source domain, on the unlabeled target domain. In this work, drawing inspiration from the concept of stability from the control theory that a robust system requires to remain consistent both externally and internally regardless of disturbances, we propose a novel framework that achieves unsupervised domain adaptive detection through stability analysis. In specific, we treat discrepancies between images and regions from different domains as disturbances, and introduce a novel simple but effective Network Stability Analysis (NSA) framework that considers various disturbances for domain adaptation. Particularly, we explore three types of perturbations including heavy and light image-level disturbances and instance-level disturbance. For each type, NSA performs external consistency analysis on the outputs from raw and perturbed images and/or internal consistency analysis on their features, using teacher-student models. By integrating NSA into Faster R-CNN, we immediately achieve state-of-the-art results. In particular, we set a new record of 52.7% mAP on Cityscapes-to-FoggyCityscapes, showing the potential of NSA for domain adaptive detection. It is worth noticing, our NSA is designed for general purpose, and thus applicable to one-stage detection model (e.g., FCOS) besides the adopted one, as shown by experiments. Code is released at https://github.com/tiankongzhang/NSA. Wenzhang Zhou, Heng Fan 0001, Tiejian Luo, Libo Zhang 0001 |
ICCV | 3 |
| 2023 | A Semantic and Structural Transformer for Code Summarization GenerationabstractCurrently most methods cast code summarization generation as a machine translation task. Wherein the Transformer framework is a representative among them. Thanks to the attention mechanism in the Transformer, such a framework has achieved the state-of-the-art performance. Unfortunately, the Transformer encounters a series of challenges when generalizing to code summarization generation domain. Compared with natural language, code sequence is characterized by more complex multi-modal features, and difficult to extract these features only by the original Transformer structure. To further improve the performance, we make full use of code semantic and structural information in abstract syntax tree to build a simple yet effective framework, which consists of self-attention and graph based module to integrate code semantic information and syntax tree structure information. Besides, to compensate for the insufficiency of Transformer in encoding local features, we present a well-designed local RNN module. Extensive experiments show that the proposed method performs on par with the state-of-the-art methods on two public benchmarks, including Java and Python datasets. The comprehensive ablation studies further demonstrate the effectiveness of architecture design choices. The source code is released at https://github.com/tzv314159/SSTrans.git. Ruyi Ji, Zhenyu Tong, Tiejian Luo, Jing Liu 0001, Libo Zhang 0001 |
IJCNN | 3 |
| 2023 | Collaborative three-stream transformers for video captioning
Hao Wang 0093, Libo Zhang 0001, Heng Fan 0001, Tiejian Luo |
Comput. Vis. Image Underst. | 4 |
| 2022 | End-to-End Compressed Video Representation Learning for Generic Event Boundary DetectionabstractGeneric event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be decoded before feeding into the network, which demands considerable computational power and storage space. To that end, we propose a new end-to-end compressed video representation learning for event boundary detection that leverages the rich information in the compressed domain, i.e., RGB, motion vectors, residuals, and the internal group of pictures (GOP) structure, withoutfully decoding the video. Specifically, we first use the Con-vNets to extract features of the I-frames in the Gaps. After that, a light-weight spatial-channel compressed encoder is designed to compute the feature representations of the P-frames based on the motion vectors, residuals and representations of their dependent I-frames. A temporal contrastive module is proposed to determine the event boundaries of video sequences. To remedy the ambiguities of annotations and speed up the training process, we use the Gaussian kernel to preprocess the ground-truth event boundaries. Extensive experiments conducted on the Kinetics-GEBD dataset demonstrate that the proposed method achieves comparable results to the state-of-the-art methods with 4.5 x faster running speed. Longyin Wen, Dexiang Hong, Tiejian Luo, Libo Zhang 0001 |
CVPR | 5 |
| 2022 | Multi-Granularity Alignment Domain Adaptation for Object DetectionabstractDomain adaptive object detection is challenging due to distinctive data distribution between source domain and target domain. In this paper, we propose a unified multi-granularity alignment based object detection framework towards domain-invariant feature learning. To this end, we encode the dependencies across different granularity perspectives including pixel-, instance-, and category-levels simultaneously to align two domains. Based on pixel-level feature maps from the backbone network, we first develop the omniscale gated fusion module to aggregate discriminative representations of instances by scale-aware convolutions, leading to robust multi-scale object detection. Meanwhile, the multi-granularity discriminators are proposed to identify which domain different granularities of samples (i.e., pixels, instances, and categories) come from. Notably, we leverage not only the instance discriminability in different categories but also the category consistency between two domains. Extensive experiments are carried out on multiple domain adaptation scenarios, demonstrating the effectiveness of our framework over state-of-the-art algorithms on top of anchor-free FCOS and anchor-based Faster R-CNN detectors with different backbones. Wenzhang Zhou, Dawei Du, Libo Zhang 0001, Tiejian Luo |
CVPR | 4 |
| 2022 | Unbiased Multi-modality Guidance for Image Inpainting
Dawei Du, Libo Zhang 0001, Tiejian Luo |
ECCV (16) | 4 |
| 2022 | High-Fidelity Image Inpainting with GAN Inversion
Libo Zhang 0001, Heng Fan 0001, Tiejian Luo |
ECCV (16) | 4 |
| 2022 | AUGER: automatically generating review comments with pre-training modelsabstractCode review is one of the best practices as a powerful safeguard for software quality. In practice, senior or highly skilled reviewers inspect source code and provide constructive comments, consider- ing what authors may ignore, for example, some special cases. The collaborative validation between contributors results in code being highly qualified and less chance of bugs. However, since personal knowledge is limited and varies, the efficiency and effectiveness of code review practice are worthy of further improvement. In fact, it still takes a colossal and time-consuming effort to deliver useful review comments. This paper explores a synergy of multiple practical review comments to enhance code review and proposes AUGER (AUtomatically GEnerating Review comments): a review comments generator with pre-training models. We first collect empirical review data from 11 notable Java projects and construct a dataset of 10,882 code changes. By leveraging Text-to-Text Transfer Transformer (T5) models, the framework synthesizes valuable knowledge in the training stage and effectively outperforms baselines by 37.38% in ROUGE-L. 29% of our automatic review comments are considered useful according to prior studies. The inference generates just in 20 seconds and is also open to training further. Moreover, the performance also gets improved when thoroughly analyzed in case study. Lingwei Li, Li Yang 0015, Huaxi Jiang, Tiejian Luo, Zihan Hua, Geng Liang, Chun Zuo |
ESEC/SIGSOFT FSE | 5 |
| 2022 | Context-guided feature enhancement network for automatic check-out
Yihan Sun 0002, Tiejian Luo, Zhen Zuo |
Neural Comput. Appl. | 2 |
| 2021 | Non-deterministic and emotional chatting machine: learning emotional conversation generation using conditional variational autoencoders
Kaichun Yao, Libo Zhang 0001, Tiejian Luo, Dawei Du |
Neural Comput. Appl. | 3 |
| 2021 | SiamCAN: Real-Time Visual Tracking Based on Siamese Center-Aware NetworkabstractIn this article, we present a novel Siamese center-aware network (SiamCAN) for visual tracking, which consists of the Siamese feature extraction subnetwork, followed by the classification, regression, and localization branches in parallel. The classification branch is used to distinguish the target from background, and the regression branch is introduced to regress the bounding box of the target. To reduce the impact of manually designed anchor boxes to adapt to different target motion patterns, we design the localization branch to localize the target center directly to assist the regression branch generating accurate results. Meanwhile, we introduce the global context module into the localization branch to capture long-range dependencies for more robustness to large displacements of the target. A multi-scale learnable attention module is used to guide these three branches to exploit discriminative features for better performance. Extensive experiments on 9 challenging benchmarks, namely VOT2016, VOT2018, VOT2019, OTB100, LTB35, LaSOT, TC128, UAV123 and VisDrone-SOT2019 demonstrate that SiamCAN achieves leading accuracy with high efficiency. Our source code is available at https://isrc.iscas.ac.cn/gitlab/research/siamcan. Wenzhang Zhou, Longyin Wen, Libo Zhang 0001, Dawei Du, Tiejian Luo |
IEEE Trans. Image Process. | 5 |
| 2021 | Iterative Knowledge Distillation for Automatic Check-OutabstractAutomatic Check-Out (ACO) provides an object detection based mechanism for retailers to process the purchases of customers automatically. However, it suffers a lot from the domain shift problem because of different data distribution between the single item in training exemplar images and mixed items in testing checkout images. In this paper, we propose a new iterative knowledge distillation method to solve the domain adaptation problem for this task. First, we develop a new augmentation data strategy to generate synthesized checkout images. It can extract segmented items from the training images by the coarse-to-fine strategy and filter items with unrealistic poses by pose pruning. Second, we propose a dual pyramid scale network (DPSNet) to exploit the multi-scale feature representation in joint detection and counting views. Third, the iterative knowledge distillation training strategy is developed to make full use of both image-level and instance-level samples to narrow the semantic gap between source domain and target domain. Extensive experiments on the large-scale Retail Product Checkout (RPC) dataset show the proposed DPSNet can achieve state-of-the-art performance compared with existing methods. The source codes can be found athttps://isrc.iscas.ac.cn/gitlab/research/dpsnet. Libo Zhang 0001, Dawei Du, Tiejian Luo |
IEEE Trans. Multim. | 5 |
| 2020 | Learning to Infer User Hidden States for Online Sequential AdvertisingabstractTo drive purchase in online advertising, it is of the advertiser's great interest to optimize the sequential advertising strategy whose performance and interpretability are both important. The lack of interpretability in existing deep reinforcement learning methods makes it not easy to understand, diagnose and further optimize the strategy.In this paper, we propose our Deep Intents Sequential Advertising (DISA) method to address these issues. The key part of interpretability is to understand a consumer's purchase intent which is, however, unobservable (called hidden states). In this paper, we model this intention as a latent variable and formulate the problem as a Partially Observable Markov Decision Process (POMDP) where the underlying intents are inferred based on the observable behaviors. Large-scale industrial offline and online experiments demonstrate our method's superior performance over several baselines. The inferred hidden states are analyzed, and the results prove the rationality of our inference. Zhaoqing Peng, Junqi Jin, Yaodong Yang 0001, Rui Luo 0001, Jun Wang 0012, Weinan Zhang 0001, Chuan Yu 0002, Tiejian Luo, Han Li 0005, Jian Xu 0015, Kun Gai |
CIKM | 11 |
| 2020 | Spatial Attention Pyramid Network for Unsupervised Domain Adaptation
Dawei Du, Libo Zhang 0001, Longyin Wen, Tiejian Luo, Pengfei Zhu 0001 |
ECCV (13) | 5 |
| 2020 | Towards interpretable and robust hand detection via pixel-wise prediction
Libo Zhang 0001, Tiejian Luo, Lili Tao |
Pattern Recognit. | 3 |
| 2020 | Dual Encoding for Abstractive Text SummarizationabstractRecurrent neural network-based sequence-to-sequence attentional models have proven effective in abstractive text summarization. In this paper, we model abstractive text summarization using a dual encoding model. Different from the previous works only using a single encoder, the proposed method employs a dual encoder including the primary and the secondary encoders. Specifically, the primary encoder conducts coarse encoding in a regular way, while the secondary encoder models the importance of words and generates more fine encoding based on the input raw text and the previously generated output text summarization. The two level encodings are combined and fed into the decoder to generate more diverse summary that can decrease repetition phenomenon for long sequence generation. The experimental results on two challenging datasets (i.e., CNN/DailyMail and DUC 2004) demonstrate that our dual encoding model performs against existing methods. Kaichun Yao, Libo Zhang 0001, Dawei Du, Tiejian Luo, Lili Tao |
IEEE Trans. Cybern. | 4 |
| 2019 | Scale Invariant Fully Convolutional Network: Detecting Hands EfficientlyabstractExisting hand detection methods usually follow the pipeline of multiple stages with high computation cost, i.e., feature extraction, region proposal, bounding box regression, and additional layers for rotated region detection. In this paper, we propose a new Scale Invariant Fully Convolutional Network (SIFCN) trained in an end-to-end fashion to detect hands efficiently. Specifically, we merge the feature maps from high to low layers in an iterative way, which handles different scales of hands better with less time overhead comparing to concatenating them simply. Moreover, we develop the Complementary Weighted Fusion (CWF) block to make full use of the distinctive features among multiple layers to achieve scale invariance. To deal with rotated hand detection, we present the rotation map to get rid of complex rotation and derotation layers. Besides, we design the multi-scale loss scheme to accelerate the training process significantly by adding supervision to the intermediate layers of the network. Compared with the state-of-the-art methods, our algorithm shows comparable accuracy and runs a 4.23 times faster speed on the VIVA dataset and achieves better average precision on Oxford hand detection dataset at a speed of 62.5 fps. Dawei Du, Libo Zhang 0001, Tiejian Luo, Feiyue Huang, Siwei Lyu |
AAAI | 4 |
| 2019 | Data Priming Network for Automatic Check-OutabstractAutomatic Check-Out (ACO) receives increased interests in recent years. An important component of the ACO system is the visual item counting, which recognizes the categories and counts of the items chosen by the customers. However, the training of such a system is challenged by the domain adaptation problem, in which the training data are images from isolated items while the testing images are for collections of items. Existing methods solve this problem with data augmentation using synthesized images, but the image synthesis leads to unreal images that affect the training process. In this paper, we propose a new data priming method to solve the domain adaptation problem. Specifically, we first use pre-augmentation data priming, in which we remove distracting background from the training images using the coarse-to-fine strategy and select images with realistic view angles by the pose pruning method. In the post-augmentation step, we train a data priming network using detection and counting collaborative learning, and select more reliable images from testing data to fine-tune the final visual item tallying network. Experiments on the large scale Retail Product Checkout (RPC) dataset demonstrate the superiority of the proposed method, i.e., we achieve 80.51% checkout accuracy compared with 56.68% of the baseline methods. The source codes can be found in https://isrc.iscas.ac.cn/gitlab/research/acm-mm-2019-ACO. Dawei Du, Libo Zhang 0001, Tiejian Luo, Qi Tian 0001, Longyin Wen, Siwei Lyu |
ACM Multimedia | 4 |
| 2019 | The optimization of sum-product network structure learning
Tiejian Luo |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Learning to Communicate via Supervised Attentional Message ProcessingabstractMany tasks in AI require the collaboration of multiple agents. Generally, these agents cooperate with each other by message-passing communication. However, agents may suffer from being overwhelmed by massive received messages and have difficulties in obtaining useful information. To this end, we use an attention-based message processing (AMP) method to model agents' interactions by considering the relevance of each received message. To improve the efficiency of learning correct interactions, a supervised variant SAMP is then proposed to directly optimize the attentional weights in AMP with a target auxiliary interaction matrix from the environment. The empirical results demonstrate our proposal outperforms other competing multi-agent methods in "predator-prey-toxin" domain, and prove the superiority of SAMP in correctly guiding the optimization of attentional weights in AMP. Zhaoqing Peng, Libo Zhang 0001, Tiejian Luo |
CASA | 3 |
| 2018 | Teaching Machines to Ask QuestionsabstractWe propose a novel neural network model that aims to generate diverse and human-like natural language questions. Our model not only directly captures the variability in possible questions by using a latent variable, but also generates certain types of questions by introducing an additional observed variable. We deploy our model in the generative adversarial network (GAN) framework and modify the discriminator which not only allows evaluating the question authenticity, but predicts the question type. Our model is trained and evaluated on a question-answering dataset SQuAD, and the experimental results shown the proposed model is able to generate diverse and readable questions with the specific attribute. Kaichun Yao, Libo Zhang 0001, Tiejian Luo, Lili Tao |
IJCAI | 3 |
| 2018 | Multi-agent Communication with Attentional and Recurrent Message IntegrationabstractEffective communication is significant for solving cooperative tasks in multi-agent domain. Agents coordinate their behaviors by appropriately modeling the communication signals or messages sent from others. To this end, agents are required to filter noise and obtain useful information from received messages, and learn to adapt to the dynamics of messages number. In this paper, we propose an attentional and recurrent message integration method (ARMI) that handles the dynamics by recurrently decoding messages, and performs attentional integration based on the relevance of each message. We evaluate our proposal on a new “predator-prey-toxin” environment where the number of agents changes, and the results outperform other competing multi-agent methods. Further investigations are also done to prove the superiority of ARMI in collaborating agents' behaviors for complex tasks and establishing interpretable communication protocol. Zhaoqing Peng, Libo Zhang 0001, Tiejian Luo |
ISCC | 3 |
| 2018 | Evaluation for Two Bloom Filters' Configuration
Chenxi Luo, Zhu Wang 0004, Tiejian Luo |
PDCAT | 3 |
| 2018 | Deep reinforcement learning for extractive document summarization
Kaichun Yao, Libo Zhang 0001, Tiejian Luo |
Neurocomputing | 3 |
| 2017 | Learning Path Generation Method Based on Migration Between Concepts
Libo Zhang 0001, Tiejian Luo |
KSEM | 3 |
| 2016 | A fast filter tracker against serious occlusionabstractMany tracking algorithms applied in medical image processing, such as observing the movement of cells, have a great improvement in accuracy and robustness. However, it is difficult to deal with the large area occlusion and complete occlusion. In this paper, we propose a fast scale adaptive tracking algorithm based on correlation filtering. Except tracking the change of the target scale quickly, our method can also deal with the problem of large area occlusion and the complete disappearance of the target. Compared with the outstanding scale adaptive tracking method, the proposed method demonstrates higher performances in terms of the accuracy of tracking the target and the real-time performance. Libo Zhang 0001, Tiejian Luo, Yihan Sun 0002 |
BIBM | 2 |
| 2016 | Semantic analysis based on human thought patternabstractSemantic analysis is an important component of recommendation systems and information retrieval in computer aided detection. Previous researches have made certain breakthroughs in disease diagnosis and drugs recommended by semantic analysis. We propose a bilateral shortest paths method for computing semantic relatedness based on the human thought patterns for making sufficient use of the hyperlink structure. The proposed novel method exploits bilateral shortest paths method to calculate word similarity, and employs the method of matrix partition to calculate text similarity. Finally, an evaluation based on WS353-Ex and Lee datasets is carried out and the result shows that we obtain effective performance. Libo Zhang 0001, Tiejian Luo, Yihan Sun 0002 |
BIBM | 2 |
| 2016 | A novel saliency detection method via manifold ranking and compactness priorabstractFor improving the performance of saliency detection, several algorithms used graph construction have achieved excellent results. This paper proposes a novel bottom-up approach of saliency detection, which takes the advantages of both prior background and compactness. At first, we optimize the image boundary selections, by removing erroneous sections with a fixed threshold, to achieve more accurate saliency estimation results. The saliency map obtained by ranking with background queries can be optimized with compactness prior. The objects of salient are connected regions which are group together, with a compact form which are spatial distributed. Compared to the 8 state-of-the-art saliency detection approaches, our experimental results which test on the three public datasets show that the proposed algorithm improves accuracy and robustness significantly. This algorithm can find its potential applications in many different areas, but it is best suit for medical science and technologies because of high accuracy requirements. It can be used in the medical imaging processing to accurately differentiate tumor from bones, muscles and fats. Libo Zhang 0001, Zakir Ullah, Yihan Sun 0002, Tiejian Luo |
BIBM | 4 |
| 2016 | MLPF algorithm for tracking fast moving target against light interferenceabstractIn order to deal with the difficulty of tracking the fast moving aerial targets with light interference, we propose an improved particle tracking algorithm named multi-layers particle filter (MLPF). In MLPF, the particles are divided into three categories: the main particles (M-particles), the subordinate particles (S-particles) and the regenerate particles (R-particles). In the phase of resampling and state estimating, only M-particles are involved, then the R-particles are generated and considered as new S-particles in the next cycle. To a certain extent, our algorithm maintains the diversity of particles and reduces the computation time. Besides, MLPF has significant improvements on overcoming the tracing error after the sudden disappearance of the target and solving the degradation of particles. We demonstrate effectiveness of our proposed algorithm through systematic experiments. Experimental results show MLPF has better tracking effect compared to the traditional particle filter (PF) when the target is moving fast and affected by light interference. In the first experiment, the running time has been reduced from 47s to 21s while the precision increased from 64% to 96%. And for the second experiment, the running time has been reduced from 237s to 121s while precision increased from 46% to 89%. Libo Zhang 0001, Yuanqiang Cai, Zakir Ullah, Tiejian Luo |
ICPR | 4 |
| 2016 | TACE: A Toolkit for Analyzing Concept Evolution in Computing CurriculaabstractEffective teaching and learning requires to knowing whether some concepts are needed to grasping.Our investigation for ACM Computing Curricula from 1991 to 2013 shows that the numbers of concepts are increased by 5 times.That phenomenon makes learners harder to distinguish which concept is update or not.In this paper, we develop a solution to explore the body of knowledge of computing.The proposed toolkit uses a graph model to represent the disciplinary knowledge structure.The analytic results by TACE give us insight for the various subjects in computing discipline.Our findings show that 61.3% concepts in CC1991 are obsolete.In CC2001, the proportion of obsolete concepts drops to 11.5%, and in CS2008 it is 16.8%.The OS, IM, DS, AL's knowledge areas are more stable than CN, NC, GV, AR.The TACE's framework is highly modular, adaptive and extendible for analyzing other discipline's curricula. Tiejian Luo, Libo Zhang 0001 |
SEKE | 1 |
| 2016 | Training query filtering for semi-supervised learning to rank with pseudo labels
Xin Zhang 0073, Ben He 0001, Tiejian Luo |
World Wide Web | 3 |
| 2015 | Selecting Training Data for Learning-Based Twitter Search
Dongxing Li, Ben He 0001, Tiejian Luo, Xin Zhang 0073 |
ECIR | 3 |
| 2013 | Clustering-based transduction for learning a ranking model with limited human labelsabstractTransductive learning is a semi-supervised learning paradigm that can leverage unlabeled data by creating pseudo labels for learning a ranking model, when there is only limited or no training examples available. However, the effectiveness of transductive learning in information retrieval (IR) can be hindered by the low quality pseudo labels. To this end, we propose to incorporate a two-step k-means clustering algorithm to select the high quality training queries for generating the pseudo labels. In particular, the first step selects the high-quality queries for which the relevant documents are highly coherent as indicated by the clustering results. The second step then selects the initial training examples for the transductive learning that iteratively aggregating the pseudo examples. Finally, the learning to rank (LTR) algorithms are applied to learn the ranking model using the pseudo training examples created by the transductive learning process. Our proposed approach is particularly suitable for applications where there is only little or no human labels available as it does not necessarily involve the use of relevance assessments information or human efforts. Experimental results on the standard TREC Tweets11 collection show that our proposed approach outperforms strong baselines, namely the conventional applications of learning to rank algorithms using human labels for the training and transductive learning using all the queries available. Xin Zhang 0073, Ben He 0001, Tiejian Luo, Dongxing Li, Jungang Xu |
CIKM | 3 |
| 2013 | Sponsored Search Ad Selection by Keyword Structure Analysis
Kai Hui 0001, Bin Gao 0001, Ben He 0001, Tiejian Luo |
ECIR | 4 |
| 2013 | Utilizing term proximity for blog post retrievalabstractTerm proximity is effective for many information retrieval (IR) research fields yet remains unexplored in blogosphere IR. The blogosphere is characterized by large amounts of noise, including incohesive, off‐topic content and spam. Consequently, the classical bag‐of‐words unigram IR models are not reliable enough to provide robust and effective retrieval performance. In this article, we propose to boost the blog postretrieval performance by employing term proximity information. We investigate a variety of popular and state‐of‐the‐art proximity‐based statistical IR models, including a proximity‐based counting model, the Markov random field (MRF) model, and the divergence from randomness (DFR) multinomial model. Extensive experimentation on the standard TREC Blog06 test dataset demonstrates that the introduction of term proximity information is indeed beneficial to retrieval from the blogosphere. Results also indicate the superiority of the unordered bi‐gram model with the sequential‐dependence phrases over other variants of the proximity‐based models. Finally, inspired by the effectiveness of proximity models, we extend our study by exploring the proximity evidence between uery terms and opinionated terms. The consequent opinionated proximity model shows promising performance in the experiments. Ben He 0001, Tiejian Luo |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2012 | Query-biased learning to rank for real-time twitter searchabstractBy incorporating diverse sources of evidence of relevance, learning to rank has been widely applied to real-time Twitter search, where users are interested in fresh relevant messages. Such approaches usually rely on a set of training queries to learn a general ranking model, which we believe that the benefits brought by learning to rank may not have been fully exploited as the characteristics and aspects unique to the given target queries are ignored. In this paper, we propose to further improve the retrieval performance of learning to rank for real-time Twitter search, by taking the difference between queries into consideration. In particular, we learn a query-biased ranking model with a semi-supervised transductive learning algorithm so that the query-specific features, e.g. the unique expansion terms, are utilized to capture the characteristics of the target query. This query-biased ranking model is combined with the general ranking model to produce the final ranked list of tweets in response to the given target query. Extensive experiments on the standard TREC Tweets11 collection show that our proposed query-biased learning to rank approach outperforms strong baseline, namely the conventional application of the state-of-the-art learning to rank algorithms. Xin Zhang 0073, Ben He 0001, Tiejian Luo, Baobin Li |
CIKM | 3 |
| 2012 | Transductive Learning for Real-Time Twitter Search
Xin Zhang 0073, Ben He 0001, Tiejian Luo |
ICWSM | 3 |
| 2011 | Relevance weighting using within-document term statisticsabstractWith the rapid development of the information technology, there exists the difficulty in deploying state-of-the-art retrieval models in environments such as peer-to-peer networks and pervasive computing, where it is expensive or even infeasible to maintain the global statistics. To this end, this paper presents an investigation in the validity of different statistical assumptions of term distributions. Based on the findings in this investigation, a variety of weighting models, called NG (standing for "no global statistics") models, are derived from the Divergence from Randomness framework, in which only the within-document statistics are used in the relevance weighting. Compared to the state-of-the-art weighting models in extensive experiments on various standard TREC test collections, our proposed NG models can provide acceptable retrieval performance in ad-hoc search, without the use of global statistics. Kai Hui 0001, Ben He 0001, Tiejian Luo, Bin Wang 0004 |
CIKM | 3 |
| 2011 | Modeling Learners and Contents in Academic-Oriented Recommendation FrameworkabstractLifelong learning is matter to knowledge society and Academic Recommendation is necessary to feed learners with the relevant and personalized contents. E-commerce recommendation system has made great successful in book suggestion, like Amazon. But these techniques are still not adapted to academic domain. Our study has found 4 factors, including learner's academic intention, social network, learning style and cognitive ability, which impact the effectiveness of AR system. This paper proposes a framework and model to build AR system. A working system based on the novel model has been constructed. This system which has explored 1099793 web page, 34737 videos, 910 experts, 13416 courses, 47390 publications, providing search, profiling, and suggestions functionality for 100k users. Tiejian Luo, Fuxing Cheng |
DASC | 2 |
| 2009 | Collaborative filtering with fine-grained trust metricabstractSimilarity-based collaborative filtering systems are vulnerable to the data sparsity, cold-start, and robustness problems. Computational trust models are promising alternative solutions to alleviate these problems by replacing similarity metric with trust metric. However, they often have some shortages that rely on users' explicit trust statements. A fine-grained model computing trust from user ratings is more reasonable and gets more nonintrusive for average users. We propose a novel trust-based recommendation model for this purpose. Experiments on a large real dataset show that the proposed model has better performance in terms of MAE, coverage, and F-metric than the conventional collaborative filtering model. Tiejian Luo, Wei Liu 0028, Yanxiang Xu |
CIDM | 2 |
| 2008 | Utilizing grid to build cyberinfrastructure for biosafety laboratoriesabstractThis paper describes work currently underway aiming to develop a distributed collaborative infrastructure to support the national science, especially the biosafety research in China. The irresistible trend of e-Science indicates that in the future more and more scientific researches will be carried out through close cooperation among collaborators. Following this trend, we need an infrastructure to satisfy the needs among institutes located across China. In this paper, we analyze the requirements for creating and maintaining a coherent Grid-enabled application management framework in a distributed environment from domain experts' and administrators' perspectives. Under such framework, we further develop computational and instrumental tools with Grid technology. As Grid is evolving over recent years towards the use of service-oriented architecture, we propose a solution with designing and implementing featured services based on the real demand of the biosafety research. This scalable architecture allows it grow into a bigger and comprehensive system with large amount of resources to better support scientific research and education. Tiejian Luo, Wei Liu 0028, Jinliang Song, Yuli Jin, Cheng Du |
CSCWD | 1 |
| 2002 | Scalable multimedia content delivery on InternetabstractWe propose a new multimedia data model, called the scalable data model (SDM), to describe how scalable pervasive Internet access with efficient data reuse can be achieved. This new data model not only provides a theoretical foundation for maximum possible reuse of transcoded data for scalable pervasive Internet access, but its supporting architecture also ensures the feasibility and practicability because of its matching to the emerging trends in HTTP and multimedia data format. Chihung Chi, Tiejian Luo |
ICME (1) | 3 |
| 2002 | An Improved Usage-Based Ranking
Chen Ding 0004, Chihung Chi, Tiejian Luo |
WAIM | 3 |
| 2000 | Defending Against Null Calls Stream Attacks by Using a Double-Threshold Dynamic Filter
Haizhi Xu, Changwei Cui, Tiejian Luo, Zhanqiu Dong |
SEC | 4 |