TY - GEN
T1 - A Cross-Domain Threat Screening and Localization Framework Using Vision Transformers and Self-supervised Learning
AU - Nasim, Ammara
AU - Akram, Muhammad Usman
AU - Khan, Asad Mansoor
AU - Khan, Muhammad Belal Afsar
AU - Hassan, Taimur
PY - 2024/9/23
Y1 - 2024/9/23
N2 - Due to the ever-changing global security landscape, various countries are constantly updating their safety protocols at designated airports and train stations. Consequently, there is an expanding inventory of prohibited items that are not permitted to be carried in luggage. This renders training of prior models on large datasets but lesser threat classes, ineffective. The scarcity of extensive, accurately labeled X-ray images with new threat items hinders the feasibility of training data-intensive, fully supervised deep models, primarily due to the high cost and time associated with manual labeling. In this paper, a self-supervised threat detection approach is discussed that is combined with vision transformer(VIT)-based pre-screening. The approach is adaptable to new classes and can easily detect new threat classes in the inference stage using very few annotated support images. The VIT based classification module showcases promising results with an accuracy of 98% and F1-score of 99%. The self-supervised threat detection module surpasses other self-supervised and weakly supervised frameworks with Mean Average Precision (mAP) of 0.89 and an Intersection over Union (IoU) value of 0.70 on GDXRAY dataset. The framework performs exceptionally well on SIXRAY dataset with an IOU score of 0.67 and mAP value of 0.865 despite being trained on GDXRAY dataset.
AB - Due to the ever-changing global security landscape, various countries are constantly updating their safety protocols at designated airports and train stations. Consequently, there is an expanding inventory of prohibited items that are not permitted to be carried in luggage. This renders training of prior models on large datasets but lesser threat classes, ineffective. The scarcity of extensive, accurately labeled X-ray images with new threat items hinders the feasibility of training data-intensive, fully supervised deep models, primarily due to the high cost and time associated with manual labeling. In this paper, a self-supervised threat detection approach is discussed that is combined with vision transformer(VIT)-based pre-screening. The approach is adaptable to new classes and can easily detect new threat classes in the inference stage using very few annotated support images. The VIT based classification module showcases promising results with an accuracy of 98% and F1-score of 99%. The self-supervised threat detection module surpasses other self-supervised and weakly supervised frameworks with Mean Average Precision (mAP) of 0.89 and an Intersection over Union (IoU) value of 0.70 on GDXRAY dataset. The framework performs exceptionally well on SIXRAY dataset with an IOU score of 0.67 and mAP value of 0.865 despite being trained on GDXRAY dataset.
KW - Cross-domain
KW - Self-supervised learning
KW - threat localization
KW - Vision transformer
UR - https://www.scopus.com/pages/publications/85206481419
UR - https://ieeexplore.ieee.org/document/10677838
U2 - 10.1109/ICPRS62101.2024.10677838
DO - 10.1109/ICPRS62101.2024.10677838
M3 - Conference publication
AN - SCOPUS:85206481419
SN - 9798350375664
T3 - 2024 14th International Conference on Pattern Recognition Systems, ICPRS 2024
BT - 2024 14th International Conference on Pattern Recognition Systems, ICPRS 2024
PB - IEEE
T2 - 14th International Conference on Pattern Recognition Systems, ICPRS 2024
Y2 - 15 July 2024 through 18 July 2024
ER -