Abstract
InfraRed Thermography (IRT), a non-invasive, non-contact, and non-destructive testing (NDT) technique, has become an established tool in the assessment of a building’s behavior and energy performance. However, the inherent low spatial resolution of thermal infrared (TIR) cameras has led recent work to fuse thermographic and geometric data to generate accurate 3D representations of buildings encapsulating temperature information. Whilst existing data fusion methods have relied on sensors in fixed relative orientation (RO), the co-registration of independent TIR and RGB blocks using ground control points (GCPs), or the reprojection of TIR images onto additional geometric or parametric models, approaches that directly match multi-modal images remain limited. In principle, if multi-modal tie points were available, it would be possible to directly align the RGB block with the TIR block; however, such matching is extremely challenging due to the substantial differences in radiometric properties. The main contribution of this paper is to demonstrate the applicability of off-the-shelf deep learning-based image matching algorithms, originally trained on mono-modal datasets, to multi-modal matching tasks for InfraRed Thermography 3D-Data Fusion (IRT-3DDF). We conduct a comparative evaluation of the principal algorithms developed in recent years, with particular emphasis on 3D accuracy and computational efficiency, under the hypothesis that, owing to the inherently local nature of the problem they address, these algorithms can generalize from a mono-modal training domain to a multi-modal application domain. The results are benchmarked against existing hand-crafted open-source multi-modal reference methods. Importantly, the proposed method is fully-automatic, obviating the need for sensor pre-calibration, manual co-registration, or associated positioning information. Results demonstrate that DL-based image matching, using pre-trained neural networks outside of their expected training domain, provides a viable approach for IRT-3DDF capable of co-registering blocks of multi-modal images across varying scales, settings, sensors, and subjects. Our results indicate accuracy in 3D is up to seven times better than multi-modal hand-crafted algorithms, while hand-crafted mono-modal methods fail to co-register images in their entirety.
IPC Classification
Keywords
€ 4.00