Deep learning-based crack segmentation through heterogeneous image fusion

Inventors

Song, WeiZhou, Shanglian

Assignees

University of Alabama at Birmingham UAB

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12146838-B2

Patent

Publication Date

2024-11-19

Expiration Date


Abstract

In an embodiment, a method for detecting cracks in road segments is provided. The method includes: receiving raw range data for a first image by a computing device from an imaging system, wherein the first image comprises a plurality of pixels; receiving raw intensity data for the first image by the computing device from an imaging system; fusing the raw range data and raw intensity data to generate fused data for the first image by the computing device; extracting a set of features from the fused data for the first image by the computing device; providing the set of features to a trained neural network by the computing device; and generating a label for each pixel of the plurality of pixels by the trained neural network, wherein a received label for a pixel indicates whether or not the pixel is associated with a crack.

Core Innovation

The invention relates to detecting cracks in road segments using a deep convolutional neural network that consumes raw 3D range data and raw 2D intensity data having pixel-to-pixel location correspondence. The pixel-to-pixel location correspondence generates spatial co-location features that uniquely identify a portion of a road segment for each pixel, and fused data are generated directly from the raw 3D range data and the raw 2D intensity data without preprocessing or filtering.

The deep convolutional neural network is an encoder-decoder network with an encoder module and a decoder module. The encoder module comprises two branches, wherein one branch receives the raw 2D intensity data to produce intensity-based features and the second branch receives the raw 3D range data to produce range-based features, and the output features from both branches are integrated through an addition operation.

Crack detection is performed by receiving labels for each pixel indicating whether the pixel is associated with a crack. The decoder module has a same number of transposed convolutional layers as either encoder branch, and the network outputs pixel-level crack/non-crack labels using the fused data containing the spatial co-location features.

Claims Coverage

The partial content provided contains two independent claims. Across these claims, the inventive features center on multi-modal raw 3D range and raw 2D intensity fusion via spatial co-location features, followed by an encoder-decoder deep convolutional neural network with two encoder branches and feature integration through an addition operation.

Pixel-to-pixel location correspondence to form spatial co-location features

Raw 2D intensity data have pixel-to-pixel location correspondence with raw 3D range data, wherein the pixel-to-pixel location correspondence generates spatial co-location features that each uniquely identify a portion of a road segment that each pixel of the plurality of pixels correspond to.

Direct fused data integration without preprocessing or filtering

Fused data are generated containing the spatial co-location features by integrating the raw 3D range data and raw 2D intensity data to generate fused data for the first image, wherein the fused data is generated directly from the raw 3D range data and the raw 2D intensity data without any preprocessing or filtering of the raw 3D range data or the raw 2D intensity data.

Encoder-decoder deep convolutional neural network with two branches and addition-based feature integration

The deep convolutional neural network is an encoder-decoder network that receives the raw 2D intensity data and raw 3D range data with the pixel-to-pixel location correspondence, wherein the encoder module comprises two branches, wherein the first branch receives the raw 2D intensity data and produces intensity-based features, wherein the second branch receives the raw 3D range data and produces range-based features, and further wherein the output features from both branches are integrated through an addition operation.

Per-pixel crack label output from deep convolutional neural network

Receiving a label for each pixel of the plurality of pixels from the deep convolutional neural network, wherein a received label for a pixel indicates whether or not the pixel is associated with a crack.

Both independent claims cover road crack detection by fusing pixel-aligned raw 3D range data and raw 2D intensity data into fused data containing spatial co-location features, and then performing per-pixel crack labeling using an encoder-decoder deep convolutional neural network with two encoder branches whose outputs are integrated via an addition operation.

Stated Advantages

Documented Applications

Detecting cracks in road segments using a vehicle-mounted laser imaging system that provides raw 3D range data and raw 2D intensity data with pixel-to-pixel location correspondence.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.