Robust roadway crack segmentation using encoder-decoder networks with range images
Inventors
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
In an implementation, a method for pixel level roadway crack segmentation is provided. The method includes: receiving a plurality of roadway range images; generating a plurality of image patches from the plurality of roadway range images; generating a crack map for the plurality of image patches by a DCNN; and generating a crack map for the plurality of roadway range images based on the generated crack map for the plurality of image patches.
Core Innovation
The invention provides pixel level roadway crack segmentation using a deep convolutional neural network (DCNN) that receives 3D roadway range image patches as inputs and produces pixel-level crack maps as outputs. The 3D roadway range images are laser-scanned perpendicular to a roadway surface by an imaging system. The approach generates 3D image patches from the 3D roadway range images without downsampling by cropping each 3D roadway range image into multiple 3D image patches.
A ground truth label for each image patch is generated, including receiving a ground truth label for the image patch from a reviewer. The DCNN is an encoder-decoder network with an encoder module and a decoder module that receive the 3D roadway range image patches. The encoder module contains a plurality of convolutional layers in a first order, and the decoder module contains a corresponding plurality of transposed convolutional layers in a second order, with residual connections connecting corresponding layers across the encoder and decoder.
The architecture includes residual connections between the plurality of 2D convolutional layers and the 2D transposed convolutional layers, where some of the transposed convolutional layers receive one or more features from one or more preceding convolutional layers through the residual connections. The invention also constrains network layer scaling, where the number of convolutional kernels in each convolutional layer equals two times the number of convolutional kernels of a directly preceding convolutional layer in the first order, and where the number of transposed convolutional kernels of each transposed convolutional layer is two times a number of transposed convolutional kernels in a directly following transposed convolutional layer in the second order.
Claims Coverage
The independent claims are clm-00001, clm-00004, and clm-00006, and they cover pixel-level roadway crack segmentation using laser-scanned 3D roadway range images, 3D image patches without downsampling, an encoder-decoder DCNN with residual connections, and output pixel-level crack maps. The claims include multiple inventive features: receiving laser-scanned perpendicular range images, patch-based processing, patch ground truth labeling, a specific encoder-decoder DCNN architecture with residual connections, and kernel-count scaling constraints, with additional quantitative processing constraints in the later claim set.
Laser-scanned perpendicular 3D roadway range images
Receiving a first plurality of 3D roadway range images, wherein each 3D roadway range image is laser-scanned perpendicular to a roadway surface.
Patch generation without downsampling by cropping 3D range images
Generating a first plurality of 3D image patches from the first plurality of 3D roadway range images without downsampling, wherein each 3D roadway range image is cropped into multiple 3D image patches.
Patch ground truth labeling for training
For each image patch of the first plurality of 3D image patches, generating a ground truth label for the image patch.
Encoder-decoder DCNN with residual connections producing pixel-level crack maps
Training a deep convolutional neural network (DCNN) using the labeled 3D image patches by the computing device, wherein the DCNN comprises a plurality of 2D convolutional layers and 2D transposed convolutional layers with residual connections between the 2D convolutional layers and the 2D transposed convolutional layers, and wherein the DCNN produces pixel-level crack maps as outputs, wherein the DCNN is an encoder-decoder network that receives the 3D roadway range image patches as inputs and comprises a decoder module and an encoder module.
Kernel-count scaling across encoder and decoder layers
A number of convolutional kernels of each convolutional layer in the encoder module equals two times a number of convolutional kernels of a directly preceding convolutional layer in the first order, and wherein a number of transposed convolutional kernels of each transposed convolutional layer in the decoder module is two times a number of transposed convolutional kernels in a directly following transposed convolutional layer in the second order.
Patch-to-full-image crack map generation for a second plurality of range images
Receiving a second plurality of 3D roadway range images, generating a second plurality of 3D image patches from them, generating a crack map for the second plurality of 3D roadway range images based on a generated crack map for the second plurality of 3D image patches.
Batch normalization and leaky rectified linear unit in encoder layers
Each encoder 2D convolutional layer applies Batch Normalization followed by a Leaky Rectified Linear Unit.
Batch normalization and leaky rectified linear unit in decoder layers
Each decoder 2D transposed convolutional layer applies Batch Normalization followed by a Leaky Rectified Linear Unit.
Across clm-00001, clm-00004, and clm-00006, the claims center on laser-scanned perpendicular 3D roadway range images, generating 3D image patches without downsampling, using patch ground truth labels, and training and applying an encoder-decoder DCNN with residual connections to produce pixel-level crack maps. The architecture further includes explicit residual layer connections and kernel-count scaling constraints between corresponding encoder and decoder layers, with optional encoder and decoder per-layer processing constraints including Batch Normalization and Leaky Rectified Linear Unit.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?