Machine learning for translation to structured computer readable representation

Inventors

Shankar, NatarajanGraham-Lengrand, StephaneElenius, DanielYeh, Chih-Hung

Assignees

SRI International Inc

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-12518089-B2

Patent

Publication Date

2026-01-06

Expiration Date


Abstract

In general, the disclosure describes techniques for machine learning for translation to structured computer readable representation. An example method to generate a training set for a natural language translation model includes receiving, by a computing system, a grammar comprising rules, one or more of the rules being associated with random biases; generating, by the computing system, at least one of random trees or random graphs based on the random biases in the grammar; for each of the random trees or random graphs, by the computing system, generating a natural language sample; and generating, by the computing system, the training set with the random trees or random graphs and the corresponding natural language samples.

Core Innovation

The invention relates to machine learning-based translation of natural-language standards text into domain-specific symbolic structured representations using a grammar and probabilistic generation of random trees and/or random graphs. A grammar comprises rules, and one or more rules are associated with respective random biases, where each random bias specifies a probability that a corresponding node for a corresponding rule is included when generating a random tree or a random graph.

Based on the random biases, at least one random tree or random graph is generated, and for each generated random tree or random graph, a corresponding natural language sample is generated. A training set is formed comprising training pairs, each training pair including one of the generated random trees or random graphs and the corresponding natural language sample, and the machine learning model is trained using the training set.

When applying the trained machine learning model, the system receives a natural language statement and applies the model to generate a translation tree. The translation tree maps the natural language statement to a symbolic language output, where the output is a structured statement defined by the structures produced from the grammar and supports translation into symbolic language outputs.

Claims Coverage

The document provides three independent claims, centered on generating training pairs from grammar-defined random trees or random graphs driven by random biases, and using a trained machine learning model to output a translation tree mapping a natural language statement to symbolic language output. The inventive features are the grammar with random biases for node inclusion, probabilistic random tree or random graph generation, generation of corresponding natural language samples, formation of training pairs, and applying the trained model to produce a translation tree for symbolic language output.

Grammar with random biases for node inclusion

Receiving a grammar comprising rules, wherein one or more of the rules are associated with respective random biases, and each random bias specifies a probability that a corresponding node for a corresponding rule is included when generating at least one random tree or random graph.

Randomly generating random trees or random graphs from biased grammar

Generating, based on the random biases in the grammar, at least one random tree or random graph.

Generating corresponding natural language samples and training pairs

For each random tree or random graph, generating a corresponding natural language sample, and generating a training set comprising training pairs, each including one of the random trees or random graphs and the corresponding natural language sample, and training the machine learning model using the training set.

Applying trained model to generate translation tree for symbolic output

Receiving the natural language statement, applying the trained machine learning model to generate a translation tree, and outputting the translation tree that maps the natural language statement to a symbolic language output.

Non-transitory computer readable medium for symbolic-language translation via translation tree

A non-transitory computer readable storage medium for causing the computing system to receive the natural language statement, apply the trained machine learning model to generate a translation tree, and output the translation tree that maps the natural language statement to a symbolic language output.

The core claim coverage is for training and using a machine learning model where training pairs are generated from grammar-defined random trees or random graphs produced using random biases that specify probabilities for node inclusion, and where applying the trained model outputs a translation tree that maps a natural language statement to symbolic language output.

Stated Advantages

Supports domain adaptation with few examples.

Enables downstream automated verification/validation.

Documented Applications

Translation from natural-language standards text into domain-specific symbolic structured representations.

Downstream automated verification/validation for the generated symbolic output.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.