Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
A voice user interface (VUI) and methods for operating the VUI are disclosed. In some embodiments, the VUI configured to receive and process linguistic and non-linguistic inputs. For example, the VUI receives an audio signal, and the VUI determines whether the audio input comprises a linguistic and/or a non-linguistic input. In accordance with a determination that the audio signal comprises a non-linguistic input, the VUI causes a system to perform an action associated with the non-linguistic input.
Core Innovation
The invention relates to a voice user interface that receives an audio signal and distinguishes linguistic input from non-linguistic input. Non-linguistic input is characterized using paralinguistic/prosodic information such as tone, cadence, and emotion, and the system determines whether the audio signal comprises non-linguistic input before triggering an action associated with the non-linguistic input.
The non-linguistic determination is performed using time- and frequency-domain feature analysis with a feature database and/or using a convolutional neural network. The resulting non-linguistic information can drive intent-based commands such as taking a picture/video, texting, and inserting an emoji, and the system modifies actions derived from linguistic input based on the detected non-linguistic input.
The voice user interface is implemented in a wearable/mixed reality head device context with a plurality of microphones. The system determines whether the audio signal is associated with the user by associating the audio source with user location, using multi-microphone spatial processing that determines respective distances between an audio source and each microphone and a location of the source; in accordance with the audio being associated with the user, it proceeds to determine non-linguistic input, and if not associated, it forgoes determining whether the audio comprises non-linguistic input.
Claims Coverage
The independent claims explicitly cover three alternative implementations: a system, a method, and a non-transitory computer-readable medium. Across these independent claims, the coverage centers on three inventive features: user association via multi-microphone source localization using distances to microphones, conditional non-linguistic input detection only when the audio is associated with the user, and performing an action associated with the non-linguistic input.
User association using microphone distances and source location
The system determines whether the audio signal is associated with a user of the wearable head device by determining respective distances between a source of the audio signal and each of the plurality of microphones, and determining, based on the respective distances, a location of the source of the audio signal.
Conditional non-linguistic input determination
In accordance with a determination that the audio signal is associated with the user of the wearable head device, the system determines whether the audio signal comprises a non-linguistic input; in accordance with a determination that the audio signal is not associated with the user of the wearable head device, the system forgoes determining whether the audio signal comprises the non-linguistic input.
Action associated with non-linguistic input
In accordance with a determination that the audio signal comprises the non-linguistic input, the system performs an action associated with the non-linguistic input.
The independent claim set is directed to a wearable head device voice user interface that first associates an audio signal to the user using distance-based source location from a plurality of microphones, and then conditionally detects non-linguistic input and performs an action associated with that non-linguistic input.
Stated Advantages
Not explicitly described in patent.
Documented Applications
Not explicitly described in patent.
Interested in licensing this patent?