Language element vision augmentation methods and devices
Inventors
Jones, Frank • Bacque, James Benson
Assignees
Interested in licensing this patent?
MTEC can help explore whether this patent might be available for licensing for your application.
Abstract
Near-to-eye displays support a range of applications from helping users with low vision through augmenting a real world view to displaying virtual environments. The images displayed may contain text to be read by the user. It would be beneficial to provide users with text enhancements to improve its readability and legibility, as measured through improved reading speed and/or comprehension. Such enhancements can provide benefits to both visually impaired and non-visually impaired users where legibility may be reduced by external factors as well as by visual dysfunction(s) of the user. Methodologies and system enhancements that augment text to be viewed by an individual, whatever the source of the image, are provided in order to aid the individual in poor viewing conditions and/or to overcome physiological or psychological visual defects affecting the individual or to simply improve the quality of the reading experience for the user.
Core Innovation
A near-to-eye (NR2I) system provides improved legibility of text within an image to a user by acquiring an original image and processing it to establish a region of a plurality of regions. Each region has a probability of character based content exceeding a threshold probability, and character based content is extracted from the region or regions whose probability exceeds the threshold probability.
The system presents the extracted character based content to the user upon a device, using rendering that places modified text in user-visible areas. Presentation occurs in predetermined portions of a field of view (FOV) and/or predetermined portions of a region of interest (ROI) other than the region where the extracted content was originally sourced, and can include modification of font, font size, foreground/background color scheme, and font effect, with support for translation to a preferred language of the user.
Processing adapts to user context and vision characteristics using gaze direction and/or head orientation, and can use preferred retinal locus (PRL) or ROI mapping. Depth mapping can be used for filtering and placing content, selectively transmissive display architectures can be used including microdisplay, free-form prism, and partial-mirroring shutter elements, and user feedback can be used when a predetermined presentation characteristic crosses a threshold between ease of comprehension and difficulty of comprehension.
Claims Coverage
The independent claims covered in the provided material relate to a near-to-eye (NR2I) system that improves legibility by selecting regions with a character-content probability above a threshold, extracting character-based content, and presenting the extracted content to a user on a device. The independent features also include modality options, rendering placement/styling/translation, and optional user-feedback thresholding.
Select regions by character-content probability threshold and extract character-based content
Processing the original image to establish a region of a plurality of regions, each region having a probability of character based content exceeding a threshold probability; processing the region of the plurality of regions to extract character-based content.
Present extracted character-based content upon a device
Presenting the extracted character-based content to the user upon a device.
Define regions using metadata or mark-up language tag data
Processing the original image to extract data associated with the original image, where the data defines the region and is at least one of meta-data and a mark-up language tag.
Render extracted character-based content in predetermined portions of FOV/ROI
Displaying the extracted character-based content in predetermined portions of the user’s field of view (FOV) and/or predetermined portions of a region of interest (ROI) other than the region where the extracted content was originally sourced, and optionally highlighting where in the original image the extracted content was extracted from.
Audible output of translated extracted character-based content via a loudspeaker
With a loudspeaker as the device, applying OCR to the region to generate recognized character-based content, translating it to the user’s preferred language, and providing the translated character-based content as an audible signal.
OCR-to-translation-to-styling rendering of extracted character-based content on a display
Applying OCR to a region to generate recognized character based content, translating recognized text to the user’s preferred language, selecting font, font size, foreground/background color scheme and font effect, and rendering the translated text on the display.
User-feedback threshold controlling a presentation limiting value
Varying a predetermined presentation characteristic for extracted character based content, receiving user feedback when the variation crosses a threshold between ease of comprehension and difficulty of comprehension, storing the feedback value, and using it as a limiting value in subsequently presenting the extracted character based content.
Overall, the claim coverage centers on probabilistic region selection for character-based content, extraction of character-based content from selected regions, and device presentation. Dependent inventive features further specify how regions are defined, how extracted content is placed in FOV/ROI and optionally highlighted, modality for audible output, OCR/translation/styling rendering on a display, and use of a user-feedback threshold to set a limiting value for subsequent presentation.
Stated Advantages
Improved legibility of text within an image for a user.
Support for user legibility through presentation of extracted character based content within predetermined portions of the user’s field of view (FOV) and/or region of interest (ROI).
Improved usability of text presentation using image processing techniques such as contrast/brightness adjustment, edge enhancement, color remapping, and binarization/greyscale conversion.
Adaptation to vision characteristics using preferred retinal locus (PRL) or ROI mapping.
User-comprehension-controlled presentation using a user-feedback threshold between ease of comprehension and difficulty of comprehension.
Documented Applications
Near-to-eye (NR2I) text enhancement for improving legibility/readability of text within an image presented to a user.
Rendering extracted text for improved readability by placing it in predetermined portions of a user’s field of view (FOV) and/or region of interest (ROI), and optionally highlighting the source region in the original image.
Providing extracted character-based content as an audible signal using a loudspeaker.
Displaying extracted character-based content with OCR and optional translation to a preferred language and with font/color styling effects.
Vision-adaptive text presentation using gaze direction/head orientation and preferred retinal locus (PRL) or ROI mapping.
Use of selectively transmissive display architectures (including microdisplay, free-form prism, and partial-mirroring shutter element) for the described text presentation.
Interested in licensing this patent?