Architecture for semantic search over encrypted data in the cloud

Inventors

Woodworth, Jason • Salehi, Mohsen Amini

Assignees

University of Louisiana at Lafayette

Interested in licensing this patent?

MTEC can help explore whether this patent might be available for licensing for your application.

Publication Number

US-11550833-B2

Patent

Publication Date

2023-01-10

Expiration Date


Abstract

An architecture for semantic search over encrypted data that improves upon existing encrypted data search techniques by providing a solution that is space-efficient on both the cloud and client sides, considers the semantic meaning of the user's query, and returns a list of documents accurately ranked by their similarity to the query. Different search schemes are presented based on S3C architecture (namely, FKSS, SKSS, and KSWF) that are fine-tuned for different types of datasets. The system requires only a single plaintext query to be entered and is easily portable to thin-clients, making it simple and quick for users to use. The system is also shown to be secure and resistant to attacks.

Core Innovation

The invention provides an architecture for semantically searching over encrypted data in the cloud using a client trusted boundary and an untrusted cloud processing server and cloud storage. Uploaded documents are parsed into indexable information and encrypted documents, where the indexable information comprises at least one keyword and the frequency at which the keyword appears in the document. The indexable information is transformed into hashed terms written to a temporary key file and added into a hashed inverted index associated with the encrypted data.

The invention performs semantic query expansion by expanding an inputted plaintext query. The expansion includes splitting the plaintext query into a split query set, inserting semantic data into the split query set to create a complete query set, and weighting the complete query set so that the plaintext query, the split query set, and the semantic data are assigned different ranking priorities. The semantic data is pulled from one or more advanced ontological networks, including key phrase extraction from sources such as Wikipedia and related terms from the ontological networks.

The weighted query set is hashed to create one or more trapdoors that are transmitted to a cloud processing server. The cloud processing server checks each trapdoor member against the hashed index and an index of the encrypted data and ranks results by using a modified ranking process based on the weighting, including a modified BM25 using inverse document frequency. The architecture includes searchable semantic schemes that vary how keywords are selected and indexed, and optionally performs topic-based clustering to search only related clusters.

Claims Coverage

Independent claim coverage is provided for a semantic searching method over encrypted data and for a corresponding computer architecture. The inventive features are centered on semantic query expansion using advanced ontological networks and on a weighted, hashed trapdoor-based interaction with a cloud server that ranks results against a hashed index using inverse document frequency.

Semantically searching over encrypted data with hashed keyword frequency indexing and weighted semantic query trapdoors

parsing said at least one uploaded document into indexable information and an encrypted document, transforming said indexable information into a hashed term, and writing said hashed term to a temporary key file, wherein said indexable information comprises at least one keyword from said at least one uploaded document and the frequency at which said at least one keyword appears in said at least one uploaded document; expanding an inputted plaintext query by splitting said plaintext query to create a split query set, inserting semantic data into said split query set to create a complete query set, and weighting said complete query set to create a weighted query set, wherein said weighting is based as follows: highest ranking to said plaint text query, next highest ranking to said split query set, and third highest ranking to said semantic data, wherein said semantic data is pulled from one or more advanced ontological networks; hashing said weighted query set members to create a trapdoor, and transmitting said trapdoor to a cloud processing server, wherein said cloud processing server checks each of said trapdoor members against said hashed index and an index of said encrypted data and ranks said trapdoor members, creating a ranked list, wherein said ranked list is generated based on said weighting.

Computer architecture for semantic searching over encrypted data using client query expansion, hashed indexing, and ranked trapdoor comparison

a file parser that parses said at least one uploaded document into indexable information, then transforms said indexable information into a hashed term, and then writes said hashed term to a temporary key file, wherein said indexable information comprises at least one keyword from said at least one uploaded document and the frequency at which said at least one keyword appears in said at least one uploaded document; an encryptor that encrypts said at least one document to create an encrypted document; a searching function that performs query modification on said query by splitting said query to create a split query consisting of said query and any individual parts of said query, together split query members, performs semantic expansion of said split query comprising performing key phrase extraction to retrieve related terms to at least one said split query members from at least one advanced ontological network and inserting said related terms into said split query to create an expanded split query; a weighting function that receives said expanded split query and adds the following weighting scheme to each member of said expanded split query so that a complete modified query set is created; and a cloud processing server that receives said temporary key file and said encrypted document and then moves said encrypted document into a cloud storage and adds said hashed terms from said temporary key file to a hashed index and performs clustering on said hashed index, and wherein said modified query set is hashed to create a trapdoor and then compares said trapdoor to said hashed index to compile a list of said at least one uploaded documents that could be considered related to said query and uses a ranking engine to rank said list of related documents using the inverse document frequency for each of said documents in said list of related documents.

Across the independent claims, the inventive coverage combines (1) hashed keyword frequency indexing of encrypted documents and (2) semantic query expansion from advanced ontological networks followed by weighting of plaintext, split query set, and semantic data. The resulting weighted query information is hashed into trapdoors that are transmitted to an untrusted cloud processing server for comparison against the hashed index and encrypted-data index to produce a ranked list using inverse document frequency consistent with the weighting.

Stated Advantages

SKSS is space efficient and scalable.

KSWF achieves higher accuracy with small overhead and a security tradeoff.

FKSS is suitable for small documents where full text matters.

Documented Applications

Semantic searching over locally encrypted cloud-stored documents using a client trusted boundary with an untrusted cloud processing server and cloud storage.

Optional topic-based clustering to search only related clusters.

JOIN OUR MAILING LIST

Stay Connected with MTEC

Keep up with active and upcoming solicitations, MTEC news and other valuable information.