Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

LLMs, Embeddings and Indexing Pipelines to Enable Natural Language Searching on Upstream Datasets

View through CrossRef
Abstract Large Language Models (LLMs) are attracting an enormous amount of interest at the moment in many domains. Their general nature, and ability to "understand" natural language, has already stimulated multiple areas of research at our company. Here we successfully demonstrate a Natural Language querying system, which is able to search a large repository of unstructured exploration data. The system supports follow up querying on the returned results, plus automatic summarization of content. The system is integrated into our novel end-to-end data-mining platform, which continuously mines our unstructured exploration data for new changes and indexes the results. Important in our method are the enrichment processes that occur prior to use of the LLM. Our approach avoids usual "chunking" techniques, which in our experience results in inferior results, especially in the multiple domain areas of Exploration. By integrating our novel ontology-model AI in the enrichment of the initial Index, we drastically boost the performance of search resulting from the LLM steps. In order to perform the search, key parts of our unstructured data, plus the query itself, need to be transformed into a vector form. This is performed using the embedding feature of the LLM. For this work, we had around 500,000 embeddings to calculate. To improve performance these were indexed in a leading Analytics Engine as a vector object, allowing fast search via cosine or Euclidian similarity. A custom dashboard was made to allow fresh searches of the vector datastore to be returned for further analysis. Our current search time across 500,000 embeddings is under 20 milli-seconds. Our custom dashboard returns the top matches for further interrogation and analysis. This includes follow-up Natural Language question support on the returned matches for summarization tasks and other customised querying. Since our exploration-specific, ontology model is able to tag each piece of data with over 40 exploration-specific labels, we are able to cross-examine the LLM returned results with the tags. Agreement on a range of queries - ranging from targeted, highly specific questions to general, open-ended queries - was surprisingly good. Natural Language based querying of our unstructured data is opening a whole new approach to data discovery in our company. Tailoring it to the exploration domain has required specific domain expertise and a novel ontology-model be used to ensure relevant prompts and query results. Obtaining search results quickly has also required expertise and fine-tuning. Future directions include ingesting more data, scaling the support infrastructure and further capability enhancement.
Title: LLMs, Embeddings and Indexing Pipelines to Enable Natural Language Searching on Upstream Datasets
Description:
Abstract Large Language Models (LLMs) are attracting an enormous amount of interest at the moment in many domains.
Their general nature, and ability to "understand" natural language, has already stimulated multiple areas of research at our company.
Here we successfully demonstrate a Natural Language querying system, which is able to search a large repository of unstructured exploration data.
The system supports follow up querying on the returned results, plus automatic summarization of content.
The system is integrated into our novel end-to-end data-mining platform, which continuously mines our unstructured exploration data for new changes and indexes the results.
Important in our method are the enrichment processes that occur prior to use of the LLM.
Our approach avoids usual "chunking" techniques, which in our experience results in inferior results, especially in the multiple domain areas of Exploration.
By integrating our novel ontology-model AI in the enrichment of the initial Index, we drastically boost the performance of search resulting from the LLM steps.
In order to perform the search, key parts of our unstructured data, plus the query itself, need to be transformed into a vector form.
This is performed using the embedding feature of the LLM.
For this work, we had around 500,000 embeddings to calculate.
To improve performance these were indexed in a leading Analytics Engine as a vector object, allowing fast search via cosine or Euclidian similarity.
A custom dashboard was made to allow fresh searches of the vector datastore to be returned for further analysis.
Our current search time across 500,000 embeddings is under 20 milli-seconds.
Our custom dashboard returns the top matches for further interrogation and analysis.
This includes follow-up Natural Language question support on the returned matches for summarization tasks and other customised querying.
Since our exploration-specific, ontology model is able to tag each piece of data with over 40 exploration-specific labels, we are able to cross-examine the LLM returned results with the tags.
Agreement on a range of queries - ranging from targeted, highly specific questions to general, open-ended queries - was surprisingly good.
Natural Language based querying of our unstructured data is opening a whole new approach to data discovery in our company.
Tailoring it to the exploration domain has required specific domain expertise and a novel ontology-model be used to ensure relevant prompts and query results.
Obtaining search results quickly has also required expertise and fine-tuning.
Future directions include ingesting more data, scaling the support infrastructure and further capability enhancement.

Related Results

Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
Hubungan Perilaku Pola Makan dengan Kejadian Anak Obesitas
<p><em><span style="font-size: 11.0pt; font-family: 'Times New Roman',serif; mso-fareast-font-family: 'Times New Roman'; mso-ansi-language: EN-US; mso-fareast-langua...
Exploiting word embeddings for modeling bilexical relations
Exploiting word embeddings for modeling bilexical relations
There has been an exponential surge of text data in the recent years. As a consequence, unsupervised methods that make use of this data have been steadily growing in the field of n...
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Exploring Large Language Models Integration in the Histopathologic Diagnosis of Skin Diseases: A Comparative Study
Abstract Introduction The exact manner in which large language models (LLMs) will be integrated into pathology is not yet fully comprehended. This study examines the accuracy, bene...
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
Učinak poučavanja razrednomu jeziku u izobrazbi nastavnika njemačkoga
The actual use of classroom language is principally limited to the classroom environment. As far as foreign language learning is concerned, the classroom often turns out to be the ...
Pelatihan Searching dan Drafting Paten Di Perguruan Tinggi Muhammadiyah Mataram
Pelatihan Searching dan Drafting Paten Di Perguruan Tinggi Muhammadiyah Mataram
Dalam pelaksanaan pengabdian ini kami memberikan pemahaman dan pelatihan akan searching dan drafting paten sesuai kebutuhan dari peserta. Untuk searching kami berikan dengan menunj...
A Review on Indexing Techniques and its application in Multilingual Information Retrieval System
A Review on Indexing Techniques and its application in Multilingual Information Retrieval System
To implement the indexing in multilingual dataset, the indexing process must know. This paper gives the brief about indexing and presents role of indexing, logical view of indexing...
Perspectives and Experiences With Large Language Models in Health Care: Survey Study (Preprint)
Perspectives and Experiences With Large Language Models in Health Care: Survey Study (Preprint)
BACKGROUND Large language models (LLMs) are transforming how data is used, including within the health care sector. However, frameworks including the Unifie...
Perspectives and Experiences With Large Language Models in Health Care: Survey Study
Perspectives and Experiences With Large Language Models in Health Care: Survey Study
Background Large language models (LLMs) are transforming how data is used, including within the health care sector. However, frameworks including the Unified Th...

Back to Top