A domain specific virtual assistant using paraphrase generation for data augmentation and Ssentence transformers on limited data

Clicks: 2
ID: 286016
2023
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This article has not been analysed, so there is no overall score — reader engagement is measured and shown alongside.
AI Quality Assessment
Not analyzed
Readership in this journal
Steady

Ranked #1,245 of 3,757 articles by views in Malay Journal

Most read Least read

Bar heights use a square-root scale. Only the 120 most-read articles are drawn; the journal has 3,757 in total.

Mint this article as an NFT
Not yet minted

Create a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.

5 SUSD one-off · no wallet required
Abstract
The use of conversational agents can be extremely beneficial in many areas such as government offices, schools, banks, malls, etc. where people often make inquiries and responses from personnel can take some time. Many of these areas, however, have inquiries that involve domain-specific vocabulary and most likely do not have a large amount of data or computational resources to properly train a complex natural language processing (NLP) model. This paper proposes a method for creating a domain-specific virtual assistant using Generative Pre-Trained Transformer-3 (GPT-3) to generate paraphrases on a relatively small dataset, and a Sentence Transformer (SBERT) model with a distilled version of BERT (DistilBERT) base, pretrained on the Quora Question Pairs dataset, and fine-tuned on the augmented dataset. This method of creating a model is evaluated on the MS MARCO, SemEval, and PubMed datasets using mean average precision (MAP), precision at k (P@k), normalized discounted cumulative gain (NDCG), and mean reciprocal rank (MRR) as performance metrics. The method was also demonstrated using a small dataset of 188 frequently asked questions from the De La Salle University website that also includes domain-specific vocabulary. The implementation of the fine-tuned model was demonstrated on a simple webpage and the results were found to be satisfactory.
Reference Key
persistent_1760657324_68f17facd04e0 Use this key to autocite in the manuscript while using SciMatic Manuscript Manager or Thesis Manager
Authors Roque, Matthew Theodore C.
Journal Malay Journal
Year 2023
DOI
DOI not found
URL
Keywords Keywords not found

Citations

No citations found. To add a citation, contact the admin at info@scimatic.org

No comments yet. Be the first to comment on this article.