Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Bi-Contextual Retrieval Augmented Generation (RAG) for Automatic Descriptive Answer Grading

View through CrossRef
Automatic Short Answer Grading (ASAG) is a well-known research task in the field of natural language processing (NLP). Its major purpose is to automatically grade descriptive answers of the students by keeping automatic grading consistent with the evaluation of human graders. Recent developments in Large Language Models (LLMs) have demonstrated a greatly enhanced performance in automated grading; however, the generalizability of the models and accuracy is still quite low because of the absence of dataset-specific grounding. We present EDURAG, a Retrieval-Augmented Generation (RAG) based model to improve contextualization of the LLM-based grading with exemplar-based grading and extra knowledge as generated by QFKE (Question Focused Knowledge Extraction) module. The proposed QFKE module provides extra layer of contextuality for the EDURAG. In contrast to conventional supervised methods, EDURAG does not need model fine-tuning. The suggested framework is tested against the ASAG2024 benchmark that consolidates seven short-answer grading datasets across various domains, educational levels, and grading scales. The benchmark protocol of measuring performance is weighted Root Mean Square Error (wRMSE). The experimental findings show that dual contextuality provided by EDURAG enhances the accuracy of grading significantly when compared to vanilla LLM grading.
Title: Bi-Contextual Retrieval Augmented Generation (RAG) for Automatic Descriptive Answer Grading
Description:
Automatic Short Answer Grading (ASAG) is a well-known research task in the field of natural language processing (NLP).
Its major purpose is to automatically grade descriptive answers of the students by keeping automatic grading consistent with the evaluation of human graders.
Recent developments in Large Language Models (LLMs) have demonstrated a greatly enhanced performance in automated grading; however, the generalizability of the models and accuracy is still quite low because of the absence of dataset-specific grounding.
We present EDURAG, a Retrieval-Augmented Generation (RAG) based model to improve contextualization of the LLM-based grading with exemplar-based grading and extra knowledge as generated by QFKE (Question Focused Knowledge Extraction) module.
The proposed QFKE module provides extra layer of contextuality for the EDURAG.
In contrast to conventional supervised methods, EDURAG does not need model fine-tuning.
The suggested framework is tested against the ASAG2024 benchmark that consolidates seven short-answer grading datasets across various domains, educational levels, and grading scales.
The benchmark protocol of measuring performance is weighted Root Mean Square Error (wRMSE).
The experimental findings show that dual contextuality provided by EDURAG enhances the accuracy of grading significantly when compared to vanilla LLM grading.

Related Results

DARE-RAG: Difficulty-Aware Retrieval Expansion for Retrieval-Augmented Generation
DARE-RAG: Difficulty-Aware Retrieval Expansion for Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) systems face a fundamental trade-off: query expansion can improve retrieval effectiveness for ambiguous or underspecified queries, yet indiscri...
JADE: jawbone lesion diagnosis and decision supporting system
JADE: jawbone lesion diagnosis and decision supporting system
Abstract Objectives To develop and evaluate JADE, a proof-of-concept retrieval-augmented generation (RAG) diagnostic assi...
Diagnosing RAG Failures: A Taxonomy of Failure Modes in Retrieval-Augmented Generation
Diagnosing RAG Failures: A Taxonomy of Failure Modes in Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) has become a standard paradigm for grounding large language models in external knowledge, yet RAG systems still fail in production in ways that...
A Systematic Literature Review of Retrieval-Augmented Generation Implementation for Enhancing Large Language Models in Education
A Systematic Literature Review of Retrieval-Augmented Generation Implementation for Enhancing Large Language Models in Education
The rapid advancement of Large Language Models (LLM) has led to the creation of increasingly adaptive intelligent learning systems. However, many educational implementations of LLM...
Summarizing Smarter: RAG vs Vanilla in Modern Transformers An Empirical Study on Retrieval-Augmented Generation Methods
Summarizing Smarter: RAG vs Vanilla in Modern Transformers An Empirical Study on Retrieval-Augmented Generation Methods
Recent advances in abstractive summarization have predominantly relied on pre-trained transformer models, which excel at generating fluent summaries but often fail to handle long o...
Automating Information Retrieval from Biodiversity Literature Using Large Language Models: A Case Study
Automating Information Retrieval from Biodiversity Literature Using Large Language Models: A Case Study
Recently, Large Language Models (LLMs) have transformed information retrieval, becoming widely adopted across various domains due to their ability to process extensive textual data...
Traditional RAG vs. Agentic RAG: A Comparative Study of Retrieval-Augmented Systems
Traditional RAG vs. Agentic RAG: A Comparative Study of Retrieval-Augmented Systems
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval to improve factual reliability. Traditional RAG employs a fix...

Back to Top