Javascript must be enabled to continue!
Keeping Elo alive: Evaluating and improving measurement properties of learning systems based on Elo ratings
View through CrossRef
The Elo Rating System which originates from competitive chess has been widely utilised in large-scale online educational applications where it is used for on-the-fly estimation of ability, item calibration, and adaptivity. In this paper, we aim to critically analyse the shortcomings of the Elo rating system in an educational context, shedding light on its measurement properties and when these may fall short in precisely and reliably capturing student abilities and item difficulties. In a simulation study, we look at the asymptotic properties of the Elo rating system. Our results show that the Elo ratings are generally not unbiased and their variances are context-dependent. Furthermore, in scenarios where items are selected adaptively based on the current ratings and the item difficulties are updated alongside the student abilities, the variance of the ratings across items and students artificially increases over time and as a result, the ratings do not converge. We propose a solution to this problem which entails using two parallel chains of ratings which remove the dependence of item selection on the current errors in the ratings.
Title: Keeping Elo alive: Evaluating and improving measurement properties of learning systems based on Elo ratings
Description:
The Elo Rating System which originates from competitive chess has been widely utilised in large-scale online educational applications where it is used for on-the-fly estimation of ability, item calibration, and adaptivity.
In this paper, we aim to critically analyse the shortcomings of the Elo rating system in an educational context, shedding light on its measurement properties and when these may fall short in precisely and reliably capturing student abilities and item difficulties.
In a simulation study, we look at the asymptotic properties of the Elo rating system.
Our results show that the Elo ratings are generally not unbiased and their variances are context-dependent.
Furthermore, in scenarios where items are selected adaptively based on the current ratings and the item difficulties are updated alongside the student abilities, the variance of the ratings across items and students artificially increases over time and as a result, the ratings do not converge.
We propose a solution to this problem which entails using two parallel chains of ratings which remove the dependence of item selection on the current errors in the ratings.
Related Results
Complex Collision Tumors: A Systematic Review
Complex Collision Tumors: A Systematic Review
Abstract
Introduction: A collision tumor consists of two distinct neoplastic components located within the same organ, separated by stromal tissue, without histological intermixing...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
Comparison of alive and dead benthic foraminiferal fauna off the Changjiang Estuary: Understanding water-mass properties and taphonomic processes
Comparison of alive and dead benthic foraminiferal fauna off the Changjiang Estuary: Understanding water-mass properties and taphonomic processes
Benthic foraminifera (BF) are utilized in palaeo-environmental reconstruction based on our understanding of how living individuals respond to environmental variations. However, the...
Hot and Easy in Florida: The Case of Economics Professors
Hot and Easy in Florida: The Case of Economics Professors
The objective of this study is threefold. First, we seek to investigate whether overall quality ratings recorded by students on RateMyProfessors.com (henceforth RMP) are positively...
Using simulations to compare the current Davis Cup ranking system to Elo
Using simulations to compare the current Davis Cup ranking system to Elo
The Davis Cup is the premier men’s team event in tennis, run by the International Tennis Federation and in which over 130 nations compete. It uses a merit-based ranking system that...
Consistency of perceiving odors: Inter- and Intra- Individual Differences in Odor Similarity Ratings
Consistency of perceiving odors: Inter- and Intra- Individual Differences in Odor Similarity Ratings
This study conducted odor pair similarity ratings twice and used the replicability of the ratings, as indicated by the correlation coefficients between the two sets of ratings, to ...
ALIVE Biofeedback HRV training for Treating Insomnia: A Pilot Randomized Controlled Study.
ALIVE Biofeedback HRV training for Treating Insomnia: A Pilot Randomized Controlled Study.
Background: Insomnia is a common sleep disorder that affects a large portion of the population. While several treatments are available, such as medication and cognitive-behavioral ...
The Influence of the Children Learning in Science (CLIS) Learning Model and Students' Learning Styles on Improving Students' Conceptual Understanding and Motivation to Learn Mathematics
The Influence of the Children Learning in Science (CLIS) Learning Model and Students' Learning Styles on Improving Students' Conceptual Understanding and Motivation to Learn Mathematics
The Children Learning in Science (CLIS) learning model emphasizes the process of reconstructing students' understanding by directing changes in initial misconceptions towards a mor...

