maandag 18 mei 2020

Assessing quality of items and tests


Quality of tests need to be assessed before they can be used to test knowledge or skills of the candidates. The RCEC review system is an analytical review system that is developed to evaluate the quality of educational exams (‘The RCEC review system for the quality of tests and exams’, n.d.). This system has six criteria that together make up the substantive and organizational aspect and the psychometric aspect. Purpose and use, test and examination material, and test administration and security combined form the first aspect. Representativeness, reliability, and standard setting and maintenance form the second aspect. To measure if the six criteria are met, questions are answered with either ‘insufficient’, ‘sufficient’, or ‘good’, this gives respectively a score of 1, 2, or 3 to the question. At the end of the questions for each criterium is checked whether enough points are gathered to be able to say if that criterium is met. In this post, a few questions of the criteria for the psychometric aspect are answered with data that was provided. This data contains the analysis of the test and items. It was a test for group 7 (grade 5), 199 students participated, and the test consisted of 40 items.

For criterium 3, representativeness, the question 3.2 ‘Is the degree of difficulty of the items and/or the actions adjusted to the intended target population?’ was selected. To be able to answer this question with sufficient, 75% - 90% of the items should have a p-value >0.20 and ≤0.80. If the percentage is lower than 75, the question is marked insufficient and if the percentage is higher than 90, the question is marked as good. When looking at this data, less than 75% of the items has a p-value between 0.20 and 0.80. Therefore, this question must be answered with insufficient and gets a score of 1.

For criterium 4, reliability, the questions 4.2 and 4.3 were selected. Question 4.2 ‘Is the reliability of the test correctly calculated?’ is answered by the number of candidates used for the calculation of the reliability. At least 200 candidates should be used for the calculation however, in this data, only 199 candidates took the test. Therefore, the answer to this question is insufficient and gets a score of 1. If there would have been 200 candidates, the score would have gone up to sufficient. Additionally, there was an objective scoring system, established in question 2.9 (criterium 2, question 9), therefore the score would go to good. So, with at least one extra candidate, the score of this question would go from 1 to 3.
The second question in this criterium is 4.3, ‘Is the reliability sufficient, considering the decisions that have to be based on the test?’. To answer this question is looked at the reliability score. A reliability between ≥0.80 and <0.90 is considered sufficient. Lower than 0.80 is insufficient and higher than 0.90 is good. In this data, the coefficient alpha is only 68%, therefore also this question is answered with insufficient and gets a score of 1.

For criterium 5, standard setting and maintenance, the questions 5.1, 5.2a, and 5.2c were selected. Question 5.1 is ‘Are norms/ standards/ cut-off scores provided?’. So, either these are/ one of these is given or not. The data shows that the Angoff method is used and the cut-off score has been set. So, this question can be marked as good and gets a score of 3.
The second question is 5.2 ‘Has the standard setting been carried out correctly?’, which is divided in three sub questions. However, only sub question a and c will be discussed.
Sub question a is ‘Has the standard setting method been carried out correctly?’. To answer this question professional consideration or argumentation to support the decision for the cut-off score needs to be considered. The Angoff method was used to set the cut-off score and seems to be carried out correctly, however, the reasoning and support of the experts is missing. Therefore, this question is answered as sufficient and gets a score of 2.
Sub question c is ‘Is there sufficient agreement between the qualified experts?’. Sufficient agreement is between 0.60 and 0.80. In this data, the agreement between the qualified experts is 89% which means that this question can be answered with good and thus gets a score of 3.

In summary, the review has strict rules with which the quality evaluation is executed. However, it is not always as straightforward as it might seem, for example, look at criterium 4. Most importantly, no conclusion can be drawn from answering a few questions since all questions must be answered to produce a reliable evaluation of the quality of the test.


Reference
The RCEC review system for the quality of tests and exams. (n.d.). Retrieved 18 May 2020, from https://www.rcec.nl/en/review-system/

maandag 11 mei 2020

Comparing standard setting methods

Introducing methods for standard setting
Tests are made and taken to see how well students understand a certain topic. When developing a test, a lot of decisions need to be made. All these decisions combined are called the standard setting process, in which is determined how well someone has to perform on a test to pass that test. It also includes setting performance standards, making exam questions, and selecting a method for setting a cut-score. The cut-sore represents the least number of items that need to be answered correctly to pass the test. In this post will be focused on explaining and comparing different methods to set a cut-score.
            There are a few methods for standard setting, and they can be divided in three subgroups, norm-referenced, criterion-referenced, and mixed. In norm-referenced methods students are compared to each other, while criterion-referenced methods are chosen when the student needs a certain level of knowledge or skills to be able to pass the test (Ertoprak & Dogan, 2016). It is also possible to use a mix of those methods.
            The passing percentage and Cohen’s method are examples of norm-referenced methods. The passing percentage method can be used when there is a desired percentage of students who need to pass. This can be due to selection or limited places available. A reason not to use this method is because the content and quality are not considered when deciding if people passed. Even if the test were made badly, people would still be able to pass. The Cohen’s method is similar; however, the best performing student is used as reference and 60-65% of that score is used to determine the cut-score. The advantage is that student ability across exams is more stable than panelist rating, additionally panelist ratings can be too expensive. On the other hand, it might be that student ability across exams fluctuates too much.
            The linear transformation and expert panels are examples of criterion-referenced methods. The linear transformation draws a straight line between the guess-score and the maximum point that can be obtained. This method can be used when there are no differences in difficulty between exams, guessing score, and maximum score. However, if there are differences, this method cannot be used. The second method entails expert panels, this is also called Angoff-method. Around ten panelists estimate the probability that a minimal competent student answers an item correctly. They do this for all items on a test. With those experts setting the cut-score, criterium reference, quality assurance, and minimal competence are clear. Additionally, professionals are engaged in the process. Since this is a time and money consuming process it might not always be the best method to choose.
            Lastly, the Hofstee-method is a combination of norm- and criterion-referenced methods. Experts decide on an acceptable passing sore, they set the minimum and the maximum failure rate and the minimum and maximum passing score. This is done to control for extreme failure rates by critical panelists. A reason not to use this method is that the ability of examinees across exams can differ a lot.

Analysing methods for setting the passing percentage and cut-score
The analysis will be performed with data from a high-stakes Mathematics exam to find similarities and differences between the previously discussed methods. The outcomes and comparisons are discussed below. A few details of the data: students could obtain 66 points in the exam, the guess-score was 16.5, and 1945 students took the exam.
            Looking at all methods, students had to answer between 40 and 44 items correctly to pass the test. The passing percentage method shows that if 51% of the students should pass, the cut-score should be set at 41. To let 57% of the students pass, the cut-score should be 40. The linear method shows that students had to answer 41.5 items correctly to get the passing grade of 5.5. Since it is not possible to get this score, the cut-score should be set to 41 or 42. A score of 41 results in a 5.0, while answering 42 items correctly results in a 6.0. There is a big difference in how is decided if students pass the test when looking at these two methods.
            When setting the cut-scores with the other methods, it becomes clear that there are less differences. The graph of the Hofstee-method shows an intersection that provides a cut-score of 41, in this case, 49% of the students will fail the test. Twelve panelists performed the Angoff-method which resulted in a cut-score of 40. Lastly, the Cohen’s method shows a cut-score of 41 and then 49,1% of the students failing. So, the Hofstee- and Cohen’s method have similar results. The Angoff-method gives a cut-score that is a little lower. If this cut-score would be used in the Hofstee- and Cohen’s method, the percentage of students failing would lower to respectively 43% and 43,2%.

Changing student ability, what happens to the passing percentage and cut-score?
If the students who would take the same test would have a lower ability level, there would be some changes in the passing percentage and cut-score. Firstly, looking at the passing percentage method and linear transformation method, there would be less items that need to be answered correctly to pass the test. If still around 60% of the students need to pass and they all score lower on the test, they will have to answer less than 40 items correctly. When looking at the linear transformation method, if the highest score of one of the students is still 66, there would not be a difference in the number of items that need to be answered correctly to pass the test. However, if the highest score on that test would not be 66, the number of items that need to be answered correctly to pass, would be lower. So, using the same exam in a group in which student’s abilities are lower, the passing percentage could be different depending on the method that is used and the highest score on the test.
            Also changes in the cut-score are dependent on the method that is used to set the cut-score. The Hofstee-method provides a range in which the cut-score can lie, the score will be different when the student’s performance is less because of the change in the cumulative graph, not because the experts have set other values for the minimum and maximum failure rate and  passing score. If the students score lower on the test, the cut-score will be lower. The cut-score determined by the Angoff-method will be different because the probability of students answering the items correct is considered. So, if the student’s ability is lower, the cut-score will be lower. Lastly, the cut-score set by the Cohen’s method can be different since it depends on the highest score in the group. So, the same holds as for the linear transformation method. If the best performing student now scores lower than in the previous group, the cut-score will also be lower. If the best performing student performs equally well as in the previous group, the cut-score will not change.

References
Ertoprak, D. G., & Dogan, N. (2016). A research on the classification validity of the decisions made according to norm and criterion-referenced assessment approaches. Anthropologist, 23(3), 612–619. https://doi.org/10.1080/09720073.2014.11891981

vrijdag 24 april 2020

Analysing data with item response theory


To gain more insight into how a test can measure someone’s knowledge it is useful to analyse the relationship between the questions of a test and the ability of respondents. This can be done with Item Response Theory (IRT). IRT represents the relationship between items in a test and the latent traits (e.g. someone’s ability in math). There are different models to describe this relationship and for this assignment three were examined, namely the Rasch, 2PL (parameter logistic), and 3PL model. The Rash model only takes the difficulty of the items into account. The 2PL model adds discrimination and the 3PL model considers difficulty, discrimination, and guessing. To learn how to find a fitting model for data, a Graduate Management Admission Test (GMAT) dataset of ShinyItemAnalysis was used (https://shiny.cs.cas.cz/ShinyItemAnalysis/). 
At this website firstly, at the Data tab the GMAT2 dataset was loaded. Secondly, at the IRT tab, the subtab ‘Rasch model’ was selected. At this page, the item characteristic curves, item information curves, test information function, table of estimated parameters, ability estimates, scatter plot of factor scores and standardized total scores, and wright map are shown. This page was inspected to learn about the characteristics of the items. The same was done for the 2PL and 3PL models. See table 1 -3 for the estimated item parameters of the various models. 
Lastly, the subtab ‘model comparison’ was selected to view the comparison of ShinyItemAnalysis of the three models that were taken a closer look at. This page shows a table of comparison statistics (see figure 1) in which the models are compared, and the best-fitting model is shown. In four out of five cases the table indicates that the 2PL model has the best fit with the data based on the comparison statistics of ShinyItemAnalysis. So, the 2PL model fits the data best. As said before, difficulty and discrimination are the two parameters of the 2PL model. Discrimination is defined as how well an item able to differentiate between people with higher and lower ability than the difficulty of the item. In figure 2, the item characteristic curves (ICC’s) are shown of the 20 items when they are analysed with the 2PL model. The ICC’s show the relationship between the difficulty in the various items and the chance a person answers the items correctly. The steeper the graph in the ICC, the more an item discriminates and is thus more informative.

Table 1

Item parameters for the Rasch model

a
SE(a)
b
SE(b)
c
SE(c)
Item 1
1.00
-
-0.11
0.07
0.00
-
Item 2
1.00
-
-0.39
0.07
0.00
-
Item 3
1.00
-
-0.93
0.07
0.00
-
Item 4
1.00
-
-1.31
0.08
0.00
-
Item 5
1.00
-
-1.49
0.08
0.00
-
Item 6
1.00
-
-0.58
0.07
0.00
-
Item 7
1.00
-
-0.65
0.07
0.00
-
Item 8
1.00
-
-0.51
0.07
0.00
-
Item 9
1.00
-
-0.32
0.07
0.00
-
Item 10
1.00
-
-0.15
0.07
0.00
-
Item 11
1.00
-
-0.75
0.07
0.00
-
Item 12
1.00
-
-0.41
0.07
0.00
-
Item 13
1.00
-
-1.26
0.08
0.00
-
Item 14
1.00
-
0.34
0.07
0.00
-
Item 15
1.00
-
0.00
0.07
0.00
-
Item 16
1.00
-
0.29
0.07
0.00
-
Item 17
1.00
-
0.17
0.07
0.00
-
Item 18
1.00
-
0.51
0.07
0.00
-
Item 19
1.00
-
0.09
0.07
0.00
-
Item 20
1.00
-
0.23
0.07
0.00
-


Table 2

Item parameters for the 2PL model

a
SE(a)
b
SE(b)
c
SE(c)
Item 1
0.70
0.11
-0.17
0.10
0.00
-
Item 2
0.82
0.12
-0.51
0.10
0.00
-
Item 3
0.25
0.10
-3.58
1.36
0.00
-
Item 4
0.41
0.11
-3.15
0.8
0.00
-
Item 5
0.67
0.12
-2.29
0.38
0.00
-
Item 6
0.69
0.11
-0.87
0.15
0.00
-
Item 7
0.44
0.10
-1.45
0.33
0.00
-
Item 8
0.49
0.10
-1.02
0.23
0.00
-
Item 9
0.23
0.09
-1.32
0.56
0.00
-
Item 10
0.44
0.09
-0.34
0.17
0.00
-
Item 11
0.49
0.10
-1.52
0.32
0.00
-
Item 12
0.35
0.09
-1.14
0.34
0.00
-
Item 13
0.46
0.11
-2.73
0.62
0.00
-
Item 14
0.76
0.11
0.47
0.11
0.00
-
Item 15
0.47
0.09
0.01
0.14
0.00
-
Item 16
0.75
0.11
0.41
0.11
0.00
-
Item 17
0.29
0.09
0.59
0.28
0.00
-
Item 18
1.03
0.13
0.56
0.09
0.00
-
Item 19
0.74
0.11
0.13
0.10
0.00
-
Item 20
0.32
0.09
0.70
0.28
0.00
-


Table 3

Item parameters for the 3PL model

a
SE(a)
b
SE(b)
c
SE(c)
Item 1
0.86
0.40
0.28
0.85
0.14
2.22
Item 2
0.82
0.12
-0.50
0.13
0.00
10.24
Item 3
0.83
0.87
1.69
0.76
0.62
0.51
Item 4
1.42
1.19
1.01
0.52
0.70
0.37
Item 5
0.66
0.13
-2.30
0.47
0.01
10.46
Item 6
1.17
0.61
0.36
0.67
0.37
0.81
Item 7
0.45
0.15
-1.28
1.76
0.04
10.57
Item 8
0.52
0.17
-0.86
1.50
0.03
11.29
Item 9
0.60
0.62
2.04
0.97
0.44
0.77
Item 10
1.25
0.84
1.37
0.32
0.41
0.40
Item 11
0.80
0.66
0.10
2.00
0.36
1.93
Item 12
0.35
0.10
-1.09
0.70
0.01
10.82
Item 13
0.47
0.12
-2.59
1.31
0.03
10.58
Item 14
1.07
0.46
0.88
0.36
0.15
1.05
Item 15
0.53
1.09
0.39
6.63
0.09
19.26
Item 16
0.74
0.11
0.43
0.14
0.00
10.41
Item 17
0.32
0.11
0.66
1.21
0.02
10.08
Item 18
2.84
1.57
0.95
0.12
0.22
0.31
Item 19
2.11
1.02
0.96
0.17
0.32
0.29
Item 20
2.62
1.92
1.72
0.22
0.40
0.12





Figure 1. Screenshot of the comparison statistics table.



Figure 2. Item characteristic curves, 2PL model.

Defining educational measurement and describing its innovations and future

In the last eight weeks, I have learned about different topics regarding educational measurement as can be seen in the previous six blogpo...