Reliability is when a particular approach (e.g., test) measures the same thing more than once and
produces the same outcomes
reliable measure is dependable, consistent, stable, trustworthy, and predictable
a reliable measure is a prerequisite for a valid measure
if a measure is not reliable is is not valid
The more reliable the measure, the more potential the measure will be valid
It is necessary to consider reliability because there is always error
If reliability is .40, the ceiling for validity is .40
Measured value = True value + Error. reduce error increases reliability
Measurement error may be divided into
systematic error (bias)
Will vary in a certain (systematic) way across levels of the independent variable
All sources of error (environment, participant, experimenter, measurement tool) have the potential
to produce systematic bias when they differ between the experimental conditions
Is a major threat to internal validity
random error
Is an unsystematicerror in measurement –it does not vary in any particular way across the levels of
the independent variable
It will influence individual scores
It might increase the measured value in some participants and decrease the measured value in
others
It has less effect on the group mean score
If error is truly random, it will balance out to zero .Will form a normal distribution
However, it will increase the variability in the scores and make the effects of the manipulation of the
IV harder to detect
sources of error
environmental
participants characteristic
experimenter
measurement tool
ways to asses reliability
test retest reliability
Used for most measurements Gives an indication of how stable the measure is over time Test
this by giving the same test/measure to the same individuals at two different times Measure the
reliability by examining the relationship between the values at Time 1 and Time 2
B Test-retest (correlation) coefficient can be calculated Simply calculate the correlation
coefficient and square it Test-retest reliability coefficient = r2 time1-time2 Reliability will vary
from 0 to +1 Adequate reliability is .60 or above (some use .70 as the cut-off
Measure administration
Administration should be standardised
Inter-observer (rater) agreement
Used when measurements are based on ratings/scores given by observers Gives an indication of
the extent to which different raters give the same measurement from the same information Test
this by having two raters independently rate the same behaviours, according to the same
measurement scale and criteria Agreement is percentage of ratings that the raters agree on Will
range from 0 to 100% The higher the bette
Inter-observer (rater) reliability
Used when measurements are based on ratings/scores given by observers Measures consistency
between raters For quantitative data, it is the square of the correlation between the two raters
Will range from 0 to 1. Indicates the percentage of relationship between the two raters. Should be
above .70
Inter-observer agreement/reliability Can be
improved by: Training raters
Simplifying their task (e.g., giving
them fewer categories, clear category
descriptions) Replacing a poor rater
Internal consistency
Used for questionnaire based measurements when multiple questions are used to measure the
same construct Gives an indication of how unified or similar (consistent) the individual items are
Several measures can be used to assess internal consistency
Average inter-item correlations
Correlate each individual item with each other and then compute the average to get the mean
inter-item correlation or median inter-item correlation. Should exceed .30
Split-half reliability coefficient
Split-half reliability coefficient Correlate the scores for one half of the items with the scores for the
other half of the items The “half” can be: First vs. second half Odd vs. even items Random
selection of items Should be above .70
Cronbach’s alpha
Is the most popular form of internal consistency Is conceptually the mean of all possible split-half
correlation coefficients Although it is not necessarily calculated that way Generally acceptable
value is .70 or better, but not too high Too high a value (> .90) indicates item redundancy Can be
calculated with SPS
Can be improved by: Increasing number of items .Refining less reliable items
Does the measurement use an established method or is it new? If established: Do they cite
evidence from the validation studies? Are the test conditions and sample in the article similar to
that in the original source? Look for calculation of reliability from own data Common with
internal consistency
When critiquing articles, look for evidence of reliability with the measures