Breaking down the question

The question has two clear parts: define reliability, then explain the different tests by which a researcher establishes it. At 10 marks the answer must be crisp and complete — a precise definition followed by a systematic account of the main reliability tests, ideally with a closing note distinguishing reliability from validity.

Reliability concerns the consistency and stability of a measure: whether a research instrument yields the same results when applied repeatedly under the same conditions. A reliable measure is dependable and free from random error, so that differences in results reflect real differences rather than the fluctuations of the tool itself.

The examiner is testing methodological precision, so the tests should be named accurately and each explained in a sentence or two.

How to approach it

  • Define reliability as consistency of measurement, contrasting it briefly with validity.
  • Explain the principal tests: test-retest, parallel-forms (equivalent-forms), split-half, and inter-rater reliability, adding internal consistency where relevant.
  • Note in closing that reliability is necessary but not sufficient — a measure can be consistent yet still wrong.

Draw on the variables, sampling, reliability and validity notes.

Model answer

Reliability refers to the consistency, stability and repeatability of a measurement. A research instrument — a scale, questionnaire, test or coding scheme — is reliable if it produces the same results on repeated applications under the same conditions, so that the findings are dependable rather than the product of random error. If a measure of, say, attitudes to work yields wildly different scores each time it is applied to the same unchanged respondents, it is unreliable and its findings cannot be trusted. Reliability should be distinguished from validity, which asks whether the instrument measures what it claims to measure; a measure can be reliable without being valid.

Social scientists use several tests to establish reliability.

The first is test-retest reliability. The same instrument is administered to the same group of respondents on two separate occasions, and the two sets of scores are correlated. A high correlation indicates that the measure is stable over time. Its limitation is the risk of memory or practice effects, and of genuine change in the respondents between the two administrations.

The second is parallel-forms or equivalent-forms reliability. Two different but equivalent versions of the instrument, designed to measure the same construct, are administered to the same respondents, and the results are correlated. A strong correlation shows that the two forms are consistent measures. The difficulty lies in constructing two genuinely equivalent forms.

The third is split-half reliability, a test of internal consistency. The items of a single instrument are divided into two halves — for example odd and even numbered items — and the scores on the two halves are correlated. A high correlation shows that the items are measuring the same underlying construct consistently. A refinement of this idea is Cronbach's alpha, developed by Lee Cronbach, which estimates internal consistency across all possible splits of the items and is the most widely used single coefficient of reliability.

The fourth is inter-rater or inter-observer reliability, important in qualitative coding and observational research. Two or more independent researchers apply the same coding scheme or observation schedule to the same material, and the degree of agreement between them is measured. High agreement indicates that the instrument yields consistent results regardless of who applies it, reducing subjective bias.

In sum, reliability is the foundation of dependable measurement, and these complementary tests — assessing stability over time, equivalence across forms, internal consistency among items, and agreement across observers — allow the researcher to demonstrate that an instrument is trustworthy. Yet reliability remains a necessary but not a sufficient condition of good measurement, for a consistently applied instrument may still be measuring the wrong thing; hence it must always be paired with validity.

Examiner's perspective

For a 10-mark answer the examiner expects a sharp definition and then the tests named correctly and explained distinctly. Vague answers that treat reliability as mere accuracy, or that list only one test, lose easy marks.

The strongest scripts cover test-retest, parallel-forms, split-half and inter-rater reliability, mention Cronbach's alpha as the standard internal-consistency measure, and close by distinguishing reliability from validity. Correct methodological vocabulary within a tight word budget is what secures the higher band.