Sampling is the procedure by which a researcher studies a part in order to say something about the whole. It rests on a simple insight with far-reaching consequences: if the part is selected by a known chance mechanism, the amount by which it is likely to differ from the whole can itself be calculated. Sampling is therefore not a regrettable compromise forced by scarce resources but a technique with its own logic, and a poorly drawn large sample is worth less than a well-drawn small one.
Four terms must be kept distinct. The population or universe is the entire set of units about which conclusions are wanted. The sampling frame is the actual list from which selection is made — an electoral roll, a village household register, a school directory. The sampling unit is the element selected, which need not be the unit of analysis: one may sample villages in order to study households. Sampling error is the discrepancy between a sample statistic and the population value arising purely from the accident of selection; it shrinks as sample size grows and can be estimated only under probability designs.
Probability designs
In probability sampling every unit in the frame has a known, non-zero chance of selection.
Simple random sampling gives every unit an equal chance, by lottery or random numbers. It is the theoretical benchmark, but it requires a complete frame and scatters the sample geographically, which is expensive in field costs.
Systematic sampling selects every kth unit after a random start, where k is the sampling interval. It is quick and easy for investigators to execute, but a periodicity in the list that coincides with the interval will produce a badly biased sample.
Stratified sampling first divides the population into homogeneous strata — region, caste category, rural and urban, income band — and samples within each. If the strata genuinely differ on the variable of interest, this yields more precise estimates than simple random sampling for the same cost. Proportionate allocation matches each stratum's share in the sample to its share in the population; disproportionate allocation deliberately over-samples small groups, such as a minority community, so that separate estimates for them are usable, and compensates with weights at the analysis stage.
Cluster sampling selects naturally occurring groups — villages, wards, schools — and studies all units within the chosen clusters. It cuts travel and frame costs dramatically, but because members of a cluster resemble each other, precision falls for a given sample size.
Multistage sampling combines these in a hierarchy: districts, then villages within selected districts, then households within selected villages. Almost every large national survey, the National Sample Survey among them, uses a stratified multistage design, because no country possesses a usable single list of all its residents.
Non-probability designs
Here the chance of selection is unknown, so sampling error cannot be estimated and statistical inference is not licensed.
Quota sampling fixes the number of respondents required in each category — so many women, so many under thirty — and leaves the interviewer to fill the quota as convenient. It reproduces the population's profile on the quota variables while allowing selection bias on everything else. Purposive or judgemental sampling picks units precisely because they are theoretically informative — a typical village, an extreme case, a critical instance — and is the appropriate design for qualitative and case-study work. Convenience sampling takes whoever is available, which is why classroom and street-corner studies rarely support general claims. Snowball sampling uses initial contacts to name further respondents, who name others in turn, so the sample accumulates through the network itself.
Snowball sampling is often the only available design. For hidden, mobile or stigmatised populations — undocumented migrants, sex workers, people living with HIV, users of illegal drugs, members of clandestine political groups — no sampling frame exists and none can be constructed, since the defining characteristic is one people conceal. The cost is a sample biased towards the well-connected and towards whichever subnetwork the researcher first entered; respondent-driven sampling attempts to correct this with structured referral chains and weighting by network size.
Why probability sampling licenses inference
The privilege of probability designs is not that they produce representative samples — any single sample may be atypical — but that the distribution of possible samples is known. From that distribution follow the standard error, the confidence interval and the tests of significance. Without a known selection probability, a percentage from a sample is a description of those respondents and nothing more. Precision depends chiefly on absolute sample size, not on the sampling fraction, which is why a national sample of a few thousand can serve a population of a billion. And sampling error is only one component of total error: coverage gaps, non-response and measurement error do not diminish with sample size at all.
Frame problems in Indian research
The frame is where Indian sampling usually breaks. Electoral rolls omit the young and the recently migrated and carry duplicates. Village household lists age quickly in areas of heavy circular migration, so seasonal labourers are systematically missed — a distortion that runs against exactly the populations sociologists most want to study. Urban slums and unauthorised colonies are under-listed or absent from municipal records. Telephone and online panels exclude those without devices, skewing against women, the elderly and the poor. Within a sampled household, the convention of interviewing the head silently substitutes his account for that of women and younger members. Hence the reliance on multistage designs anchored in census enumeration blocks, on fresh household listing in the field, and on explicit weighting.
For the UPSC answer
Open by defining population, frame, sampling unit and sampling error, because these four terms let you discuss everything else precisely. Present probability designs as a family — simple random, systematic, stratified, cluster, multistage — and say what each buys and what it costs, then set the non-probability designs against them, noting that they are not inferior but answer a different question. The central theoretical point to state plainly is that probability selection is what makes statistical inference legitimate, and that snowball sampling is nonetheless the only option for hidden populations. Finish with Indian frame problems — migrant workers, unlisted settlements, the household head as proxy — which shows methodological judgement rather than memorised lists.
References & further reading
- Kish, L. (1965). Survey Sampling. Wiley.
- Cochran, W. G. (1977). Sampling Techniques. Wiley.
- Moser, C. A., and Kalton, G. (1971). Survey Methods in Social Investigation. Heinemann.
- Mahalanobis, P. C. (1944). On Large-Scale Sample Surveys. Philosophical Transactions of the Royal Society of London, Series B, 231, 329–451.
- Goodman, L. A. (1961). Snowball Sampling. Annals of Mathematical Statistics, 32(1), 148–170.
- Heckathorn, D. D. (1997). Respondent-Driven Sampling: A New Approach to the Study of Hidden Populations. Social Problems, 44(2), 174–199.