Введение
Процесс статистического отбора
В статистике, выборочный опрос описывает процесс отбора выборки элементов из генеральной совокупности для проведения опроса. Термин «опрос» может относиться к различным типам или методам наблюдения. В выборочном опросе он чаще всего включает в себя анкету, используемую для измерения характеристик и/или установок людей. Различные способы установления контакта с участниками выборки после их отбора являются предметом сбора данных опроса. Цель выборки – сократить стоимость и/или объем работы, необходимый для опроса всей генеральной совокупности. Опрос, охватывающий всю генеральную совокупность, называется переписью. Выборка относится к группе или части совокупности, из которой требуется получить информацию. Выборки опросов можно разделить на два основных типа: вероятностные и сверхвыборки. Вероятностная выборка предполагает реализацию плана выборки с заданными вероятностями (возможно, адаптированными вероятностями, определенными адаптивной процедурой). Вероятностная выборка позволяет проводить выводы о генеральной совокупности на основе дизайна. Выводы основаны на известном объективном распределении вероятностей, указанном в протоколе исследования. Выводы, полученные в результате вероятностных опросов, все еще могут быть подвержены различным видам систематических ошибок. Опросы, не основанные на вероятностной выборке, сложнее оценить с точки зрения систематической ошибки или ошибки выборки. Опросы, основанные на не вероятностных выборках, часто не отражают состав генеральной совокупности. В академических и государственных исследованиях вероятностная выборка является стандартной процедурой. В Соединенных Штатах «Список стандартов статистических обследований» Управления по бюджету и управлению гласит, что финансируемые из федерального бюджета опросы должны проводиться путем: отбора выборок с использованием общепринятых статистических методов (например, вероятностных методов, которые могут предоставить оценки ошибки выборки). Любое использование не вероятностных методов выборки (например, усеченных или модельных выборок) должно быть статистически обосновано и обеспечивать возможность измерения ошибки оценки. Случайная выборка и выводы на основе дизайна дополняются другими статистическими методами, такими как выборка с использованием модели и выборка на основе модели. Например, во многих опросах наблюдается значительный уровень отказа от ответов. Несмотря на то, что единицы первоначально отбираются с известными вероятностями, механизмы отказа от ответов неизвестны. Для опросов с существенным уровнем отказа от ответов статистиками предложены статистические модели, используемые для анализа наборов данных. Вопросы, связанные с выборочными опросами, обсуждаются в различных источниках, включая Salant и Dillman (1994).
In statistics, survey sampling describes the process of selecting a sample of elements from a target population to conduct a survey. The term "survey" may refer to many different types or techniques of observation. In survey sampling it most often involves a questionnaire used to measure the characteristics and/or attitudes of people. Different ways of contacting members of a sample once they have been selected is the subject of survey data collection. The purpose of sampling is to reduce the cost and/or the amount of work that it would take to survey the entire target population. A survey that measures the entire target population is called a census. A sample refers to a group or section of a population from which information is to be obtained. Survey samples can be broadly divided into two types: probability samples and super samples. Probability based samples implement a sampling plan with specified probabilities (perhaps adapted probabilities specified by an adaptive procedure). Probability based sampling allows design based inference about the target population. The inferences are based on a known objective probability distribution that was specified in the study protocol. Inferences from probability based surveys may still suffer from many types of bias. Surveys that are not based on probability sampling have greater difficulty measuring their bias or sampling error. Surveys based on non probability samples often fail to represent the people in the target population. In academic and government survey research, probability sampling is a standard procedure. In the United States, the Office of Management and Budget's "List of Standards for Statistical Surveys" states that federally funded surveys must be performed:
selecting samples using generally accepted statistical methods (e. g., probabilistic methods that can provide estimates of sampling error). Any use of nonprobability sampling methods (e. g., cut off or model based samples) must be justified statistically and be able to measure estimation error. Random sampling and design based inference are supplemented by other statistical methods, such as model assisted sampling and model based sampling. For example, many surveys have substantial amounts of nonresponse. Even though the units are initially chosen with known probabilities, the nonresponse mechanisms are unknown. For surveys with substantial nonresponse, statisticians have proposed statistical models with which the data sets are analyzed. Issues related to survey sampling are discussed in several sources, including Salant and Dillman (1994).
Вероятностный отбор
В вероятностной выборке (также называемой «научной» или «случайной» выборкой) каждый член целевой популяции имеет известную и ненулевую вероятность включения в выборку. Опрос, основанный на вероятностной выборке, теоретически может давать статистические оценки целевой популяции, не имеющие систематической ошибки, поскольку математическое ожидание среднего значения выборки равно среднему значению генеральной совокупности, E(ȳ)=μ, либо иметь измеримую ошибку выборки, которую можно выразить в виде доверительного интервала или погрешности. Вероятностная выборка для опроса создается путем составления списка целевой популяции, называемого базой выборки, рандомизированной процедуры отбора единиц из базы выборки, называемой процедурой отбора, и метода установления контакта с отобранными единицами для участия в опросе, называемого методом или способом сбора данных. Для некоторых целевых популяций этот процесс может быть простым; например, выборка сотрудников компании по спискам заработной платы. Однако в больших, неорганизованных популяциях само создание подходящей базы выборки часто является сложной и дорогостоящей задачей. Распространенными методами проведения вероятностной выборки населения домохозяйств в Соединенных Штатах являются выборка по областям, выборка с помощью случайного набора телефонных номеров и, в последнее время, выборка по адресам. В рамках вероятностного отбора существуют специализированные методы, такие как стратифицированная и кластерная выборки, которые повышают точность или эффективность процесса выборки, не изменяя фундаментальные принципы вероятностного отбора. Стратификация – это процесс разделения членов популяции на однородные подгруппы перед отбором, основанный на дополнительной информации о каждой единице выборки. Страты должны быть взаимоисключающими: каждый элемент в популяции должен быть отнесен только к одному страту. Страты также должны быть полными: ни один элемент популяции не должен быть исключен. Затем в каждом страте могут применяться такие методы, как простая случайная или систематическая выборка. Стратификация часто повышает репрезентативность выборки за счет снижения ошибки выборки.
Склонности в вероятностном выборе образцов
Пристрастия в опросах нежелательны, но часто неизбежны. Основные типы систематической ошибки, которые могут возникнуть в процессе выборки:
Ошибка, связанная с отказом от участия: когда отдельные лица или домохозяйства, отобранные в выборке опроса, не могут или не хотят принять участие в опросе, существует вероятность возникновения систематической ошибки из-за этого отказа. Ошибка, связанная с отказом от участия, возникает, когда наблюдаемое значение отклоняется от параметра генеральной совокупности из-за различий между респондентами и нереспондентами. Ошибка, связанная с ответами: это не противоположность ошибке, связанной с отказом от участия, а скорее отражает возможную тенденцию респондентов давать неточные или ложные ответы по разным причинам. Ошибка отбора: возникает, когда у некоторых единиц выборки вероятность отбора отличается и это не учитывается исследователем. Например, некоторые домохозяйства имеют несколько телефонных номеров, что повышает вероятность их отбора в телефонном опросе по сравнению с домохозяйствами, имеющими только один номер. Эта ошибка отбора может быть исправлена путем применения веса опроса, равного [1/(# телефонных номеров)] для каждого домохозяйства. Самоотбор: тип систематической ошибки, при котором люди добровольно включаются в группу, что потенциально искажает результаты этой группы. Ошибка участия: систематическая ошибка, возникающая из-за характеристик тех, кто решил принять участие в опросе. Ошибка охвата: возникает, когда члены генеральной совокупности отсутствуют в базе выборки (неполный охват). Ошибка охвата возникает, когда наблюдаемое значение отклоняется от параметра генеральной совокупности из-за различий между охваченными и неохваченными единицами. Телефонные опросы подвержены известному источнику ошибки охвата, поскольку они не могут включать домохозяйства без телефонов.
Non response bias: When individuals or households selected in the survey sample cannot or will not complete the survey there is the potential for bias to result from this non response. Nonresponse bias occurs when the observed value deviates from the population parameter due to differences between respondents and nonrespondents. Response bias: This is not the opposite of non response bias, but instead relates to a possible tendency of respondents to give inaccurate or untruthful answers for various reasons. Selection Bias: Selection bias occurs when some units have a differing probability of selection that is unaccounted for by the researcher. For example, some households have multiple phone numbers making them more likely to be selected in a telephone survey than households with only one phone number. This selection bias would be corrected by applying a survey weight equal to [1/(# of phone numbers)] to each household. Self selection bias: A type of bias in which individuals voluntarily select themselves into a group, thereby potentially biasing the response of that group. Participation bias: Bias that arises due to the characteristics of those who choose to participate in a survey or poll. Coverage bias: Coverage bias can occur when population members do not appear in the sample frame (undercoverage). Coverage bias occurs when the observed value deviates from the population parameter due to differences between covered and non covered units. Telephone surveys suffer from a well known source of coverage bias because they cannot include households without telephones.
Невероятностный отбор проб
Многие опросы не основываются на вероятностных выборках, а на поиске подходящей группы респондентов для участия в опросе. Некоторые распространенные примеры невыборочных методов выборки:
Judgement Samples: A researcher decides which population members to include in the sample based on his or her judgement. The researcher may provide some alternative justification for the representativeness of the sample. The underlying assumption is that the investigator will select units that are characteristic of the population. This method can be subjected to researcher's biases and perception. Snowball Samples: Often used when a target population is rare. Members of the target population recruit other members of the population for the survey. Quota Samples: The sample is designed to include a designated number of people with certain specified characteristics. For example, 100 coffee drinkers. This type of sampling is common in non probability market research surveys. Convenience Samples: The sample is composed of whatever persons can be most easily accessed to fill out the survey. In non probability samples the relationship between the target population and the survey sample is immeasurable and potential bias is unknowable. Sophisticated users of non probability survey samples tend to view the survey as an experimental condition, rather than a tool for population measurement, and examine the results for internally consistent relationships.
Выборки на основе экспертной оценки: Исследователь решает, каких членов генеральной совокупности включить в выборку, основываясь на своем экспертном суждении. Исследователь может предоставить альтернативное обоснование репрезентативности выборки. Основное предположение заключается в том, что исследователь выберет единицы, типичные для генеральной совокупности. Этот метод может быть подвержен предвзятости и субъективному восприятию исследователя.
Judgement Samples: A researcher decides which population members to include in the sample based on his or her judgement. The researcher may provide some alternative justification for the representativeness of the sample. The underlying assumption is that the investigator will select units that are characteristic of the population. This method can be subjected to researcher's biases and perception. Snowball Samples: Often used when a target population is rare. Members of the target population recruit other members of the population for the survey. Quota Samples: The sample is designed to include a designated number of people with certain specified characteristics. For example, 100 coffee drinkers. This type of sampling is common in non probability market research surveys. Convenience Samples: The sample is composed of whatever persons can be most easily accessed to fill out the survey. In non probability samples the relationship between the target population and the survey sample is immeasurable and potential bias is unknowable. Sophisticated users of non probability survey samples tend to view the survey as an experimental condition, rather than a tool for population measurement, and examine the results for internally consistent relationships.
Выборки по снежному кому: Часто используются, когда целевая популяция немногочисленна. Члены целевой популяции привлекают других членов этой популяции для участия в опросе.
Judgement Samples: A researcher decides which population members to include in the sample based on his or her judgement. The researcher may provide some alternative justification for the representativeness of the sample. The underlying assumption is that the investigator will select units that are characteristic of the population. This method can be subjected to researcher's biases and perception. Snowball Samples: Often used when a target population is rare. Members of the target population recruit other members of the population for the survey. Quota Samples: The sample is designed to include a designated number of people with certain specified characteristics. For example, 100 coffee drinkers. This type of sampling is common in non probability market research surveys. Convenience Samples: The sample is composed of whatever persons can be most easily accessed to fill out the survey. In non probability samples the relationship between the target population and the survey sample is immeasurable and potential bias is unknowable. Sophisticated users of non probability survey samples tend to view the survey as an experimental condition, rather than a tool for population measurement, and examine the results for internally consistent relationships.
Квотные выборки: Выборка формируется таким образом, чтобы включить заданное количество людей с определенными характеристиками. Например, 100 любителей кофе. Этот тип выборки распространен в невыборочных маркетинговых исследованиях.
Judgement Samples: A researcher decides which population members to include in the sample based on his or her judgement. The researcher may provide some alternative justification for the representativeness of the sample. The underlying assumption is that the investigator will select units that are characteristic of the population. This method can be subjected to researcher's biases and perception. Snowball Samples: Often used when a target population is rare. Members of the target population recruit other members of the population for the survey. Quota Samples: The sample is designed to include a designated number of people with certain specified characteristics. For example, 100 coffee drinkers. This type of sampling is common in non probability market research surveys. Convenience Samples: The sample is composed of whatever persons can be most easily accessed to fill out the survey. In non probability samples the relationship between the target population and the survey sample is immeasurable and potential bias is unknowable. Sophisticated users of non probability survey samples tend to view the survey as an experimental condition, rather than a tool for population measurement, and examine the results for internally consistent relationships.
Выборки удобства: Выборка состоит из тех лиц, к которым легче всего получить доступ для заполнения анкеты.
Judgement Samples: A researcher decides which population members to include in the sample based on his or her judgement. The researcher may provide some alternative justification for the representativeness of the sample. The underlying assumption is that the investigator will select units that are characteristic of the population. This method can be subjected to researcher's biases and perception. Snowball Samples: Often used when a target population is rare. Members of the target population recruit other members of the population for the survey. Quota Samples: The sample is designed to include a designated number of people with certain specified characteristics. For example, 100 coffee drinkers. This type of sampling is common in non probability market research surveys. Convenience Samples: The sample is composed of whatever persons can be most easily accessed to fill out the survey. In non probability samples the relationship between the target population and the survey sample is immeasurable and potential bias is unknowable. Sophisticated users of non probability survey samples tend to view the survey as an experimental condition, rather than a tool for population measurement, and examine the results for internally consistent relationships.
В невыборочных выборках связь между генеральной совокупностью и выборкой для опроса неизмерима, а потенциальная предвзятость неизвестна. Опытные пользователи невыборочных выборок склонны рассматривать опрос как экспериментальное условие, а не как инструмент для измерения численности генеральной совокупности, и анализировать результаты на предмет внутренней согласованности.
Judgement Samples: A researcher decides which population members to include in the sample based on his or her judgement. The researcher may provide some alternative justification for the representativeness of the sample. The underlying assumption is that the investigator will select units that are characteristic of the population. This method can be subjected to researcher's biases and perception. Snowball Samples: Often used when a target population is rare. Members of the target population recruit other members of the population for the survey. Quota Samples: The sample is designed to include a designated number of people with certain specified characteristics. For example, 100 coffee drinkers. This type of sampling is common in non probability market research surveys. Convenience Samples: The sample is composed of whatever persons can be most easily accessed to fill out the survey. In non probability samples the relationship between the target population and the survey sample is immeasurable and potential bias is unknowable. Sophisticated users of non probability survey samples tend to view the survey as an experimental condition, rather than a tool for population measurement, and examine the results for internally consistent relationships.