1
00:00:01,639 --> 00:00:02,640
Speaker 1: Welcome back everyone.

2
00:00:02,680 --> 00:00:04,919
Speaker 2: Today, we're going to be looking at a lot of

3
00:00:05,759 --> 00:00:08,720
more statistics here in the next few podcasts, and then

4
00:00:08,759 --> 00:00:10,800
we're going to get into legal and ethical issues.

5
00:00:12,199 --> 00:00:14,119
Speaker 1: Sampling is the one we're going to be talking about today.

6
00:00:14,439 --> 00:00:17,600
Speaker 2: Before you analyze anything, you must run any single statistic.

7
00:00:17,679 --> 00:00:19,760
You have to decide who gets included in your study,

8
00:00:20,199 --> 00:00:23,679
and that decision shapes everything. Here's the key idea that

9
00:00:23,719 --> 00:00:26,879
each tripop wants you to hold. Choose the wrong sample

10
00:00:26,920 --> 00:00:29,440
and even perfect statistics. When this leads you, you can

11
00:00:29,519 --> 00:00:31,920
run the most sophisticated regression in the world. But if

12
00:00:31,920 --> 00:00:35,679
your sample doesn't represent the population you care about, your

13
00:00:35,719 --> 00:00:39,679
conclusions are built on sand. So when you're reading a study,

14
00:00:40,399 --> 00:00:43,359
you need to ask yourself who was sampled, how were

15
00:00:43,399 --> 00:00:46,320
they recruited, does this sample reflect the target population? And

16
00:00:46,359 --> 00:00:48,719
are there demographic or cultural groups that limit where we

17
00:00:48,759 --> 00:00:52,880
can conclude that last question. Matt is enormously in clinical

18
00:00:52,920 --> 00:00:56,159
psych where so much foundational research was conducted on narrow

19
00:00:56,920 --> 00:01:02,039
Western educated, industrialized, rich, and democratic ample. If a therapy

20
00:01:02,079 --> 00:01:05,920
works in that population, we cannot assume it generalizes to

21
00:01:05,920 --> 00:01:10,719
an unhoused individual. That bridge from study to real world

22
00:01:10,799 --> 00:01:16,359
application is called external validity. We've covered this before, so

23
00:01:16,560 --> 00:01:20,799
probability samplings. Our first one means every individual the population

24
00:01:20,879 --> 00:01:24,200
has a known and non zero chance of being selected.

25
00:01:25,120 --> 00:01:28,400
This is what supports generalizability, and there are four types

26
00:01:28,439 --> 00:01:31,879
of probability sampling. Simple random, every person in the population

27
00:01:31,959 --> 00:01:33,959
has an equal chance of being chosen.

28
00:01:34,680 --> 00:01:37,200
Speaker 1: Think of it like drawing names out of a hat.

29
00:01:37,280 --> 00:01:40,079
Speaker 2: You need a complete population list to do this, which

30
00:01:40,120 --> 00:01:42,920
is often impractical for large populations, but when you can

31
00:01:42,920 --> 00:01:46,000
pull it off, it's the gold standard. Imagine you're studying

32
00:01:46,000 --> 00:01:49,599
depression rates among the licensed psychologists. If you have access

33
00:01:49,599 --> 00:01:52,599
to the full board list and randomly seared names from it,

34
00:01:52,719 --> 00:01:57,799
that's simple random sampling every licensed psychologist an equal shot.

35
00:01:58,920 --> 00:02:00,560
Speaker 1: Stratified divide the.

36
00:02:00,519 --> 00:02:05,040
Speaker 2: Population into subgroups called strata, things like gender, ethnicity, or

37
00:02:05,079 --> 00:02:07,920
diagnostic category, and then randomly sample.

38
00:02:07,680 --> 00:02:08,439
Speaker 1: Within each group.

39
00:02:09,039 --> 00:02:14,280
Speaker 2: The goal is ensuring proportional representation across key demographics. So

40
00:02:14,439 --> 00:02:17,520
you're studying CBT outcomes across racial groups. If you just

41
00:02:17,639 --> 00:02:19,840
randomly sampled from a clinic, you might end up with

42
00:02:19,879 --> 00:02:23,479
a sample it's eighty five percent white because that reflects

43
00:02:23,560 --> 00:02:30,319
clinic demographics. Stratified sampling insures you have adequate representation from Black, Latino,

44
00:02:30,360 --> 00:02:35,240
and Asian clients, making your finds more clinically meaningful and generalizable.

45
00:02:36,280 --> 00:02:40,599
Cluster sampling, instead of sampling individuals, you randomly select entire

46
00:02:40,639 --> 00:02:45,719
groups like schools, hospitals, or neighborhoods, and then same individuals

47
00:02:45,719 --> 00:02:52,039
within those clusters. This is efficient for geographically spread populations,

48
00:02:52,080 --> 00:02:54,960
but it produces more sampling error than stratified sampling because

49
00:02:54,960 --> 00:02:58,639
of group level variants, people within the same cluster tend

50
00:02:58,719 --> 00:03:01,159
to be more similar to each other the broader population.

51
00:03:02,840 --> 00:03:03,280
Speaker 1: You want to.

52
00:03:03,240 --> 00:03:07,159
Speaker 2: Study anxiety prevalence in community mental health centers across a

53
00:03:07,240 --> 00:03:10,120
large state, rather than sampling every center, you randomly select

54
00:03:10,120 --> 00:03:13,280
ten and access all clients within them.

55
00:03:13,719 --> 00:03:14,879
Speaker 1: That's cluster sampling.

56
00:03:16,360 --> 00:03:21,960
Speaker 2: Systematic sampling you select every individual from a list after

57
00:03:22,039 --> 00:03:25,520
a random starting point. So if you look at if

58
00:03:25,599 --> 00:03:28,280
K equals ten, you start at a random number, say

59
00:03:28,360 --> 00:03:31,080
number four, then take number fourteen, twenty four, thirty four,

60
00:03:31,080 --> 00:03:33,960
and so forth. It's easy to implement, but it carries

61
00:03:34,000 --> 00:03:37,800
one specific risk, what they call periodicity bias. If there

62
00:03:37,879 --> 00:03:39,840
is a repeating pattern in the list that aligns with

63
00:03:39,919 --> 00:03:43,159
your sampling interval, you wire systematically over under select certain

64
00:03:43,199 --> 00:03:46,599
types of people. Example, would be you have a list

65
00:03:46,639 --> 00:03:49,319
of therapy and take appointments organized by day of week.

66
00:03:51,240 --> 00:03:54,039
If you take every seventh appointment and the list cycles weekly,

67
00:03:54,080 --> 00:03:56,319
you might only ever select clients who come in on Mondays,

68
00:03:56,360 --> 00:03:59,360
potentially missing weekend to only available clients who may differ

69
00:03:59,360 --> 00:04:04,520
in employments, status, or severity. Another one now is non

70
00:04:04,599 --> 00:04:09,639
probability sampling. We looked at probability sampling methods that included

71
00:04:09,680 --> 00:04:14,719
systematic sampling, cluster sampling, stratified, and simple random. Now we

72
00:04:14,759 --> 00:04:17,560
look at non probability and it's used when probability sampling

73
00:04:17,600 --> 00:04:20,759
isn't feasible. The trade off is that generalizability is limited.

74
00:04:21,279 --> 00:04:24,439
You have to know that's going into your interpretation. The

75
00:04:24,439 --> 00:04:28,240
first one is convenience. Sampling participants are selected based on

76
00:04:28,240 --> 00:04:32,079
availability or proximity. Fast, low cost, easy to do, and

77
00:04:32,160 --> 00:04:36,920
highly vulnerable to selection bias. The classic examples recruiting undergrad

78
00:04:36,920 --> 00:04:39,519
psych students because they're sitting in the building common in

79
00:04:39,560 --> 00:04:41,800
academic research, but you have to be cautious about what

80
00:04:41,920 --> 00:04:45,839
you can claim from those findings. Clinical example, a researcher

81
00:04:45,879 --> 00:04:49,319
studies the effect of mindfulness on stress by recruiting from

82
00:04:49,319 --> 00:04:53,120
a university psychology department waiting list. These are people who

83
00:04:53,160 --> 00:04:56,439
already sought help, are likely college educated, are probably not

84
00:04:56,519 --> 00:05:01,279
representative of community health mental health populations. The researcher deliberately

85
00:05:01,319 --> 00:05:05,160
selects participants who meets specific criteria. This is especially useful

86
00:05:05,160 --> 00:05:09,600
in qualitative research. This is purpose of sampling. So again,

87
00:05:09,680 --> 00:05:13,040
the researcher deliberately selects participants who meets specific criteria. This

88
00:05:13,120 --> 00:05:17,160
is especially useful in qualitative research. For instance, you're studying

89
00:05:17,160 --> 00:05:19,680
the lived experience of psychologist who have treated clients with

90
00:05:19,759 --> 00:05:27,399
Capgras syndrome, so you purposely selected or recruited clinicians we

91
00:05:27,439 --> 00:05:33,439
have direct treatment experience. Snowball sampling, existing participants recruit future

92
00:05:33,439 --> 00:05:35,800
participants from their social networks. This is used when you're

93
00:05:35,800 --> 00:05:40,240
studying hidden or stigmatized populations where a sampling frame doesn't exist.

94
00:05:40,279 --> 00:05:45,120
It's hard to get individuals like distress and undocumented immigrants

95
00:05:45,120 --> 00:05:47,680
who are unlikely to respond to traditional recruitment. You start

96
00:05:47,720 --> 00:05:50,800
with two or three participants through a community organization and

97
00:05:50,839 --> 00:05:52,600
ask them to refer other as they trust, and then

98
00:05:52,639 --> 00:05:53,160
it grows.

99
00:05:54,160 --> 00:05:56,439
Speaker 1: Lastly, it's quota sampling. These are like.

100
00:05:56,439 --> 00:06:00,720
Speaker 2: Stratified sampling and structure, but non random and execution. You

101
00:06:00,800 --> 00:06:04,519
identify categories you want represented and feel predetermined quotas for

102
00:06:04,600 --> 00:06:07,560
each is fast and structured, but it lacks true randomness,

103
00:06:08,079 --> 00:06:11,920
so you can't apply probability based statistical inferences.

104
00:06:11,399 --> 00:06:12,680
Speaker 1: The same way.

105
00:06:13,480 --> 00:06:18,000
Speaker 2: Another area sampling distributions, the central limit theorem and standard error.

106
00:06:20,120 --> 00:06:23,040
Even good samples vary from the population, there's always some

107
00:06:23,160 --> 00:06:25,639
degree of sampling error. But here's the insight. If we

108
00:06:25,720 --> 00:06:29,839
understand how samples behave across many repetitions, we can estimate

109
00:06:29,879 --> 00:06:32,839
population parameters from a single sample. That's the whole game.

110
00:06:33,600 --> 00:06:36,560
This is the theoretical distribution of a statistics, say the

111
00:06:36,639 --> 00:06:40,319
mean across many random samples drawn from the same population.

112
00:06:40,439 --> 00:06:42,680
You're not looking at one sample. You're imagining what the

113
00:06:42,720 --> 00:06:44,959
distribution of that statistic would look like if you do

114
00:06:45,040 --> 00:06:47,839
samples over and over again. This is one of the

115
00:06:47,839 --> 00:06:50,399
most important theorems in all of statistics, and the e

116
00:06:50,480 --> 00:06:54,199
triple p tested directly. Here's what it says for the

117
00:06:54,240 --> 00:06:57,439
central limit theorem. With a large enough sample, typically thirty

118
00:06:57,519 --> 00:07:01,720
or more, the sampling distribution of the mean becomes approximately normal,

119
00:07:01,920 --> 00:07:06,639
regardless of the shape of the population distribution. Why doesn't

120
00:07:06,639 --> 00:07:10,519
matter because it justifies using what we call parametric tests

121
00:07:10,560 --> 00:07:13,720
even when your underlying population isn't normally distributed.

122
00:07:14,360 --> 00:07:15,480
Speaker 1: If you're studying.

123
00:07:15,079 --> 00:07:18,360
Speaker 2: Trauma scores in a community sample and a distribution is scored,

124
00:07:18,800 --> 00:07:21,040
you can still use a T test or a NOVA

125
00:07:21,519 --> 00:07:24,399
if your sample is large enough, because the central limit

126
00:07:24,480 --> 00:07:26,560
theorem tells you the sampling.

127
00:07:26,199 --> 00:07:32,240
Speaker 1: Distribution of the mean will be normal. Standard error this.

128
00:07:32,120 --> 00:07:37,600
Speaker 2: Is the standard deviation of the sampling distribution. It tells

129
00:07:37,639 --> 00:07:40,240
you how much sample means vary from sample to sample.

130
00:07:40,279 --> 00:07:43,480
The formula's se equals the standard deviation divided by the

131
00:07:43,480 --> 00:07:47,360
square root of N. An example would be a sample

132
00:07:47,439 --> 00:07:51,240
sized increases standard error decreases, your estimate of the population

133
00:07:51,959 --> 00:07:55,360
mean becomes more precise. This is why large critical trials

134
00:07:55,360 --> 00:07:58,800
are more trustworthy than small pilot studies, not because the

135
00:07:58,839 --> 00:08:02,160
researchers are smarter, because the math of sampling gives them

136
00:08:02,199 --> 00:08:07,240
a more stable estimate. There's five factors that affect sampling size.

137
00:08:08,519 --> 00:08:10,839
First is the effects size. This is the magnitude of

138
00:08:10,839 --> 00:08:15,800
the difference or relationship you expect to find. Cohen's benchmarks

139
00:08:15,800 --> 00:08:18,480
are d equals zero point two zero for small, zero

140
00:08:18,519 --> 00:08:21,199
point five for medium, and zero point eight for large.

141
00:08:21,680 --> 00:08:26,680
Smaller effects require larger samples to detect. If you're studying

142
00:08:27,240 --> 00:08:31,000
a subtle early intervention effect on subclinical anxiety, you need

143
00:08:31,120 --> 00:08:33,159
more participants than if you're studying the effect of a

144
00:08:33,159 --> 00:08:42,120
major trauma event on PTSD. Severity power is the probability

145
00:08:42,159 --> 00:08:46,559
of detecting a real effect when one truly exists. The

146
00:08:46,600 --> 00:08:49,639
conventional target is zero point eight, meaning on eighty percent

147
00:08:49,720 --> 00:08:53,240
chance of detecting a true effect. Power increases with larger

148
00:08:53,279 --> 00:08:57,399
sample size, larger effects size, and higher alpha on the

149
00:08:57,519 --> 00:09:00,320
h If a study fails to find a significant result

150
00:09:00,399 --> 00:09:03,480
with low power, you should be suspicious. That might be

151
00:09:03,519 --> 00:09:06,120
a type two error false negative, not a true null

152
00:09:06,240 --> 00:09:12,639
result type one error false positive. This is the threshold

153
00:09:12,720 --> 00:09:16,279
you used to decide whether it's a result is statistically significant,

154
00:09:16,360 --> 00:09:19,879
also called the alpha level. Alpha level type one error

155
00:09:19,960 --> 00:09:22,960
rate typically set at zero point zero five. If you

156
00:09:23,039 --> 00:09:25,279
lower alpha to zero point zero one to be more

157
00:09:25,279 --> 00:09:27,840
conservative to reduce the risk of false positives, you need

158
00:09:27,840 --> 00:09:31,159
a much larger sample to maintain your power. There's always

159
00:09:31,200 --> 00:09:35,000
going to be a trade off design complexity. More groups

160
00:09:35,000 --> 00:09:37,279
and more predictors mean you need a larger sample. A

161
00:09:37,360 --> 00:09:41,639
study comparing five treatment conditions requires more participants than a

162
00:09:41,679 --> 00:09:46,679
two group design. Within subject designs where the same person

163
00:09:46,720 --> 00:09:51,120
is measured multiple times, typically fewer participants are needed because

164
00:09:51,159 --> 00:09:54,159
you reduce vary variants by using each person on their

165
00:09:54,200 --> 00:09:59,120
own control. There are also constraints budget, time, and access

166
00:09:59,200 --> 00:10:03,320
to participants often limit ideal sample sizes. The researcher studying

167
00:10:03,320 --> 00:10:05,639
a rare personality sort of and I'm able to recruit

168
00:10:06,639 --> 00:10:09,120
may not be able to recruit three hundred participants, no

169
00:10:09,120 --> 00:10:15,519
matter how clean the design is. As section six is

170
00:10:15,600 --> 00:10:18,360
talk about power analysis, which is the form of process

171
00:10:19,720 --> 00:10:22,360
used to calculate the minimum sample size and needed to

172
00:10:22,399 --> 00:10:29,480
detect an effect, and you run it before data collection.

173
00:10:29,519 --> 00:10:32,919
It's a planning tool, not an afterthought. Effect size using

174
00:10:32,960 --> 00:10:36,639
Cohen's d f orr. Depending on the design, alpha level

175
00:10:36,720 --> 00:10:39,799
desired power usually set at zero point eight and a

176
00:10:39,879 --> 00:10:44,039
number of groups of predictors. The standard software used in

177
00:10:44,080 --> 00:10:47,879
psychology's g power, which is free and widely used. Failing

178
00:10:47,919 --> 00:10:51,200
to conduct the power analysis before a study creates two problems. First,

179
00:10:51,200 --> 00:10:53,440
you might end up underpowered and miss a real effect

180
00:10:54,039 --> 00:10:56,879
that's a false negative type two error. Second, you might

181
00:10:56,919 --> 00:11:00,600
over recruit, which waste resources and raises unnecessary ethical concerns

182
00:11:01,120 --> 00:11:05,159
about participant burden. For instance, the researcher wants to study

183
00:11:05,159 --> 00:11:10,039
whether trauma focus CBT reduces PTSD symptoms more than supportive counseling.

184
00:11:10,480 --> 00:11:13,279
They expect the medium effects size of D equals zero

185
00:11:13,320 --> 00:11:17,039
point five. They set alpha zero point five and zero

186
00:11:17,600 --> 00:11:20,120
point zero five and one power of zero point eight.

187
00:11:20,799 --> 00:11:23,480
Running g power tells them they need approximately sixty four

188
00:11:23,559 --> 00:11:27,279
participants per group, so one hundred and twenty eight total.

189
00:11:27,679 --> 00:11:31,240
If they only recruit forty, they're severely underpowered, and a

190
00:11:31,320 --> 00:11:36,159
null result tells us almost nothing. So let's talk about

191
00:11:36,159 --> 00:11:39,279
what goes wrong. Sampling bias is introduced when some members

192
00:11:39,320 --> 00:11:42,159
of the population are systematically less likely to be included

193
00:11:42,159 --> 00:11:46,919
in the sample. The key word is systematically self selection bias.

194
00:11:47,039 --> 00:11:51,519
Volunteers differ from non volunteers in ways that are often

195
00:11:51,559 --> 00:11:52,639
clinically relevant.

196
00:11:53,039 --> 00:11:54,279
Speaker 1: People who volunteer for.

197
00:11:54,279 --> 00:11:57,159
Speaker 2: Depression treatment studies tend to be more motivated, more resourced,

198
00:11:57,200 --> 00:12:02,120
and less severely impaired than the average with depression. If

199
00:12:02,159 --> 00:12:04,600
your results are based entirely on volunteers, you may not

200
00:12:04,639 --> 00:12:10,080
be systematically overestimating, or you may be systematically overestimating treatment effectiveness.

201
00:12:11,360 --> 00:12:13,360
Speaker 1: Non response by certain groups.

202
00:12:13,120 --> 00:12:16,519
Speaker 2: Are less likely to complete surveys or follow ups if

203
00:12:16,559 --> 00:12:18,840
clients with the most severe pathology drop out of your

204
00:12:18,840 --> 00:12:22,000
longitudinal study because they're too disregulated to complete it. Your

205
00:12:22,080 --> 00:12:28,120
data on chronic compairment is compromised. Undercoverage, certain populations are

206
00:12:28,200 --> 00:12:32,840
left out entirely. Classic example, unhoused individuals. If you're studying

207
00:12:32,840 --> 00:12:35,919
mental health outcomes in a city and your sampling frame

208
00:12:36,039 --> 00:12:39,320
is a clinic registry, you've already excluded the population most

209
00:12:39,399 --> 00:12:44,279
likely to have severe untreated comobid conditions. Selection effects are

210
00:12:44,399 --> 00:12:47,200
related but distinct concept. They occur when the process of

211
00:12:47,240 --> 00:12:51,440
assigning participants to conditions or groups is not random. This

212
00:12:51,559 --> 00:12:55,559
affects internal validity, your ability to claim that the independent variable,

213
00:12:55,720 --> 00:12:59,320
not some pre existing difference between groups.

214
00:12:59,159 --> 00:13:00,360
Speaker 1: Cause the outcome.

215
00:13:03,240 --> 00:13:06,879
Speaker 2: Selection effects are why randomized control trials are the gold standard.

216
00:13:07,480 --> 00:13:11,200
Random assignment breaks the link between pre existing characteristics and

217
00:13:11,240 --> 00:13:16,440
group measurement membership. Poor sampling equals limited to not generalizability

218
00:13:16,519 --> 00:13:21,120
external validity, which is the degree to which fundings from

219
00:13:21,120 --> 00:13:23,519
a study can have generalize to other people. Sampling is

220
00:13:23,519 --> 00:13:28,799
the primary drive driver of external validity. If your sample

221
00:13:28,840 --> 00:13:32,039
is narrow, your conclusions are narrow. When reading research like

222
00:13:32,080 --> 00:13:35,600
a scientist, run through the checklist. Was the sampling method

223
00:13:35,639 --> 00:13:41,039
is described. Was randomization used? Was the sample diverse and representative?

224
00:13:41,120 --> 00:13:43,840
How large was it it was? A power analysis reported

225
00:13:44,440 --> 00:13:47,720
whether high attrition rates are missing data. These are not

226
00:13:47,799 --> 00:13:50,799
just academic questions and clinical practice that determine whether the

227
00:13:50,799 --> 00:13:54,159
intervention you're considering for your client was actually tested on

228
00:13:54,240 --> 00:13:55,799
someone like your client.

229
00:13:58,120 --> 00:13:59,320
Speaker 1: That's it for now, folks,

