1
00:00:00,840 --> 00:00:02,000
Speaker 1: Welcome back everybody.

2
00:00:02,040 --> 00:00:05,200
Speaker 2: Today we're going to be looking at validity concepts and

3
00:00:05,280 --> 00:00:10,320
psychological research. We're taking a deeper dive into research and

4
00:00:10,439 --> 00:00:14,800
we'll be looking at validity. Some of them must know definitions.

5
00:00:14,839 --> 00:00:18,679
Internal validity asked whether the independent variable truly caused the

6
00:00:18,800 --> 00:00:23,440
change and the dependent variable. So, if you remember research

7
00:00:23,559 --> 00:00:26,440
we are looking for here, is there really the independent variable?

8
00:00:26,440 --> 00:00:30,039
Do we have other confounding variables playing a role? External

9
00:00:30,079 --> 00:00:33,479
validity asks whether findings can be generalized beyond the study

10
00:00:33,520 --> 00:00:37,719
to other people, other settings and times, So between lab

11
00:00:37,759 --> 00:00:42,079
work compared to the real world. Construct validity asked whether

12
00:00:42,119 --> 00:00:45,880
we actually measured what we intended to measure, and we'll look.

13
00:00:45,759 --> 00:00:46,880
Speaker 1: At these deeper in a minute.

14
00:00:46,920 --> 00:00:50,799
Speaker 2: Statistical conclusion validity ask whether the relationships between variables.

15
00:00:50,479 --> 00:00:51,600
Speaker 1: Are accurate and reliable.

16
00:00:52,280 --> 00:00:56,119
Speaker 2: Threats to ability undermine the trustworthiness of conclusions and must

17
00:00:56,159 --> 00:01:01,840
be addressed in study design, interpretation, and application. Strong research

18
00:01:01,880 --> 00:01:06,640
design uses tools like randomization, manipulation checks, multi method approaches,

19
00:01:06,640 --> 00:01:11,040
and sampling strategies. So back to internal validity, the backbone

20
00:01:11,040 --> 00:01:15,000
of causal inference. The core question can be can we

21
00:01:15,120 --> 00:01:17,519
be confident the independent of aerial cause a change in

22
00:01:17,560 --> 00:01:21,400
the dependent variable. Common threats to internal validity include to history.

23
00:01:21,480 --> 00:01:25,920
For instance, external events occurring during the study that influence

24
00:01:26,000 --> 00:01:30,239
results independently of the intervention. It's an example of history.

25
00:01:31,040 --> 00:01:33,280
Here are another example of it. A therapist is studying

26
00:01:33,319 --> 00:01:35,400
the effect of a UCBT.

27
00:01:34,799 --> 00:01:36,079
Speaker 1: Protocol on depression.

28
00:01:36,159 --> 00:01:39,200
Speaker 2: Halfway through, a major natural disaster occurs in the region,

29
00:01:39,719 --> 00:01:43,840
and participants mood scores worse than regardless of treatment. The disaster,

30
00:01:44,439 --> 00:01:48,159
not the treatment, is the driving change. Another factor could

31
00:01:48,159 --> 00:01:51,719
be maturation, the natural development of or biological changes and

32
00:01:51,760 --> 00:01:56,560
participants that occur over time independent of the treatment. Clinical

33
00:01:56,599 --> 00:01:59,920
examples A school psychologists tests of reading intervention was second greater.

34
00:02:00,120 --> 00:02:02,599
Over eight months, kids improve and much of the game

35
00:02:02,719 --> 00:02:06,680
reflects normal reading development that would have occurred anyway, not

36
00:02:06,840 --> 00:02:11,280
the program itself. Another one it's testing effects. Exposure to

37
00:02:11,319 --> 00:02:16,360
a pretest influences performance on the post test independent of

38
00:02:16,439 --> 00:02:20,879
the intervention. So a researcher administers the BDI before and

39
00:02:20,960 --> 00:02:24,919
after an eight week mindfulness program. Participants score lower on

40
00:02:24,960 --> 00:02:27,240
the post test, partly because they have already seen the

41
00:02:27,319 --> 00:02:30,479
questions and know how to respond, not solely because of

42
00:02:30,520 --> 00:02:37,479
mindfulness practice instrumentation changes in measurement tools raiders for scoring

43
00:02:37,520 --> 00:02:40,919
criteria over the course of a study. Clinical example, in

44
00:02:40,919 --> 00:02:44,599
a longitudinal outcome studies, supervisors rating patient progress become more

45
00:02:44,680 --> 00:02:48,759
lenient in their scoring after several months of familiarity with participants,

46
00:02:48,840 --> 00:02:54,560
producing artificially inflated improvement scores. Another one in statistical regression

47
00:02:54,919 --> 00:02:57,400
definition and participants are selected because.

48
00:02:57,199 --> 00:02:58,280
Speaker 1: Of extreme scores.

49
00:02:58,840 --> 00:03:02,840
Speaker 2: Scores naturally drift toward the mean or retesting regardless of intervention.

50
00:03:04,240 --> 00:03:06,960
The clinician recruits clients scoring in the severe range on

51
00:03:07,000 --> 00:03:09,560
an anxiety measure for a new exposure therapy trial.

52
00:03:10,400 --> 00:03:11,960
Speaker 1: At protests post.

53
00:03:11,840 --> 00:03:16,599
Speaker 2: Tests, many show improvement simply because extreme core scores statistically

54
00:03:16,639 --> 00:03:21,000
regress toward average, not necessarily because of treatment efficacy.

55
00:03:21,080 --> 00:03:22,360
Speaker 1: You can see this in athletes too.

56
00:03:22,439 --> 00:03:25,439
Speaker 2: Just to make it if you're inclined to watching sports.

57
00:03:25,479 --> 00:03:28,599
If you look at a basketball player and they average

58
00:03:29,120 --> 00:03:31,240
thirty five thirty points a game, but one game they

59
00:03:31,319 --> 00:03:34,039
hit fifty, more than likely there's going to be games

60
00:03:34,080 --> 00:03:37,080
in the future that they're going to score under thirty and.

61
00:03:37,039 --> 00:03:38,000
Speaker 1: The average will bounce.

62
00:03:38,280 --> 00:03:40,280
Speaker 2: The average will drop right into that thirty range over

63
00:03:40,319 --> 00:03:41,759
a course of a series of games.

64
00:03:43,639 --> 00:03:45,120
Speaker 1: Selection bias is another one.

65
00:03:45,280 --> 00:03:49,400
Speaker 2: Non random group assignment produces pre existing differences between groups

66
00:03:49,400 --> 00:03:53,000
that confound results. An example would be a hospital assigns

67
00:03:53,159 --> 00:03:56,759
patients who voluntarily request therapy to the treatment group and

68
00:03:56,840 --> 00:03:59,000
those on a wait list to the control group. The

69
00:03:59,039 --> 00:04:02,840
treatment grouping are ready be more motivated, making any outcome

70
00:04:02,840 --> 00:04:04,680
difference attribute to motivation.

71
00:04:04,639 --> 00:04:05,879
Speaker 1: Rather than the therapy itself.

72
00:04:07,639 --> 00:04:11,639
Speaker 2: Attrition definition is systematic drop out of participants that skews

73
00:04:11,680 --> 00:04:15,039
the remaining sample and distorted results. An example would be

74
00:04:15,039 --> 00:04:17,800
a twelve month substance abuse treatment study. The most severe

75
00:04:17,879 --> 00:04:21,879
users drop out early. The remaining completers show high success rates,

76
00:04:21,879 --> 00:04:26,399
but the findings reflect survivor bias but not true treatment effectiveness.

77
00:04:27,839 --> 00:04:31,040
Another one is a diffusion of treatment, where control group

78
00:04:31,120 --> 00:04:34,879
participants become exposed in the intervention, contaminating the comparison. So

79
00:04:35,600 --> 00:04:38,600
in a DBT's skill study, participants in the control group

80
00:04:38,600 --> 00:04:41,920
begin learning distress tolerance techniques from friends in the treatment group,

81
00:04:42,519 --> 00:04:47,600
reducing the observable difference between the conditions. Strategies of strengthen

82
00:04:48,079 --> 00:04:53,199
internal validity or random assignment. It distributes participant characteristics evenly

83
00:04:53,240 --> 00:04:57,439
across groups, eliminating selection bias at the outset. Control groups

84
00:04:57,439 --> 00:05:00,720
provide a comparison baseline so that orth changes can be

85
00:05:00,759 --> 00:05:05,319
attributed to the independent variable rather than extraneous variables. Consistent

86
00:05:05,360 --> 00:05:10,240
measurement protocols reduce instrumentation drift by standardizing how when data

87
00:05:10,279 --> 00:05:14,959
is collected across old time points. Blinding raiders or participants

88
00:05:15,000 --> 00:05:19,439
prevents expectancy effects from influencing data collection or participant behavior.

89
00:05:20,800 --> 00:05:24,439
Pretesting to assess baseline equivalence confirms that groups start from

90
00:05:24,439 --> 00:05:29,879
the same point before the intervention begins. External validity can

91
00:05:29,920 --> 00:05:31,439
we generalize these fightings through.

92
00:05:31,319 --> 00:05:33,759
Speaker 1: The real world or other people or other settings.

93
00:05:34,319 --> 00:05:37,639
Speaker 2: Some of the major threats to external validity include population validity,

94
00:05:37,759 --> 00:05:41,360
the degree to which findings generalize from the study sample

95
00:05:41,399 --> 00:05:46,439
to the broader intended population. The landmark study on social

96
00:05:46,439 --> 00:05:49,839
anxiety treatment was conducted entirely with white college aged women

97
00:05:50,279 --> 00:05:53,680
recruited from US a university psychology pool. Applying those findings

98
00:05:53,680 --> 00:05:59,000
to middle aged Latino men settings stretches a sample beyond

99
00:06:00,079 --> 00:06:01,279
on its original boundaries.

100
00:06:04,240 --> 00:06:06,319
Speaker 1: Another one is ecological validity.

101
00:06:06,279 --> 00:06:08,680
Speaker 2: The degree to which the study environment mirrors the real

102
00:06:08,720 --> 00:06:13,959
world conditions. For instance, a memory intervention tested in a quiet,

103
00:06:14,000 --> 00:06:16,839
well lit lab with no distractions may show strong effects,

104
00:06:16,879 --> 00:06:19,480
but when applied and an assisted living facility with ambient

105
00:06:19,519 --> 00:06:21,079
noise and caregiver interruptions.

106
00:06:21,720 --> 00:06:23,759
Speaker 1: The benefits diminish substantially.

107
00:06:25,040 --> 00:06:28,120
Speaker 2: Temporal validity the degree to which findings hold across different

108
00:06:28,160 --> 00:06:31,759
historical time periods and cultural moments, For instance, and.

109
00:06:31,800 --> 00:06:32,720
Speaker 1: Stress and oculation.

110
00:06:32,800 --> 00:06:36,000
Speaker 2: Protocol validated during a period of economic stability may not

111
00:06:36,079 --> 00:06:40,720
generalize the populations experiencing a pandemic, mass employment, or a

112
00:06:40,759 --> 00:06:41,759
political crisis.

113
00:06:42,600 --> 00:06:44,319
Speaker 1: Strengthening the external validity.

114
00:06:44,240 --> 00:06:49,040
Speaker 2: Well representative sampling ensures the study sample reflects the demographics

115
00:06:49,519 --> 00:06:54,199
and characteristics of the target population. Also, realistic settings or

116
00:06:54,240 --> 00:06:56,240
field studies move research out of labs and into the

117
00:06:56,279 --> 00:07:02,639
environments where findings will actually be applied. Replication across samples

118
00:07:02,639 --> 00:07:06,480
and settings builds confidence that an effect is robust rather

119
00:07:06,519 --> 00:07:12,240
than an artifact of one particular context. Constructibilities Are we

120
00:07:12,279 --> 00:07:14,639
measuring the theoretical construct who claim to be measuring?

121
00:07:15,240 --> 00:07:15,360
Speaker 1: It?

122
00:07:15,399 --> 00:07:17,560
Speaker 2: Is about the fit between concept and method. If you

123
00:07:17,560 --> 00:07:20,360
claim to measure anxiety, are you capturing anxiety?

124
00:07:20,920 --> 00:07:22,439
Speaker 1: Or are you capturing general.

125
00:07:22,120 --> 00:07:27,560
Speaker 2: Distress, social discomfort, or physiological arousal. Here are some threats

126
00:07:27,560 --> 00:07:31,680
to construct validity monomethod bias. Relying on only one measurement

127
00:07:31,720 --> 00:07:34,959
method such a self report alone limits the scope of

128
00:07:34,959 --> 00:07:40,720
what is captured and introduces shared method variance that inflates correlations.

129
00:07:41,240 --> 00:07:44,240
A researcher assesses both depression severity and quality of life

130
00:07:44,360 --> 00:07:48,040
using only self report questionnaires administered in the same session.

131
00:07:48,240 --> 00:07:52,480
Any correlation between scores may partially reflect the participant's general

132
00:07:52,600 --> 00:07:54,920
negative mood in the state, not a true relationship between

133
00:07:54,920 --> 00:08:00,480
the constructs. Inadequate operationalization, poorly defined, or mismat which measures

134
00:08:00,519 --> 00:08:03,720
that do not adequately represent the theoretical construct being studied.

135
00:08:04,639 --> 00:08:07,160
The researcher claims to measure empathy, but uses only a

136
00:08:07,160 --> 00:08:11,959
cognitive perspective taking scale. Affect of empathy that felt resonance

137
00:08:12,000 --> 00:08:18,040
with another's emotions goes entirely unmeasured, leaving the operationalization incomplete.

138
00:08:18,759 --> 00:08:26,120
Experimental expectancy, Subtle behavioral cues from researchers unconsciously influenced participant responses,

139
00:08:26,160 --> 00:08:30,079
contaminating the measurement. A therapist researcher who believes a new

140
00:08:30,120 --> 00:08:33,559
intervention will work may ask questions in a warmer, more

141
00:08:33,600 --> 00:08:38,879
affirming tone during treatment sessions, inadvertently shaping participant responses in

142
00:08:38,919 --> 00:08:44,879
the direction of the hypothesis. Hypothesis Guessing participants offer and

143
00:08:45,960 --> 00:08:49,679
sorry explain that participants infer the study's purpose and alter

144
00:08:49,799 --> 00:08:55,919
their behavior to conform to perceived expectations, producing demand characteristics. Example,

145
00:08:56,879 --> 00:08:59,919
participants in a mindfulness study figure out that researchers expected

146
00:09:00,120 --> 00:09:04,679
to report reduce stress. They minimize stress complaints on questionnaires,

147
00:09:04,720 --> 00:09:09,639
even when their actual experience has not changed meaningfully. Confounding constructs,

148
00:09:09,759 --> 00:09:12,759
overlapping variables are not separated from one another, making an

149
00:09:12,759 --> 00:09:14,879
impossible determine which construct.

150
00:09:14,399 --> 00:09:15,639
Speaker 1: Is actually driving the effect.

151
00:09:16,559 --> 00:09:18,639
Speaker 2: An example would be a study claims to demonstrate the

152
00:09:18,679 --> 00:09:23,159
effects of self esteem on academic performance, but the measure

153
00:09:23,200 --> 00:09:26,639
of self esteem is heavily saturated with general positive affect

154
00:09:26,799 --> 00:09:28,759
it becomes unclear whether self esteem or.

155
00:09:28,679 --> 00:09:30,679
Speaker 1: Mood is the active ingredient.

156
00:09:31,720 --> 00:09:34,799
Speaker 2: So how do we improve construct validity? Multi method assessments

157
00:09:34,799 --> 00:09:39,919
combined self report, physiological measures, and behavioral observation to triangulate

158
00:09:39,919 --> 00:09:41,480
a construct from multiple angles.

159
00:09:42,120 --> 00:09:44,360
Speaker 1: Right, you want to understand this from multiple sides.

160
00:09:45,440 --> 00:09:49,879
Speaker 2: Manipulation of checks confirm that the independent variable actually produce

161
00:09:49,960 --> 00:09:53,120
the intended psychological state before examining.

162
00:09:52,639 --> 00:09:53,799
Speaker 1: Its downstream effects.

163
00:09:54,559 --> 00:09:57,840
Speaker 2: Pilot testing measures for clarity and precision catches ambiguous or

164
00:09:57,879 --> 00:10:03,279
double barreled items before the full study launches. Established a

165
00:10:03,399 --> 00:10:07,080
validated scales carry a history or psychometric evidence supporting their

166
00:10:07,080 --> 00:10:11,879
ability to measure the intended construct. Another idea is statistical

167
00:10:11,919 --> 00:10:16,000
conclusion validity. The question are the statistical conclusions accurate to

168
00:10:16,080 --> 00:10:19,039
the data analysis, preserve a real relationship, or miss one

169
00:10:19,080 --> 00:10:23,120
that existed. Some of the threats are low statistical power.

170
00:10:23,240 --> 00:10:26,799
An underpowered study lacks the sensitivity to detect the real effect,

171
00:10:26,960 --> 00:10:28,519
increasing the risk of type.

172
00:10:28,279 --> 00:10:30,519
Speaker 1: Two error false negatives.

173
00:10:30,759 --> 00:10:34,320
Speaker 2: A pilot study comparing two therapy modalities uses only twelve

174
00:10:34,360 --> 00:10:38,720
participants per group. A genuine medium sized treatment difference exists,

175
00:10:38,720 --> 00:10:41,320
but fails to reach significance because the sample size is

176
00:10:41,320 --> 00:10:46,679
too small to detected reliably. Speaking of that, unreliable measures

177
00:10:46,679 --> 00:10:50,240
is another one. Measures with poor reliability introduce random noise

178
00:10:50,279 --> 00:10:54,440
that obscures true relationships between variables. For example, a researcher

179
00:10:54,559 --> 00:10:58,720
uses a hastily assembled depression scale of poor internal consistency.

180
00:10:59,519 --> 00:11:02,440
The noise is in the measure swamps the signal, producing

181
00:11:02,440 --> 00:11:06,480
a non significant result even though the treatment genuinely reduced

182
00:11:06,960 --> 00:11:13,519
depressive symptoms. How about violation of assumptions applying parametric statistical

183
00:11:13,559 --> 00:11:16,200
test to data that does not meet the required distributional

184
00:11:16,279 --> 00:11:17,200
or scaling assumptions.

185
00:11:17,279 --> 00:11:18,720
Speaker 1: Let me give you an example of this one.

186
00:11:19,000 --> 00:11:22,840
Speaker 2: A therapists researcher runs a standard A NOVA, so you're

187
00:11:22,840 --> 00:11:26,279
looking at multiple independent variables on a heavily skewed trauma

188
00:11:26,360 --> 00:11:32,039
symptom scale without transformation, producing misleading f statistics and potentially

189
00:11:32,080 --> 00:11:38,600
false conclusions about group differences. So let me try to

190
00:11:38,639 --> 00:11:40,559
explain this a little bit more simple because it kind

191
00:11:40,559 --> 00:11:44,279
of gets into the weeds when it comes to parametric

192
00:11:44,360 --> 00:11:48,679
statistical tests. So imagine you're entering a high performance sports

193
00:11:48,720 --> 00:11:51,240
car into a race. Before you can race, the officials

194
00:11:51,240 --> 00:11:53,840
give you a strict rule book. The car must use

195
00:11:54,000 --> 00:11:57,559
premium racing fuel, they must race on a smooth, paved track,

196
00:11:58,039 --> 00:12:00,840
and the tires must be properly inflamed. If you ignore

197
00:12:00,879 --> 00:12:04,240
the rules, say you fill the tank with cheap gas

198
00:12:04,240 --> 00:12:06,240
and try to drive the car through a muddy swamp,

199
00:12:06,679 --> 00:12:08,559
the car is going to stall, sputter, and give you

200
00:12:08,679 --> 00:12:13,279
terrible performance and statistics. The violation of assumptions is exactly

201
00:12:13,320 --> 00:12:15,559
like trying to drive a sports car through the mud.

202
00:12:15,840 --> 00:12:19,399
You're using a highly powerful, sensitive mathematical.

203
00:12:18,679 --> 00:12:21,000
Speaker 1: Tool on data that doesn't fit the rules.

204
00:12:22,639 --> 00:12:25,919
Speaker 2: So parametric statistical tests, these are the high performance sports

205
00:12:25,919 --> 00:12:28,759
cars of math, like a lenova or a T test.

206
00:12:28,799 --> 00:12:30,960
They're incredibly accurate, but the only work if your data

207
00:12:31,039 --> 00:12:35,240
follows strict rules distributional assumptions. The biggest rules that your

208
00:12:35,279 --> 00:12:38,879
data must form a perfect symmetrical Barel curve called the

209
00:12:38,919 --> 00:12:41,240
normal distribution. Most people should score in the middle, with

210
00:12:41,279 --> 00:12:43,159
a few scoring very high and few little low.

211
00:12:45,320 --> 00:12:46,360
Speaker 1: Scaling assumptions, the.

212
00:12:46,399 --> 00:12:48,720
Speaker 2: Data must be measured on a consistent even scale like

213
00:12:48,799 --> 00:12:53,080
height or temperature, not just random categories. If violation happens

214
00:12:53,120 --> 00:12:55,960
when a researcher gets impatient or careless ignores these rules

215
00:12:56,039 --> 00:12:59,039
and forces messy lopsided data into one of these strict

216
00:12:59,080 --> 00:13:00,559
math formulas.

217
00:13:00,600 --> 00:13:04,000
Speaker 1: Anyway, so here plays out.

218
00:13:04,039 --> 00:13:06,759
Speaker 2: So a therapist runs a standard in nova we talked

219
00:13:06,799 --> 00:13:10,960
about in a heavily skewed trauma symptom scale without transformation,

220
00:13:11,080 --> 00:13:15,519
producing misleading f statistics and potentially false conclusions. The messy

221
00:13:15,639 --> 00:13:18,679
data the mud. The researcher is measuring trauma symptoms and patience.

222
00:13:18,720 --> 00:13:21,440
Because trauma severe, almost everyone in the study scores incredibly

223
00:13:21,519 --> 00:13:24,720
high on the symptoms scale, while almost no one scores low.

224
00:13:24,799 --> 00:13:27,000
Speaker 1: This means the data is heavily skewed lopsided.

225
00:13:27,080 --> 00:13:30,080
Speaker 2: It looks like a giant cliff, not a smooth, symmetrical

226
00:13:30,120 --> 00:13:34,039
Bell curve. The violation the researcher wants to compare three

227
00:13:34,080 --> 00:13:37,600
different therapy groups to which one is the best. Instead

228
00:13:37,600 --> 00:13:40,279
of fixing the lopside of data, they immediately run a

229
00:13:40,360 --> 00:13:43,519
standard of nova, a test that mathematically demands a perfect

230
00:13:43,519 --> 00:13:46,200
bell curve that we talked about. So this is the violation.

231
00:13:47,600 --> 00:13:50,279
Because the data broke the rules. The math formula breaks down,

232
00:13:50,759 --> 00:13:54,080
the test spits out a warped, incorrect number and misleading

233
00:13:54,320 --> 00:13:57,919
f statistic. The test might say, Wow, therapy A is

234
00:13:57,960 --> 00:14:00,799
a miracle cure. In reality, therapy didn't work at all.

235
00:14:01,440 --> 00:14:04,840
The broken math created an illusion. It's created a false

236
00:14:04,879 --> 00:14:09,279
positive type one error. So you have to transform the

237
00:14:09,360 --> 00:14:11,960
data with algebra and other things.

238
00:14:11,960 --> 00:14:14,799
Speaker 1: But we'll leave it out of that. Out of it

239
00:14:14,840 --> 00:14:15,120
for now.

240
00:14:15,159 --> 00:14:18,960
Speaker 2: So let's go back another when it's fishing and multiple comparisons.

241
00:14:19,000 --> 00:14:23,279
Running numerous statistical tests without correction inflates the teeth, inflates

242
00:14:23,279 --> 00:14:26,519
the type one error rate, increasing the probability of false positives.

243
00:14:27,159 --> 00:14:30,639
The researcher tests twenty different outcome variables in a single

244
00:14:30,639 --> 00:14:33,360
study without applying it. What they call it bond Feroni

245
00:14:33,440 --> 00:14:36,799
or false discovery or rate correction. By chance alone one

246
00:14:36,879 --> 00:14:39,320
or two will reach the P value less than point

247
00:14:39,440 --> 00:14:43,679
zero five. But these are likely fake findings, what they

248
00:14:43,679 --> 00:14:47,320
call spurious findings. So you can keep running statistical tests

249
00:14:47,399 --> 00:14:51,159
and eventually your false positive will appear. Restriction of range,

250
00:14:51,840 --> 00:14:56,200
limited variability and variable reduces the ability to detect true

251
00:14:56,200 --> 00:14:59,200
correlations or effects. So let me see example of study

252
00:14:59,200 --> 00:15:03,200
examining the relation between therapist experience and client outcomes. Recruits

253
00:15:03,240 --> 00:15:06,279
only licensed clinicians with ten to twenty five years of experience.

254
00:15:07,480 --> 00:15:10,919
The narrow range of experience prevents detection of a gradient

255
00:15:10,960 --> 00:15:13,399
that might be visible of trainees and seasoned veterans were

256
00:15:13,440 --> 00:15:20,679
both included. Error statistical analysis well, misuse of tests, incorrect

257
00:15:20,759 --> 00:15:25,720
model specification, or misinterpretation of P values will affect the conclusions,

258
00:15:25,759 --> 00:15:28,320
and researcher treats a P value zero point four to

259
00:15:28,440 --> 00:15:31,279
nine as definitive proof of a real effect without reporting

260
00:15:31,279 --> 00:15:35,080
effects sizes or confidence intervals, ignoring that defining may be

261
00:15:35,080 --> 00:15:37,720
statistically significant but clinically trivial.

262
00:15:37,960 --> 00:15:38,440
Speaker 1: That's where the.

263
00:15:38,559 --> 00:15:43,200
Speaker 2: Effect sizes and confidence in themals come. In power analysis,

264
00:15:43,360 --> 00:15:47,399
this is for strengthening statistical conclusion validity. Power analysis determines

265
00:15:47,399 --> 00:15:50,240
the minimum sample size needed to reliably detect and effect

266
00:15:50,320 --> 00:15:54,559
on the expected magnitude. Correct statistical models match the analytic

267
00:15:54,559 --> 00:15:55,960
approach to the data structure.

268
00:15:56,440 --> 00:15:57,200
Speaker 1: Research question.

269
00:15:57,320 --> 00:16:00,919
Speaker 2: Measurement level reliability of measures insures that it iss are internally

270
00:16:00,960 --> 00:16:06,399
consistent and stable across time. Adjusting for comparisons using Bonferoni

271
00:16:06,480 --> 00:16:09,639
correction or similar procedures. Controls the family wise air rate,

272
00:16:10,320 --> 00:16:16,039
reporting effects sizes, and confidence intervals alongside p values. Conveys

273
00:16:16,039 --> 00:16:19,200
the magnitude and precision of findings. Using validity concepts to

274
00:16:19,200 --> 00:16:22,840
evaluate research. When reading or designing research, use these four

275
00:16:22,919 --> 00:16:27,360
validity questions as a filter. Does this study rule out

276
00:16:27,360 --> 00:16:31,399
alternative explanations? That's internal validity? Would these findings apply to

277
00:16:31,440 --> 00:16:34,600
my client or setting external validity? Do the measures match

278
00:16:34,639 --> 00:16:38,840
the construct construct validity? And can I trust the data analysis?

279
00:16:39,080 --> 00:16:48,200
Statistical conclusion validity question one to practice for you, The

280
00:16:48,279 --> 00:16:51,559
researcher conducts a six month study evaluating a new cognitive

281
00:16:51,600 --> 00:16:56,039
behavioral intervention for generalized anxiety disorder. Participants are recruited from

282
00:16:56,039 --> 00:16:59,159
a single university counseling center and assessed using only a

283
00:16:59,240 --> 00:17:04,680
self reporting aliety questionnaire. Halfway through the studies, several participants

284
00:17:04,720 --> 00:17:06,799
begin comparing notes with each other in the waiting room.

285
00:17:06,640 --> 00:17:08,599
Speaker 1: And start adjusting their questionnaire responses.

286
00:17:09,480 --> 00:17:12,119
Speaker 2: Which two validity threats are most directly illustrated in this

287
00:17:12,200 --> 00:17:18,920
scenario history and statistical regression, attrition and ecological validity or

288
00:17:19,200 --> 00:17:27,839
hypothesis guessing and monomethod bias. The answer is b B.

289
00:17:28,079 --> 00:17:32,480
Hypothesis guessing and monomethod bias. Participants Altering responses based on

290
00:17:32,559 --> 00:17:35,759
perceived expectations is the definition of hypothesis guessing.

291
00:17:37,160 --> 00:17:38,960
Speaker 1: It constructs validity threat.

292
00:17:39,039 --> 00:17:43,079
Speaker 2: The exclusive use of a single self report measure represents

293
00:17:43,119 --> 00:17:46,920
monomethod bias. Population validiti is also arguably compromised by the

294
00:17:47,039 --> 00:17:50,359
university sample that wasn't given in your choices.

295
00:17:51,559 --> 00:17:52,079
Speaker 1: Another one.

296
00:17:52,079 --> 00:17:56,799
Speaker 2: A clinical researcher publishes findings showing a new mindfulness based

297
00:17:56,799 --> 00:18:01,680
interventions significantly reduced burnout scores sample of emergency room nurses.

298
00:18:01,799 --> 00:18:05,839
The study used random assignment, a control group, standardized measures,

299
00:18:05,880 --> 00:18:08,880
and appropriate statistical modeling. However, all data were collected in

300
00:18:08,920 --> 00:18:13,559
a single large urban hospital in the Pacific Northwest.

301
00:18:15,200 --> 00:18:17,759
Speaker 1: During a period of unusually high staffing levels.

302
00:18:18,000 --> 00:18:22,200
Speaker 2: A colleague wants to implement the intervention which validity concern

303
00:18:22,279 --> 00:18:25,200
is most relevant to the colleague's hesitation because they want

304
00:18:25,200 --> 00:18:28,400
to implement in a rural community health clinic in the

305
00:18:28,440 --> 00:18:33,480
Deep South. So as an A internal validity due to

306
00:18:33,599 --> 00:18:38,759
lack of random assignment, B statistical conclusion validity due to

307
00:18:38,759 --> 00:18:42,759
small sample size, or C external validity due to population

308
00:18:42,960 --> 00:18:48,559
and ecological differences. If you said C, you're right. The

309
00:18:48,599 --> 00:18:51,599
original study had strong internal validity given random assignment and

310
00:18:51,640 --> 00:18:54,160
control conditions. The concern here's where they're finding some urban

311
00:18:55,559 --> 00:19:00,000
Pacific Northwest nurses with high staffing generalized to rural Southern clinicians.

312
00:19:00,839 --> 00:19:04,640
This directly invokes population validity and ecological validity.

313
00:19:05,319 --> 00:19:06,839
Speaker 1: That's it for now. It's a little longer than

314
00:19:06,880 --> 00:19:10,759
Speaker 2: Normal, but hopefully you enjoyed the podcast and have a

315
00:19:10,799 --> 00:19:11,200
good week.

