1
00:00:00,600 --> 00:00:02,000
Speaker 1: All right, welcome back to everybody.

2
00:00:02,040 --> 00:00:05,919
Speaker 2: It's continuing and doing statistics now, so it hasn't really

3
00:00:05,960 --> 00:00:07,960
gone away yet and we're diving.

4
00:00:07,679 --> 00:00:09,759
Speaker 1: Into advanced statistics. But don't panic.

5
00:00:10,320 --> 00:00:12,640
Speaker 2: The E triple P is not expecting you to calculate

6
00:00:12,679 --> 00:00:15,320
these analysis. Instead, the exam wants you to recognize when

7
00:00:15,320 --> 00:00:19,359
each procedure is appropriate. So think of today's lesson as

8
00:00:19,440 --> 00:00:22,839
learning the right tool for the right research question. First

9
00:00:22,839 --> 00:00:26,000
one off the gate is MANOVA. It stands for multivariate

10
00:00:26,039 --> 00:00:30,839
analysis of variants. It compares groups on multiple dependent variables simultaneously.

11
00:00:31,679 --> 00:00:35,079
Think one independent variable with multiple dependent variables.

12
00:00:35,119 --> 00:00:38,000
Speaker 1: For instance, if you want.

13
00:00:37,759 --> 00:00:43,920
Speaker 2: To assign clients to three therapy groups CBT, psychodynamic, and medication,

14
00:00:44,039 --> 00:00:48,399
instead of measuring only depression, they measure depression, anxiety, and

15
00:00:48,439 --> 00:00:52,719
stress at the same time. Rather than running three adova's,

16
00:00:53,000 --> 00:00:55,079
they perform one Manova.

17
00:00:58,560 --> 00:01:00,759
Speaker 1: So why use it?

18
00:01:00,840 --> 00:01:04,239
Speaker 2: Because running several Andova's increases the chance of a type

19
00:01:04,239 --> 00:01:10,439
one or false positive simply by chance. Manova helps control

20
00:01:11,120 --> 00:01:16,719
that inflation. So a Manova again is one independent variable,

21
00:01:16,959 --> 00:01:22,120
multiple dependent variables that compares groups on multiple dependent variables simultaneously.

22
00:01:23,599 --> 00:01:26,200
All right, heading over to do I have one of

23
00:01:26,200 --> 00:01:27,560
the tips you can think of is do I have

24
00:01:27,599 --> 00:01:31,560
one independent variable and multiple dependent variables? If yes, think

25
00:01:31,920 --> 00:01:37,799
Manova and Kova analysis of covariance and Cova compares group

26
00:01:38,200 --> 00:01:42,959
means while statistically controlling for another variable called a covariate.

27
00:01:43,840 --> 00:01:46,519
It's a variable that might influence the dependent variable, but

28
00:01:46,640 --> 00:01:50,560
isn't the main focus of the study. An example be

29
00:01:50,599 --> 00:01:55,560
researchers compare CBT versus medication for depression. However, patients begin

30
00:01:55,640 --> 00:02:00,680
treatment with different levels of depression baseline depression could affect results.

31
00:02:00,840 --> 00:02:05,680
Researchers statistically remove that influence using Ancova. Think of Ancouva

32
00:02:05,680 --> 00:02:12,639
as leveling the playing field before comparing treatments. Discriminate function

33
00:02:12,759 --> 00:02:16,560
analysis This predicts which group someone belongs to using several

34
00:02:16,639 --> 00:02:21,439
predictor variables. So let me give you an example. Can

35
00:02:21,520 --> 00:02:26,520
depression scores, anxiety scores, sleep quality, and trauma history correctly

36
00:02:26,560 --> 00:02:32,439
classify someone as having PTSD or generalized anxiety disorder. The

37
00:02:32,520 --> 00:02:39,719
dependent variables categorical, The predictors are continuous, perfect situation for

38
00:02:39,800 --> 00:02:45,120
discriminate functional or function analysis again predicts which group someone

39
00:02:45,159 --> 00:02:49,080
belongs to using several predictor variables. Right, the predictor variables

40
00:02:49,080 --> 00:02:53,759
being anxiety scores, depression scores, and then, of course, the

41
00:02:53,800 --> 00:02:56,280
dependent variables categorical.

42
00:02:56,759 --> 00:02:59,879
Speaker 1: So as, PTSD or generalized anxieties.

43
00:03:01,400 --> 00:03:07,520
Speaker 2: NeXT's factor analysis, which discovers hidden psychological traits called latent variables.

44
00:03:08,719 --> 00:03:12,879
A latent variables a psychological characteristic. It cannot be directly observed,

45
00:03:12,919 --> 00:03:18,840
but can be measured indirectly. Examples include intelligence, depression, anxiety,

46
00:03:18,919 --> 00:03:22,919
self esteem. You cannot see self esteem, you measure it

47
00:03:23,039 --> 00:03:26,479
using questionnaires. So I imagine a personality test with one

48
00:03:26,560 --> 00:03:29,879
hundred questions. Many of these questions actually measure only five

49
00:03:30,000 --> 00:03:38,240
personality traits. Factor analysis uncovers those hidden dimensions EFA the

50
00:03:38,280 --> 00:03:42,599
exploratory factor analysis. It explores data without having a prior theory.

51
00:03:42,800 --> 00:03:45,599
Researchers ask what factors naturally appear.

52
00:03:47,000 --> 00:03:48,479
Speaker 1: So if you do a study, a.

53
00:03:48,439 --> 00:03:52,199
Speaker 2: Psychologist develops a new trauma questionnaire, nobody knows how many

54
00:03:52,199 --> 00:03:57,120
dimensions exist. Yet maybe there's three, maybe five. Exploratory factor

55
00:03:57,159 --> 00:04:04,599
analysis lets the data decide. Another important vocabulary word is eigenvalue.

56
00:04:04,680 --> 00:04:07,120
Speaker 1: That's E I G E and V A l u E.

57
00:04:07,199 --> 00:04:12,039
Speaker 2: It measures how much variance a factor explains. Higher eigenvalues

58
00:04:12,120 --> 00:04:17,839
usually indicate more important factors. A screen plot s c

59
00:04:18,279 --> 00:04:21,800
R E e is a graph showing where meaningful factors

60
00:04:21,839 --> 00:04:23,000
begin to level off.

61
00:04:23,600 --> 00:04:26,959
Speaker 1: Imagine a mountain. The point where it flatteness tells you

62
00:04:27,000 --> 00:04:28,519
how many factors to keep.

63
00:04:30,399 --> 00:04:33,600
Speaker 2: Factor loading the correlation between a question and a factor.

64
00:04:34,319 --> 00:04:38,800
Large loadings mean that that question strongly measures that factor.

65
00:04:40,800 --> 00:04:44,800
Commonality the pro proportion of variants in a variable explained

66
00:04:44,800 --> 00:04:46,160
by all factors together.

67
00:04:49,399 --> 00:04:51,759
Speaker 1: It might be easier to actually go back as I

68
00:04:51,800 --> 00:04:54,279
was reading this to give you example.

69
00:04:54,360 --> 00:04:56,879
Speaker 2: So the eigenvalue again, the higher the more important the

70
00:04:56,920 --> 00:05:01,920
factor is. So imagine you create a forty questions personality questionnaire.

71
00:05:02,560 --> 00:05:06,360
After running an exploratory factor analysis that we talked about earlier,

72
00:05:06,800 --> 00:05:09,920
you discover four factors. Factor one has an eigenvalue of

73
00:05:09,959 --> 00:05:14,360
eight point five, factor three one point four. Factor one

74
00:05:14,399 --> 00:05:18,680
explains the largest amount of personality differences among participants. Most

75
00:05:18,720 --> 00:05:22,240
researchers keep factors with an eigenvalue greater than one, So

76
00:05:22,360 --> 00:05:25,040
factors one through three in the example which I didn't

77
00:05:25,040 --> 00:05:27,000
give you all the numbers would likely be retained well.

78
00:05:27,079 --> 00:05:30,040
Factor four at a zero point six would probably be discarded.

79
00:05:30,600 --> 00:05:33,800
Think of of an eigenvalue as a factor's importance score.

80
00:05:34,720 --> 00:05:37,199
The screen plot is a graph that helps researchers decide

81
00:05:37,240 --> 00:05:39,600
how many factors should be kept. Researchers look for the

82
00:05:39,639 --> 00:05:42,040
point where the graph suddenly levels off like.

83
00:05:42,000 --> 00:05:44,120
Speaker 1: I was telling you about the mountains. Is called the elbow.

84
00:05:44,759 --> 00:05:48,759
Speaker 2: Suppose your questionnaire initially identifies eight possible factors. When you

85
00:05:48,839 --> 00:05:52,639
examine the screen plot, the first three factors drop sharply,

86
00:05:53,160 --> 00:05:54,639
and after the third factor.

87
00:05:54,360 --> 00:05:55,560
Speaker 1: The line becomes flat.

88
00:05:56,160 --> 00:05:59,879
Speaker 2: That tells you that only three meaningful psychological factors exist,

89
00:06:00,120 --> 00:06:05,639
while the remaining factors mostly represent noise. Factor loading, as

90
00:06:05,680 --> 00:06:07,959
we told you earlier, as the correlation between an individual

91
00:06:08,000 --> 00:06:12,600
test item and a particular factor. Suppose one factor's label depression.

92
00:06:12,680 --> 00:06:16,160
One questionnaire item says I feel hopeless most days. It's

93
00:06:16,279 --> 00:06:19,040
factor loading is zero point eighty seven. Another one says

94
00:06:19,040 --> 00:06:22,759
I enjoy watching sports. It's factor loading a zero point nine.

95
00:06:23,600 --> 00:06:27,040
The first question strongly measures depression, in the second almost

96
00:06:27,279 --> 00:06:31,439
nothing to that factor. Factor loadings are like magnets. The

97
00:06:31,480 --> 00:06:33,959
stronger the loading, the more stronger the questions pulled toward

98
00:06:34,000 --> 00:06:34,600
that factor.

99
00:06:35,720 --> 00:06:36,879
Speaker 1: Another one we did not discuss.

100
00:06:36,959 --> 00:06:40,600
Speaker 2: Yes, communality represents the percentage of a question's variance explained

101
00:06:40,639 --> 00:06:44,800
by all of the retained factors combined. Higher communalities we

102
00:06:44,879 --> 00:06:48,600
mean that the factor model explains that question well. Suppose

103
00:06:48,680 --> 00:06:51,959
the questionnaire item I feel nervous in crowds has a

104
00:06:52,000 --> 00:06:55,920
communality of zero point eight two. That means that eighty

105
00:06:55,959 --> 00:06:58,439
two percent of the differences in people's responses can be

106
00:06:58,480 --> 00:07:02,879
explained by the psychological fing identified in your model, such

107
00:07:02,879 --> 00:07:08,000
as anxiety and social fear. Imagine a different one, I

108
00:07:08,079 --> 00:07:12,560
prefer chocolate over vanilla. Its communality is zero point one eight.

109
00:07:13,240 --> 00:07:16,800
Only eighteen percent of the variation is explained by your factors,

110
00:07:16,879 --> 00:07:19,079
suggesting this question probably doesn't.

111
00:07:18,839 --> 00:07:21,160
Speaker 1: Belong to an anxiety questionnaire.

112
00:07:22,720 --> 00:07:27,240
Speaker 2: Another one is orth orthogonal rotation. Orthogonal rotation it's O

113
00:07:27,480 --> 00:07:31,639
R thho g O n A L or verimax assumes

114
00:07:31,639 --> 00:07:35,079
that the factors are completely independent and unrelated to one another.

115
00:07:36,199 --> 00:07:41,360
It's a researcher develops a vocational aptitude test measuring mathematical ability,

116
00:07:42,000 --> 00:07:46,839
musical ability, mechanical ability. The researcher assumes these abilities are unrelated.

117
00:07:47,759 --> 00:07:51,759
Verimax rotation keeps the factors independent and easier to interpret.

118
00:07:52,959 --> 00:07:56,680
Orthogonical orthogonal equals orthodox and dependence.

119
00:07:56,720 --> 00:07:58,800
Speaker 1: The factors stay separate.

120
00:08:01,000 --> 00:08:05,000
Speaker 2: Oblique rotation or promax allows factors to be correlated. Because

121
00:08:05,000 --> 00:08:10,680
many psychological traits naturally overlap, for instance, depression, anxiety, and stress,

122
00:08:10,720 --> 00:08:12,600
these constructs often occur together.

123
00:08:12,879 --> 00:08:14,920
Speaker 1: Some estimate over seventy percent of the time.

124
00:08:15,639 --> 00:08:21,000
Speaker 2: Someone with severe anxiety frequently expresses experiences depression as well. Promax,

125
00:08:21,000 --> 00:08:24,279
which a rotation allows those factors to correlate.

126
00:08:23,879 --> 00:08:25,759
Speaker 1: Making the results much more realistic.

127
00:08:26,560 --> 00:08:29,560
Speaker 2: Oblique means the factors are allowed to lean toward one

128
00:08:29,560 --> 00:08:35,000
another instead of standing apart. Another quick rule to help

129
00:08:35,039 --> 00:08:37,840
you is I can value is how important is the factor?

130
00:08:38,399 --> 00:08:42,000
Screenplot is? How many factors should I keep? Factor loading?

131
00:08:42,080 --> 00:08:44,120
Which questions belong to each factor?

132
00:08:44,720 --> 00:08:45,519
Speaker 1: Communality?

133
00:08:45,679 --> 00:08:50,399
Speaker 2: How well do the factors explain each question? Verimax factors

134
00:08:50,399 --> 00:08:53,279
stay independent promax the factors.

135
00:08:52,879 --> 00:08:54,519
Speaker 1: Are allowed to correlate.

136
00:08:56,080 --> 00:08:58,320
Speaker 2: Yes, this is going to be something that's probably pretty

137
00:08:58,360 --> 00:09:03,399
Foreign's covered quite a bit in classes that I remember.

138
00:09:04,879 --> 00:09:08,759
All right, so we move along past that now and

139
00:09:08,799 --> 00:09:12,120
we're heading into a different territory. So now we're going

140
00:09:12,200 --> 00:09:21,159
into what they call the structured sorry, what they call

141
00:09:21,200 --> 00:09:26,440
the structural equation modeling the SEM. This combines several statistical

142
00:09:26,600 --> 00:09:31,480
techniques into one large model. It can test regression path analysis,

143
00:09:31,639 --> 00:09:37,720
confirmatory factor analysis, direct effects, indirect effects latent variables all

144
00:09:37,799 --> 00:09:42,720
at once. The structural equation modeling combines several statistical techniques

145
00:09:42,720 --> 00:09:48,559
into one large model. So, for example, researchers propose parenting style,

146
00:09:49,559 --> 00:09:51,840
self esteem, and academic performance.

147
00:09:52,120 --> 00:09:53,240
Speaker 1: That's what they're going to look at.

148
00:09:53,519 --> 00:10:00,960
Speaker 2: SEM tests the entire model simultaneously. One analysis, multiple relationship.

149
00:10:02,879 --> 00:10:06,559
NeXT's mediation. A mediator explains how or why one variable

150
00:10:06,600 --> 00:10:10,080
affects another. So if you look at mindfulness improves emotion

151
00:10:10,200 --> 00:10:14,480
and regulation, which then reduces anxiety. Emotion regulation explains why

152
00:10:14,720 --> 00:10:20,480
mindfulness works. It's the mediator, right, mindfulness anxiety the two players.

153
00:10:20,960 --> 00:10:25,759
But why does mindfulness reduce anxiety because emotional regulation improves

154
00:10:27,080 --> 00:10:34,200
Remember mediator answers how. Moderation as a moderator tells us when,

155
00:10:34,440 --> 00:10:39,039
for whom or under what conditions the relationship exists. CBT

156
00:10:39,240 --> 00:10:42,360
reduces oppression. Now you're looking for the questions when, for

157
00:10:42,440 --> 00:10:45,919
whom and under what? But only among younger adults. That's

158
00:10:45,960 --> 00:10:50,440
for whom age changes the strength of treatment under what conditions?

159
00:10:50,440 --> 00:11:01,559
And age is the moderator. Remember moderator answers when or whom?

160
00:11:02,879 --> 00:11:07,399
Non parametric tests these are used when assumptions like normality

161
00:11:07,440 --> 00:11:11,279
or violated remember that from the.

162
00:11:10,080 --> 00:11:11,720
Speaker 1: Distribution that we talked about.

163
00:11:11,960 --> 00:11:16,120
Speaker 2: When they're evenly distributed, non parametric version of the independent

164
00:11:16,159 --> 00:11:20,120
samples T tests you use Man Whitney you so non

165
00:11:20,720 --> 00:11:23,799
parametric version of the independent samples T test you use

166
00:11:23,879 --> 00:11:29,120
for two independent groups, so comparing depression scores between males

167
00:11:29,120 --> 00:11:33,559
and females. When scores are highly skewed, they're not equally

168
00:11:33,600 --> 00:11:38,240
distributed like the Bell curve, then you would go to

169
00:11:38,320 --> 00:11:43,240
the Man Whitney You A non parametric equivalent of the

170
00:11:43,279 --> 00:11:48,519
paired samples T test is will Coxin signed ranked test,

171
00:11:48,759 --> 00:11:52,440
well Coxon signed ranked test. It's used for two related groups,

172
00:11:52,799 --> 00:11:57,639
so depression before therapy versus after therapy.

173
00:11:57,799 --> 00:12:01,720
Speaker 1: But again they're skewed. Equivalent of a one way on nova.

174
00:12:01,720 --> 00:12:04,759
Speaker 2: It's something called the Cruse Call Wallace test k r

175
00:12:04,879 --> 00:12:12,399
U SKA l WA LI s so Cruse Call Wallace test.

176
00:12:12,840 --> 00:12:16,279
It looks at more than two independent groups, so it's

177
00:12:16,320 --> 00:12:22,120
comparing three therapy programs using ordinal data right in order.

178
00:12:23,600 --> 00:12:30,240
So again it's a non parametric freedoman test FRI E

179
00:12:30,360 --> 00:12:32,759
D M A N is the equivalent of repeated measures

180
00:12:32,799 --> 00:12:37,679
a NOVA more than two repeated measurements. Example, anxiety measured

181
00:12:37,759 --> 00:12:43,080
before treatment, mid treatment, and after treatment using ordinal rankings again.

182
00:12:44,480 --> 00:12:47,919
Now head over to hierarchical linear modeling also called multi

183
00:12:48,000 --> 00:12:49,120
level modeling.

184
00:12:49,320 --> 00:12:51,000
Speaker 1: Use whenever data are nested.

185
00:12:52,039 --> 00:12:57,559
Speaker 2: Nested data means participants naturally belong inside larger groups. Examples

186
00:12:57,600 --> 00:13:04,039
of students in side classrooms, patients inside hospitals, employees inside companies.

187
00:13:04,960 --> 00:13:06,600
Speaker 1: You even think about it as clients in.

188
00:13:06,559 --> 00:13:11,600
Speaker 2: Site therapists psychology example, fifty therapists each treat ten clients.

189
00:13:11,720 --> 00:13:16,080
Clients treated by the same therapists resemble one another. Ordinary

190
00:13:16,120 --> 00:13:18,519
regression assumes everyone is independent.

191
00:13:19,159 --> 00:13:20,720
Speaker 1: That assumption is violated.

192
00:13:21,840 --> 00:13:27,639
Speaker 2: HLM properly analyzes the nested data right, because these individuals

193
00:13:27,639 --> 00:13:31,320
are seeing the therapists for particular reasons, similar reasons right,

194
00:13:31,360 --> 00:13:36,919
psychiatric issues, depression, anxiety, whatnot. Statistical control that means you

195
00:13:36,960 --> 00:13:39,440
remove the influence of confounding variables.

196
00:13:39,600 --> 00:13:40,799
Speaker 1: Probably heard this word before.

197
00:13:40,840 --> 00:13:43,120
Speaker 2: It's a variable related to both the predictor and outcome

198
00:13:43,159 --> 00:13:48,519
that may falsely explain an unobserved relationship. So exercise appears

199
00:13:48,559 --> 00:13:53,000
to reduce suppression. However, younger people exercise more age could

200
00:13:53,000 --> 00:13:54,000
explain the finding.

201
00:13:54,519 --> 00:13:56,879
Speaker 1: Researchers statistically control.

202
00:13:56,519 --> 00:14:00,559
Speaker 2: For age, so they eliminate the young people or they

203
00:14:00,600 --> 00:14:03,519
separate the groups to track to see if there's differences.

204
00:14:04,120 --> 00:14:07,840
Effect size measures the magnitude of an effect, not merely

205
00:14:07,879 --> 00:14:13,159
whether it's statistically significant.

206
00:14:14,679 --> 00:14:16,399
Speaker 1: So we use Coen's D for this.

207
00:14:16,600 --> 00:14:19,759
Speaker 2: Used with T tests, small as zero point two zero,

208
00:14:19,919 --> 00:14:23,159
medium is zero point five, and then large as zero point.

209
00:14:22,879 --> 00:14:23,440
Speaker 1: Eight to zero.

210
00:14:24,159 --> 00:14:28,639
Speaker 2: Partial Eta squared us in n Nova and Manova is

211
00:14:28,720 --> 00:14:32,360
small zero point point zero one, medium point zero six,

212
00:14:32,440 --> 00:14:36,320
and largest point one four. For the partial Eta square,

213
00:14:36,360 --> 00:14:39,679
these again I'm measuring the magnitude of an effect. The

214
00:14:39,720 --> 00:14:43,360
last one is Cohen's F using regression analysis, small as

215
00:14:43,440 --> 00:14:47,240
point zero two, medium point one five, and largest point

216
00:14:47,519 --> 00:14:50,039
three five. The most common ones are the ones used

217
00:14:50,039 --> 00:14:53,399
with T tests, which is Cohen's D, and every so

218
00:14:53,480 --> 00:14:56,519
often you'll probably see partial Eta squared on a nova

219
00:14:56,639 --> 00:15:00,440
in Manova. Our square represents the percentage of variants explained

220
00:15:00,440 --> 00:15:03,320
by the predictors. Example, are squared equals point six to

221
00:15:03,360 --> 00:15:06,519
zero means the model explains sixty percent of the variability

222
00:15:06,559 --> 00:15:12,879
and depression scores. All right, rapid fire review, ask yourself

223
00:15:12,960 --> 00:15:20,279
multiple dependent variables Boom Manova control for baseline differences, and

224
00:15:20,559 --> 00:15:24,759
COVA if you want to find hidden personality traits EFA,

225
00:15:24,960 --> 00:15:31,240
exploratory factor analysis, confirm an existing theory, confirmatory factor analysis

226
00:15:31,360 --> 00:15:33,720
right in the word right, So confirm an existing theory.

227
00:15:33,720 --> 00:15:40,799
Confirmatory factor analysis, predict group membership, discriminate function analysis, and

228
00:15:40,960 --> 00:15:48,840
model complex relationships. Sem explain how something works. Mediation explain

229
00:15:49,080 --> 00:15:55,120
when it works. Moderation or for whom Nested data, hierarchical

230
00:15:55,559 --> 00:16:01,840
linear modeling, and assumptions are violent. You go to non

231
00:16:02,360 --> 00:16:08,399
parametric tests, all right, So let's get some practice questions

232
00:16:08,399 --> 00:16:13,559
and wrap it up today. Psychologist compares three different therapy programs,

233
00:16:13,720 --> 00:16:20,399
measures both depression and anxiety after treatment, which statistical test

234
00:16:20,480 --> 00:16:29,480
is most appropriate and Cova Manova or multiple regression. If

235
00:16:29,480 --> 00:16:32,559
you said Manova, you're right. There is one independent variable,

236
00:16:32,600 --> 00:16:35,039
which is therapy type. Even though they're three different therapy

237
00:16:35,080 --> 00:16:40,960
programs and multiple dependent variables depression and anxiety manova and

238
00:16:41,039 --> 00:16:44,639
analyzes them simultaneously. Are ConTroll controlling for inflated type one

239
00:16:44,720 --> 00:16:45,600
false positives?

240
00:16:46,919 --> 00:16:47,240
Speaker 1: All right.

241
00:16:47,360 --> 00:16:51,799
Speaker 2: Researchers discover that mindfulness reduces anxiety because it approves emotion regulation.

242
00:16:51,879 --> 00:16:55,919
Emotion regulation is best described as what moderator, mediator or

243
00:16:55,960 --> 00:17:04,000
covariate right explains how and why an independent variable influences

244
00:17:04,079 --> 00:17:07,599
and outcome. That's it for now. Hopefully enjoyed this podcast.

245
00:17:07,640 --> 00:17:08,680
We'll see you all next time.

