1
00:00:01,720 --> 00:00:02,560
Speaker 2: Welcome back everyone.

2
00:00:02,560 --> 00:00:04,679
Speaker 1: Today we're gonna get back into the weeds of still

3
00:00:04,679 --> 00:00:09,199
into research. We're covering psychometrics, reliability, and validity.

4
00:00:10,679 --> 00:00:13,880
Speaker 2: So every test you've ever given a.

5
00:00:13,800 --> 00:00:18,079
Speaker 1: Client, every questionnaire, every checklist, is making a silent promise

6
00:00:18,079 --> 00:00:19,879
and it is promising that the number it spets out

7
00:00:20,079 --> 00:00:22,760
actually means something. Today we're pulling that promise apart. So

8
00:00:22,760 --> 00:00:25,280
we're going to be talking about psychometrics, the mathema logic

9
00:00:26,039 --> 00:00:29,280
behind whether your tools can be trusted. So let's look

10
00:00:29,280 --> 00:00:30,359
at some of the things we're going to look at.

11
00:00:30,359 --> 00:00:34,079
Scale of measurement. First, nominal scales or categorical with no order.

12
00:00:34,600 --> 00:00:39,759
Think diagnosis categories, gender, eye color. There's no more then

13
00:00:39,960 --> 00:00:43,880
or less then blue eyes. Here are just different buckets.

14
00:00:43,920 --> 00:00:47,079
Because there's no order, you're allowed analysis are limited to

15
00:00:47,119 --> 00:00:49,719
frequency counts and what they call Ki square tests, which

16
00:00:49,759 --> 00:00:53,840
we've talked about before. The ordinal scales rank categories, but

17
00:00:53,960 --> 00:00:56,840
the distance between ranks is not equal. So the classic

18
00:00:56,960 --> 00:01:00,920
psychic example is a LIKEERD scale that strongly agree or

19
00:01:01,240 --> 00:01:05,439
strongly disagree you see on almost every self report measure.

20
00:01:05,840 --> 00:01:08,680
A client moving from moderate to severe depression on a

21
00:01:08,799 --> 00:01:12,239
rating scale tells you direction, but not really the magnitude

22
00:01:12,920 --> 00:01:17,400
allowed analysis. Here are medians, percentiles, and non parametric tests

23
00:01:17,480 --> 00:01:21,560
like the Man Whitney you, which will be important to understand.

24
00:01:22,480 --> 00:01:26,560
Interval scales of equal intervals, but no true zero, So IQ.

25
00:01:26,400 --> 00:01:28,280
Speaker 2: Scores are the TechBook textbook.

26
00:01:27,879 --> 00:01:30,640
Speaker 1: Example zero and IQ score does not mean the absence

27
00:01:30,680 --> 00:01:33,760
of intelligence. It just means a raw score of zero.

28
00:01:33,879 --> 00:01:36,480
You can compute means and standard deviations and run T

29
00:01:36,680 --> 00:01:40,560
tests and Anova's, but you cannot meaningfully say someone with

30
00:01:40,599 --> 00:01:43,840
a one forty RQ is twice as smart as someone

31
00:01:43,840 --> 00:01:48,280
with a seventy because there's no true zero. Now, if

32
00:01:48,319 --> 00:01:52,000
you have a ratio, which is have equal intervals and

33
00:01:52,040 --> 00:01:56,079
a true zero, reaction time and number of errors on

34
00:01:56,120 --> 00:02:00,000
a task are ratio data. So here all mathematical operations

35
00:02:00,079 --> 00:02:03,640
since they're fair game, including ratios. So when you say

36
00:02:03,640 --> 00:02:07,319
a test score doubled that's true, Ask yourself, is this

37
00:02:07,560 --> 00:02:10,360
actually a racial scale? Am I assuming one where it

38
00:02:10,400 --> 00:02:13,400
does not exist? I accuse the perfect trap here because

39
00:02:13,439 --> 00:02:16,199
it feels like it should be ratio, but it's not

40
00:02:16,319 --> 00:02:24,080
because there's no true zero. Segment two reliability consistency of

41
00:02:24,240 --> 00:02:27,360
measurement reliabilities about consistency, it does not ask whether a

42
00:02:27,439 --> 00:02:29,599
test measures the right thing only whether it measures the

43
00:02:29,639 --> 00:02:33,639
same thing in the same way every time. Test retest

44
00:02:33,680 --> 00:02:35,759
reliability measures temporal stability.

45
00:02:35,759 --> 00:02:37,280
Speaker 2: You administer the same test.

46
00:02:37,080 --> 00:02:40,879
Speaker 1: Twice and look for a high correlation between the two administrations.

47
00:02:41,680 --> 00:02:43,879
Let me give you a clinical example, giving a client

48
00:02:44,000 --> 00:02:46,599
the Beck depression inventory at intake and again two weeks

49
00:02:46,680 --> 00:02:50,199
later to see if scores hold steady absent in any intervention.

50
00:02:50,719 --> 00:02:53,520
The threats here are practice effects and maturation, meaning the

51
00:02:53,560 --> 00:02:56,360
person might get better at taking the test or might

52
00:02:56,400 --> 00:02:58,800
actually change between sessions for reasons.

53
00:02:58,560 --> 00:02:59,879
Speaker 2: Unrelated to the construct.

54
00:03:00,879 --> 00:03:04,240
Speaker 1: In tonal consistency most often measured with something called the

55
00:03:04,319 --> 00:03:08,919
Kronback's alpha. This asks whether items within the test measure

56
00:03:09,000 --> 00:03:12,759
the same construct. An alpha above zero point seven is

57
00:03:12,800 --> 00:03:17,080
generally considered acceptable. Here's a trap you gotta be careful with.

58
00:03:17,439 --> 00:03:21,280
An alpha above ninety is not automatically better. It can't

59
00:03:21,360 --> 00:03:26,120
actually indicate item redundancy, meaning your items are just restating

60
00:03:26,159 --> 00:03:29,800
each other rather than each contributing something useful. So split

61
00:03:29,840 --> 00:03:32,520
half reliability is a related method where you divide the

62
00:03:32,560 --> 00:03:37,360
test and two and correlate the halves. Another one is

63
00:03:37,400 --> 00:03:41,879
intirator reliability. It's as agreement between observers or scores calculating

64
00:03:42,319 --> 00:03:46,560
calculator using something called the Kappa Kappa or inner intra

65
00:03:46,639 --> 00:03:51,080
class correlation the ICC. This matter is enormously in behavioral coding,

66
00:03:51,199 --> 00:03:56,120
clinical diagnosis, and qualitative scoring. Picture two trained clinicians independently

67
00:03:56,159 --> 00:03:59,960
coding videotape sessions for therapists empathy. Their agreement or lack

68
00:04:00,879 --> 00:04:03,199
is inturrator reliability.

69
00:04:02,599 --> 00:04:04,919
Speaker 2: And action parallel forms.

70
00:04:04,960 --> 00:04:08,479
Speaker 1: Reliability uses two versions of a test measuring the same construct,

71
00:04:08,520 --> 00:04:12,240
administered to the same people, where scores should correlate highly.

72
00:04:12,639 --> 00:04:16,120
This controls for form specific area, which is why standardized

73
00:04:16,160 --> 00:04:22,639
tests like the SAT have multiple versions across administrations. Another

74
00:04:22,680 --> 00:04:25,519
part is validity, the accuracy of measurement. Here's a line

75
00:04:25,519 --> 00:04:28,879
that should be tattooed in every person's forearm. You can

76
00:04:28,920 --> 00:04:35,399
have reliability without validity, but you cannot have validity without reliability.

77
00:04:35,959 --> 00:04:37,759
Speaker 2: A broken scale and it always.

78
00:04:37,519 --> 00:04:41,959
Speaker 1: Reads five pounds too heavy is perfectly reliable and completely invalid.

79
00:04:42,959 --> 00:04:46,720
Content Validity asks whether the test items addequately cover the construct.

80
00:04:47,040 --> 00:04:48,199
Speaker 2: It is evaluated by.

81
00:04:48,040 --> 00:04:51,040
Speaker 1: Expert judgment, not statistics, and it is key in achievements

82
00:04:51,040 --> 00:04:54,959
and personality scores. An example would be a generalized anxiety scale.

83
00:04:54,959 --> 00:05:00,160
Should cover worry, physical tensions, sleep disturbance, and concentration difficulty,

84
00:05:00,879 --> 00:05:02,639
not just one slice.

85
00:05:02,199 --> 00:05:03,079
Speaker 2: Of the construct.

86
00:05:04,040 --> 00:05:07,160
Speaker 1: Criterion validity asks how well the test correlates with an

87
00:05:07,160 --> 00:05:08,120
external standard.

88
00:05:08,399 --> 00:05:09,519
Speaker 2: There are two favors.

89
00:05:09,560 --> 00:05:12,959
Speaker 1: Flavors Concurrent validity is measured at the same time, like

90
00:05:13,000 --> 00:05:16,399
comparing a new depression screener against the same day clinical diagnosis.

91
00:05:17,000 --> 00:05:20,079
Predictive validity is about the test predicting a future outcome,

92
00:05:20,759 --> 00:05:24,920
the classic example being SAT scores predicting college gpa.

93
00:05:25,120 --> 00:05:27,360
Speaker 2: Both are measured using correlation coefficients.

94
00:05:27,480 --> 00:05:31,639
Speaker 1: Are Construct validity asks whether the test is measuring the

95
00:05:31,639 --> 00:05:35,160
theoretical construct that claims to measure. That splits into convergent

96
00:05:35,240 --> 00:05:38,920
validity where the test correlates with measures of similar constructs,

97
00:05:38,920 --> 00:05:42,720
and discriminate validity whether it where it does not correlate

98
00:05:42,759 --> 00:05:47,279
with unrelated constructs. An example would be a new social

99
00:05:47,279 --> 00:05:50,639
anxiety scale should correlate with an established social anxiety measure

100
00:05:50,800 --> 00:05:54,160
that's convergent, but should not correlate strongly with a measure

101
00:05:54,199 --> 00:05:56,560
of an unrelated trait like extraversion.

102
00:05:56,959 --> 00:05:59,879
Speaker 2: That would be a discriminate validity.

103
00:06:00,199 --> 00:06:05,360
Speaker 1: This is demonstrated through factor analysis and multi trade method matrixes.

104
00:06:06,240 --> 00:06:08,319
Speaker 2: Face validity simply.

105
00:06:07,920 --> 00:06:10,519
Speaker 1: Asked whether the test looks like it measures with it

106
00:06:10,519 --> 00:06:13,800
should It is not scientific. It can, but it can't

107
00:06:13,839 --> 00:06:17,759
influence client engagement and social acceptability. And item like I

108
00:06:17,800 --> 00:06:20,720
feel sad most days has high face validity for depression,

109
00:06:20,759 --> 00:06:24,120
even if it's statistical properties still need to be tested.

110
00:06:25,279 --> 00:06:28,480
Speaker 2: Another aspect is measurement issues.

111
00:06:28,959 --> 00:06:31,519
Speaker 1: This is where good instruments can go wrong in practice.

112
00:06:31,680 --> 00:06:35,639
Response sets are systematic ways people respond, regardless of item content.

113
00:06:36,560 --> 00:06:41,560
Acquiescence is always agreeing. Extreme responding means only using the

114
00:06:41,720 --> 00:06:45,439
end points of a scale. Central tendency bias means avoiding

115
00:06:45,480 --> 00:06:48,920
extremes altogether and clustering in the middle. You mitigate these

116
00:06:48,959 --> 00:06:53,519
with reverse coded items and balance scales. Social desirability this

117
00:06:53,639 --> 00:06:55,680
is the tendency to answer in a way that looks

118
00:06:55,680 --> 00:06:58,480
good and can inflate or obscure a client's true traits.

119
00:07:00,600 --> 00:07:03,000
I think of a client under reporting substance use or

120
00:07:03,040 --> 00:07:05,800
minimizing anger because they want to be see seen favorably.

121
00:07:06,399 --> 00:07:09,439
This is exactly why instruments like the MMMPI two build

122
00:07:09,480 --> 00:07:12,759
in validity scales, and why clinicians sometimes use indirect items

123
00:07:12,800 --> 00:07:16,079
to detect it. Ceiling and floor effects happen when the

124
00:07:16,120 --> 00:07:18,240
test is too easy or too hard for the population.

125
00:07:18,879 --> 00:07:22,240
A ceiling effect means scores bunch at the top, like

126
00:07:22,279 --> 00:07:24,920
giving a standard i Q test to a gifted population

127
00:07:24,959 --> 00:07:28,600
where everyone tops out. A floor means scores bunch at

128
00:07:28,639 --> 00:07:31,079
the bottom, like using an adult cognitive measure where the

129
00:07:31,079 --> 00:07:36,360
population has significant intellectual disability. Both reduced variability, which weakens

130
00:07:36,439 --> 00:07:41,800
correlations and discrimination. Cultural bias occurs when items reflect the

131
00:07:41,839 --> 00:07:47,079
dominant culture's norm or language, potentially disadvantaging non native speakers

132
00:07:48,240 --> 00:07:53,920
or neuroindividuent in nerve neurodivergent individuals. Measurement models, we look

133
00:07:53,959 --> 00:07:58,160
at classical tests. Theory CTTT arrests on one core equation.

134
00:07:59,079 --> 00:08:05,800
Every observes score equals the true score plus error. Reliability

135
00:08:05,839 --> 00:08:08,839
under this model reflects the proportion of true score variance,

136
00:08:08,839 --> 00:08:11,279
and its limitation is that it treats eras random and

137
00:08:11,319 --> 00:08:17,079
does not account for individual item difficulty. Item response theory

138
00:08:17,120 --> 00:08:19,959
goes deeper. It examines the relationship between a person's latent

139
00:08:20,160 --> 00:08:23,480
trait level and their probability of responding correctly to each

140
00:08:23,519 --> 00:08:28,560
specific item. Each item carries its own parameters difficulty, how

141
00:08:28,560 --> 00:08:31,319
hard it is, discrimination, how well it separates people of

142
00:08:31,360 --> 00:08:34,080
different ability levels, and guessing the chance of getting it

143
00:08:34,200 --> 00:08:38,720
right without actually having knowledge about it. IRT allows for

144
00:08:38,759 --> 00:08:41,799
adaptive testing and better scale precision, which is why you

145
00:08:41,840 --> 00:08:44,279
see it in large scale assessments of the GRE and

146
00:08:44,279 --> 00:08:49,559
in certain adaptations of the MMPI. The six steps of

147
00:08:49,559 --> 00:08:52,720
scale development. If you ever had to build a clinical

148
00:08:52,720 --> 00:08:55,919
measure from scratch, here's the order of operation. To find

149
00:08:55,919 --> 00:08:58,480
the construct is Number one, what exactly are you measuring

150
00:08:58,679 --> 00:09:02,879
based on theory, literature, clinical input? Number two generate items.

151
00:09:02,919 --> 00:09:05,639
These are questions using experts, clients, and focus groups that

152
00:09:05,679 --> 00:09:09,960
create a broad item pool. Number three pilot testing deministering

153
00:09:09,960 --> 00:09:14,440
items to a sample and analyzing item total, correlations, difficulty,

154
00:09:14,480 --> 00:09:19,120
and discrimination. Number four factor analysis, which splits into exploratory

155
00:09:19,159 --> 00:09:24,279
factor analysis, uncovering underlying dimensions you did not specify in advance.

156
00:09:24,320 --> 00:09:28,919
In confirmatory factor analysis, testing a hypothesized structure you already

157
00:09:28,960 --> 00:09:32,039
believe is true. The goal is to identify and retain

158
00:09:32,159 --> 00:09:37,519
strongly loading items. Number five establish reliability and validity, checking

159
00:09:37,600 --> 00:09:43,000
internal consistency, test retests, constructing criteria validity, and number six norming,

160
00:09:43,039 --> 00:09:46,519
which is administering the final instrument to a large representative sample.

161
00:09:46,759 --> 00:09:49,039
Speaker 2: To create standard scores and cutoffs.

162
00:09:49,600 --> 00:09:52,799
Speaker 1: When selecting a sec test, match into the referral question

163
00:09:52,879 --> 00:09:55,480
and the population in front of you, check its norms

164
00:09:55,600 --> 00:10:01,879
reliability and validity, and considered language, cultural context ability. When

165
00:10:01,919 --> 00:10:04,759
interpreting a test to always factor in measurement, error and

166
00:10:04,879 --> 00:10:08,559
confidence intervals. Avoid over reliance on a single score by

167
00:10:08,600 --> 00:10:11,960
looking for patterns across the data and straight transparent about

168
00:10:12,000 --> 00:10:17,120
the limits of your interpretation. That's it for now, folks. Oh, actually,

169
00:10:17,120 --> 00:10:18,480
I forgot some practice questions.

170
00:10:18,519 --> 00:10:19,799
Speaker 2: A researcher.

171
00:10:20,840 --> 00:10:25,519
Speaker 1: Develops a new generalized anxiety scale to validated. She administers

172
00:10:25,559 --> 00:10:28,840
the new scale in a structured clinical interview for anxiety

173
00:10:28,879 --> 00:10:31,720
disorders to the same group of participants in the same day,

174
00:10:32,159 --> 00:10:34,799
then compares the results. Which type of validity is it

175
00:10:35,279 --> 00:10:43,759
as a predictive A, content B or concurrent C. If

176
00:10:43,799 --> 00:10:45,080
you said concurrent.

177
00:10:44,720 --> 00:10:45,799
Speaker 2: Validity, you're right.

178
00:10:46,799 --> 00:10:49,480
Speaker 1: This is the validity measured at the same point in time.

179
00:10:49,600 --> 00:10:52,000
Comparing the new instrument against an external standard of the

180
00:10:52,000 --> 00:10:57,120
clinical interview. Question number two, the clinical researcher trains two

181
00:10:57,320 --> 00:11:01,159
raids to an independently code video tape therapy sessions for

182
00:11:01,159 --> 00:11:03,840
a level of therapist empathy. She wants to know how

183
00:11:03,879 --> 00:11:07,279
consistently the two vadors agree with each other which reliability

184
00:11:07,519 --> 00:11:13,440
MU should she calculate, Test retest reliability, parallel forms reliability

185
00:11:13,519 --> 00:11:19,600
or interrator reliability. If you said C, you're right interrator reliability.

186
00:11:19,919 --> 00:11:25,759
Typically calculator using Kappa or interraclass correlation ICC to determine this.

187
00:11:26,159 --> 00:11:27,279
Speaker 2: That's it for now, folks,

