1
00:00:01,480 --> 00:00:03,399
Speaker 1: Welcome back today. We're going to be looking at assessment

2
00:00:03,439 --> 00:00:06,679
and diagnosis, So psychometric concepts we're going to be looking

3
00:00:06,679 --> 00:00:13,080
first is reliability and validity. Reliability refers to consistency of

4
00:00:13,119 --> 00:00:15,519
a measure, and validity refers to whether it measures what

5
00:00:15,560 --> 00:00:20,519
it's supposed to measure. Key types of reliability include test retest,

6
00:00:20,719 --> 00:00:25,359
internal consistency. Such things are such as Cronback's alpha that

7
00:00:25,399 --> 00:00:30,120
we'll talk about later in split half inturrator in parallel forms.

8
00:00:30,480 --> 00:00:34,520
Key types of validity include content criterion, concurrent, predictive, and

9
00:00:34,600 --> 00:00:38,000
construct and we'll look at what these mean later. Reliability

10
00:00:38,079 --> 00:00:42,159
is necessary for validity, but it's not sufficient by itself.

11
00:00:43,000 --> 00:00:45,719
Standard error of measurement will look at that and confidence

12
00:00:45,759 --> 00:00:51,200
intervals as well, So let's get started on this reliability.

13
00:00:52,920 --> 00:00:58,320
A reliable test produces similar results under consistent conditions. It

14
00:00:58,359 --> 00:01:01,759
answers the question how dependable is this measure. One way

15
00:01:01,759 --> 00:01:05,200
to look at to test it is test retest measure reliability.

16
00:01:05,480 --> 00:01:09,959
It measures temporal stability. So over time, administer the same

17
00:01:10,040 --> 00:01:12,640
test to the same group at two points in time,

18
00:01:13,120 --> 00:01:15,359
they correlate to two sets of scores, and a high

19
00:01:15,439 --> 00:01:19,599
correlation equals strong reliability over time. This is best for

20
00:01:19,680 --> 00:01:22,400
traits assumed to be stable, like intelligence, but not for

21
00:01:22,480 --> 00:01:26,400
mood or anxiety, which are much more transient. An example

22
00:01:26,439 --> 00:01:29,280
of an assessment would be the ways WAIS scores one

23
00:01:29,280 --> 00:01:33,120
month apart should remain consistent unless a major event intervenes.

24
00:01:34,480 --> 00:01:38,560
Another way of testing reliability is internal consistency measures how

25
00:01:38,599 --> 00:01:41,519
well items or questions on a test measure the same

26
00:01:41,680 --> 00:01:47,920
construct best for tests with homogeneous content like a depression inventory.

27
00:01:49,480 --> 00:01:54,480
Main indices include Kronback's alpha so average correlation among items

28
00:01:55,000 --> 00:01:57,560
split half, which correlates scores from two halves of the

29
00:01:57,599 --> 00:02:01,599
same tests odd versus even items, the KR twenty, which

30
00:02:01,640 --> 00:02:05,280
is for dichotomous items true or false type things. Rule

31
00:02:05,319 --> 00:02:09,000
of thumb and alpha greater or equal to seventy is

32
00:02:09,039 --> 00:02:12,199
acceptable for early research, but for clinical settings or should

33
00:02:12,199 --> 00:02:15,039
be higher than point eight zero. Point nine zero may

34
00:02:15,080 --> 00:02:19,319
indicate redundancy. Inter Radar reliability is another way to measure

35
00:02:19,360 --> 00:02:24,120
agreement between raiders scoring subjective responses. Use them behavioral observations

36
00:02:24,240 --> 00:02:31,199
or diagnostic interviews, or the psychoanalytic projective tests calculated using

37
00:02:31,400 --> 00:02:37,319
Coen's Kappa for categorical decisions and in intra class correlation

38
00:02:37,439 --> 00:02:43,000
for continuous scales. The parallel forms reliability measure is creates

39
00:02:43,039 --> 00:02:45,439
two versions of a test measuring the same construct. The

40
00:02:45,520 --> 00:02:47,919
minister both to the same group and correlate the results

41
00:02:48,639 --> 00:02:51,599
ideal but really rare due to difficulty in making truly

42
00:02:51,680 --> 00:02:56,120
equivalent forms, but it can be helpful with practiced effects

43
00:02:56,159 --> 00:02:59,680
are concerned validity. What is this? It's accuracy of what

44
00:02:59,759 --> 00:03:02,680
is being measured. So validity is about truth. Does the

45
00:03:02,800 --> 00:03:06,719
test measure what it claims to measure? And let's look

46
00:03:06,719 --> 00:03:10,520
at the different types of measuring validity for a test.

47
00:03:11,280 --> 00:03:14,240
What is content validity? How well test items represent the

48
00:03:14,280 --> 00:03:17,639
full range of the construct judged by expert review, not

49
00:03:17,719 --> 00:03:23,840
statistics important for achievement, test, clinical symptom checklists, and diagnostic criteria. Example,

50
00:03:23,879 --> 00:03:26,800
a depression inventory missing items on sleep or appetite has

51
00:03:26,840 --> 00:03:33,560
poor content validity. Criterion validity assesses the relationship between test

52
00:03:33,560 --> 00:03:37,680
scores and an external benchmark concurrent validity. It correlates with

53
00:03:37,759 --> 00:03:44,800
current performance. Example, depression inventory scores correlate with clinical interview diagnosis.

54
00:03:45,719 --> 00:03:49,560
That's a concurrent validity. So again you're assessing the relationship

55
00:03:49,560 --> 00:03:55,560
between test scores and an external benchmark criteria. You're also

56
00:03:55,599 --> 00:03:59,599
looking for predictive validity under the criterion category category correlates

57
00:03:59,639 --> 00:04:06,479
with future performance like GRE scores predicting GPA scores in college.

58
00:04:07,039 --> 00:04:09,680
Effect size is going to be important, so point one

59
00:04:09,759 --> 00:04:12,400
zero is small, point three to zero is moderate, anything

60
00:04:12,520 --> 00:04:16,120
above point five zero is a strong prediction. The next

61
00:04:16,160 --> 00:04:19,199
type of validity you can measure is construct validity. This

62
00:04:19,279 --> 00:04:21,720
one's goods. This one kind of gets confusing a little bit,

63
00:04:22,000 --> 00:04:24,399
but let's look at it showing that the test accurately

64
00:04:24,480 --> 00:04:30,639
reflects the theoretical concept convergent. Convergent validity correlates positively with

65
00:04:30,720 --> 00:04:34,079
measures of the same construct. So you're creating a new

66
00:04:34,079 --> 00:04:37,839
anxiety scale, how does it correlate with a GAD seven?

67
00:04:38,639 --> 00:04:41,920
A discriminate validity is the opposite. It correlates weekly or

68
00:04:41,959 --> 00:04:44,959
not at all, with unrelated constructs. In other words, if

69
00:04:44,959 --> 00:04:48,000
you're creating a depression scale, it should not strongly correlate

70
00:04:48,040 --> 00:04:53,759
with IQ. Construct validity, again, often evolves over multiple studies

71
00:04:53,839 --> 00:04:59,560
using factor analysis, theoretical modeling, and multi trade multimethm matrices.

72
00:05:01,000 --> 00:05:03,319
Face validity is the final one. Do the items look

73
00:05:03,360 --> 00:05:06,519
like they measure what they're supposed to. It's not scientific required.

74
00:05:06,560 --> 00:05:10,360
It affects clients, trust and tests engagement. A test with

75
00:05:10,480 --> 00:05:13,439
high face validity may be easier to fake or manipulate.

76
00:05:14,560 --> 00:05:18,959
Reliability versus validity. Tests must be reliable to be valid,

77
00:05:19,399 --> 00:05:22,399
but reliability alone does not ensure validity. You could be validity,

78
00:05:22,439 --> 00:05:25,600
you could be reliably wrong. Think of a dartboard. Tight

79
00:05:25,639 --> 00:05:29,800
cluster off center equals reliable but not valid, scattered hits

80
00:05:29,839 --> 00:05:32,560
near the bulls eyes low reliability, and tight cluster on

81
00:05:32,639 --> 00:05:37,519
the bulls eyes high reliability and validity. Standard of measurement

82
00:05:37,560 --> 00:05:41,360
of error and confidence intervals. No test score is exact,

83
00:05:41,600 --> 00:05:44,319
and the standard error of measurement smates how much a

84
00:05:44,319 --> 00:05:47,680
person's observed score might deviate from their true score. The

85
00:05:47,720 --> 00:05:52,120
formula is SEM equals SD times the square root of

86
00:05:52,199 --> 00:05:56,680
one minus reliability coefficient. The smaller the SEM means more

87
00:05:56,800 --> 00:06:01,800
precise measurements. Confidence intervals use SEM to create a range

88
00:06:01,839 --> 00:06:05,759
around the observed score. So sixty eight percent confidence interval

89
00:06:06,160 --> 00:06:11,160
equals plus or minus one SEM. Ninety five percent is

90
00:06:11,199 --> 00:06:14,920
plus or minus two SEM. So if a client scores

91
00:06:14,920 --> 00:06:17,240
one hundred on a nine Q test with an SCM

92
00:06:17,319 --> 00:06:21,959
equaling three, and the ninety five percent confidence sivitable would

93
00:06:21,959 --> 00:06:26,720
be one Well November ninety five percent confidence and of

94
00:06:26,720 --> 00:06:29,120
what equals plus or minus two sem. So with an

95
00:06:29,160 --> 00:06:32,519
SEM of three, you're looking at negative six plus six,

96
00:06:32,600 --> 00:06:34,600
so the range will become a ninety four to one

97
00:06:34,639 --> 00:06:37,959
oh six. If the client scores one hundred. You wouldn't

98
00:06:37,959 --> 00:06:40,480
make a rigid clinical adjustment based on a single point

99
00:06:40,519 --> 00:06:46,759
though factors affecting reliability, coefficients or test lengths. Longer tests

100
00:06:46,759 --> 00:06:51,399
have often higher reliability, more data, less noise, item homogenety,

101
00:06:51,519 --> 00:06:56,680
more consistent content improves internal consistency, time between administrations, Shorter

102
00:06:56,759 --> 00:07:02,480
intervals improved test retest reliability, respond variability, more variability between

103
00:07:02,519 --> 00:07:09,800
subjects improves reliability. Testing conditions, noise distractions, unclear instructions reduce reliability.

104
00:07:10,720 --> 00:07:13,399
So how do you look at this clinically? Well, you

105
00:07:13,480 --> 00:07:16,160
use test selections, so you use tests with strong reliability

106
00:07:16,160 --> 00:07:19,199
and validity evidence for the target population. Check test manuals

107
00:07:19,199 --> 00:07:25,600
for norms, sample characteristics and limitations. Avoid tests for nondiverse populations.

108
00:07:25,600 --> 00:07:29,240
Interpreting scores always report confidence intervals, don't rely on single

109
00:07:29,279 --> 00:07:34,519
item scores. For complex constructs, use multiple data sources, assessment

110
00:07:34,560 --> 00:07:37,879
tools and diagnoses. Use structured interviews like SKID with stronger

111
00:07:38,120 --> 00:07:42,519
strong innervator reliability for diagnostic clarity. Select symptom inventories with

112
00:07:42,600 --> 00:07:47,160
high internal consistency and predictive validity ensure tools align with

113
00:07:47,240 --> 00:07:51,959
DSM criteria and cultural presentation of symptoms. Next time, we'll

114
00:07:51,959 --> 00:07:55,480
look at the intelligence assessment instruments and theories about them.

115
00:07:55,639 --> 00:07:56,399
That's it for now

