WEBVTT

1
00:00:00.120 --> 00:00:02.919
<v Speaker 1>What's up everyone, and welcome back to the program. In

2
00:00:03.000 --> 00:00:05.799
<v Speaker 1>this episode, we're headed back up to Moscow where we're

3
00:00:05.839 --> 00:00:11.400
<v Speaker 1>diving right back into the unredacted IgG transcripts, and once

4
00:00:11.400 --> 00:00:14.359
<v Speaker 1>again it's Ann Taylor doing the questioning, and this time

5
00:00:14.400 --> 00:00:19.120
<v Speaker 1>the witness is Daniel Hellwig. And like in the previous episodes,

6
00:00:19.920 --> 00:00:23.679
<v Speaker 1>when I say question, that's Ann Taylor and answer is

7
00:00:23.760 --> 00:00:29.359
<v Speaker 1>mister Hellwig, Miss Taylor. The next witness is Daniel Hellwig

8
00:00:29.800 --> 00:00:32.320
<v Speaker 1>the clerk. Do you solemnly swear or affirm that the

9
00:00:32.359 --> 00:00:35.479
<v Speaker 1>testimony you're about to give now before the court will

10
00:00:35.520 --> 00:00:38.039
<v Speaker 1>be the truth, the whole truth, and nothing but the truth.

11
00:00:38.640 --> 00:00:42.039
<v Speaker 1>Mister Hellwig? I do. The court go ahead, Miss Taylor.

12
00:00:42.200 --> 00:00:47.240
<v Speaker 1>Thank you, Daniel Hellwig having been duly sworn testified as follows,

13
00:00:47.799 --> 00:00:51.960
<v Speaker 1>Miss Taylor. Good morning, answer, Good morning. Question. Will you

14
00:00:52.000 --> 00:00:55.920
<v Speaker 1>please state your full name? Answer? My name is Daniel Hellwig.

15
00:00:56.439 --> 00:00:59.159
<v Speaker 1>Question what do you do for a living Answer? I'm

16
00:00:59.200 --> 00:01:02.399
<v Speaker 1>currently the friend director at Inner Mountain Forensics. We're a

17
00:01:02.439 --> 00:01:05.359
<v Speaker 1>five oh one three c nonprofit and our main mission

18
00:01:05.599 --> 00:01:09.799
<v Speaker 1>is pushing forward new and cutting edge DNA technology, assisting

19
00:01:09.840 --> 00:01:14.560
<v Speaker 1>different agencies, law enforcement and otherwise with education, training and

20
00:01:14.680 --> 00:01:18.799
<v Speaker 1>consultation on those. Our main focus right now is to

21
00:01:18.840 --> 00:01:22.400
<v Speaker 1>help fund these cases specifically, and most of our mission

22
00:01:22.480 --> 00:01:28.079
<v Speaker 1>is revolved around forensic investigative genetic genealogy or FIG. Question

23
00:01:28.640 --> 00:01:31.280
<v Speaker 1>when you say help fund these cases as a nonprofit,

24
00:01:31.319 --> 00:01:34.640
<v Speaker 1>what do you mean? Answer a variety of different We

25
00:01:34.840 --> 00:01:37.760
<v Speaker 1>actually allow people to submit cases that they may be

26
00:01:37.840 --> 00:01:41.560
<v Speaker 1>having problems getting the revenue and resources to move it forward,

27
00:01:41.840 --> 00:01:44.920
<v Speaker 1>and then we evaluate it. So this could be We've

28
00:01:45.000 --> 00:01:47.879
<v Speaker 1>worked and tried to assist, and mainly I think most

29
00:01:47.920 --> 00:01:50.799
<v Speaker 1>of our work is in smaller agencies that don't have

30
00:01:50.920 --> 00:01:54.000
<v Speaker 1>nearly as many resources to further their case work with

31
00:01:54.079 --> 00:01:57.799
<v Speaker 1>this technology. But we've done work with law enforcement, medical

32
00:01:57.799 --> 00:02:03.480
<v Speaker 1>examiners offices, and defense, especially the Innocence Project. Question what

33
00:02:03.560 --> 00:02:06.519
<v Speaker 1>did you do before you worked at this nonprofit? Answer?

34
00:02:06.959 --> 00:02:10.439
<v Speaker 1>I've been a forensic DNA for twenty plus years. I

35
00:02:10.479 --> 00:02:14.560
<v Speaker 1>have a bachelor's in biology and chemistry from Viturbo University

36
00:02:14.800 --> 00:02:18.560
<v Speaker 1>and a master's in forensic science from Marshall University. I

37
00:02:18.599 --> 00:02:22.240
<v Speaker 1>started my career at the Armed Forces DNA Identification Laboratory

38
00:02:22.439 --> 00:02:25.879
<v Speaker 1>as an intern there and moved on to various public laboratories.

39
00:02:26.479 --> 00:02:29.520
<v Speaker 1>I worked in forensic DNA at the New Mexico Department

40
00:02:29.560 --> 00:02:32.879
<v Speaker 1>of Public Safety and the Minnesota Bureau of Criminal Apprehension.

41
00:02:33.800 --> 00:02:36.639
<v Speaker 1>I then started in the private sector. I worked for

42
00:02:36.639 --> 00:02:39.680
<v Speaker 1>Sorens and Forensics in Salt Lake City, Utah, in a

43
00:02:39.719 --> 00:02:42.879
<v Speaker 1>variety of different jobs there. Initially I was a DNA

44
00:02:42.960 --> 00:02:47.039
<v Speaker 1>technical leader, essentially the quality manager of the DNA section,

45
00:02:47.520 --> 00:02:52.039
<v Speaker 1>but moved into executive management as the laboratory director. In

46
00:02:52.039 --> 00:02:55.680
<v Speaker 1>twenty nineteen. I began with the inter Mountain Forensics as

47
00:02:55.719 --> 00:02:59.319
<v Speaker 1>a founder. I'll refer to that sometimes as IMF. We

48
00:02:59.319 --> 00:03:02.120
<v Speaker 1>were up until July of last year. Our mission was

49
00:03:02.159 --> 00:03:06.159
<v Speaker 1>not only funding cases, education outreach, but also we established

50
00:03:06.159 --> 00:03:11.439
<v Speaker 1>and operated a fully functioning forensic DNA lab our. Mission

51
00:03:11.560 --> 00:03:14.840
<v Speaker 1>in that realm was to again continue with cutting edge

52
00:03:14.919 --> 00:03:18.919
<v Speaker 1>DNA technology and specifically we sought to operate and utilize

53
00:03:18.960 --> 00:03:23.840
<v Speaker 1>a forensic investigative genetic genealogy support laboratory. Question what does

54
00:03:23.879 --> 00:03:27.759
<v Speaker 1>that mean? Answer? So our goal in this We had

55
00:03:27.759 --> 00:03:30.360
<v Speaker 1>a fully function in laboratory in that we did what

56
00:03:30.520 --> 00:03:34.919
<v Speaker 1>I'll call traditional forensics STRs as mentioned previously, but we

57
00:03:34.960 --> 00:03:39.479
<v Speaker 1>also wanted to implement this new technology forensic investigative genetic

58
00:03:39.520 --> 00:03:44.560
<v Speaker 1>genealogy and laboratory processes behind it. Specifically in this case,

59
00:03:44.599 --> 00:03:50.680
<v Speaker 1>you're talking about forensic snips single nucleotide polymorphisms the court,

60
00:03:51.280 --> 00:03:53.159
<v Speaker 1>so you need to slow down a little bit. You

61
00:03:53.159 --> 00:03:56.039
<v Speaker 1>don't enunciate particularly well and so you need to go

62
00:03:56.120 --> 00:03:59.479
<v Speaker 1>slower the witness. I promise question. Let me see if

63
00:03:59.479 --> 00:04:01.800
<v Speaker 1>I can understand and your work right before the five

64
00:04:01.840 --> 00:04:05.360
<v Speaker 1>oh one C three getting into the forensic investigative genetic

65
00:04:05.400 --> 00:04:08.719
<v Speaker 1>genealogy world. Do I understand that your work was to

66
00:04:08.879 --> 00:04:11.400
<v Speaker 1>do the part all the long steps until you get

67
00:04:11.439 --> 00:04:13.840
<v Speaker 1>to the SNP and then it would be handed off

68
00:04:13.919 --> 00:04:17.079
<v Speaker 1>for a genealogist to do the other part of the work. Answer.

69
00:04:17.439 --> 00:04:21.879
<v Speaker 1>Beyond identifying SNPs, it's to generate and upload file an

70
00:04:22.000 --> 00:04:24.439
<v Speaker 1>end product that would then be handed off to an

71
00:04:24.480 --> 00:04:29.279
<v Speaker 1>investigative genetic genealogist to continue the research. So essentially our

72
00:04:29.399 --> 00:04:33.120
<v Speaker 1>laboratory started had the ability to start from sample from

73
00:04:33.160 --> 00:04:38.040
<v Speaker 1>evidentary item to DNA extraction, which is essentially popping cells open,

74
00:04:38.360 --> 00:04:41.800
<v Speaker 1>pulling DNA out and washing all the residual material away.

75
00:04:42.120 --> 00:04:45.680
<v Speaker 1>And then traditional forensics where you're doing short tandem repeats

76
00:04:45.879 --> 00:04:49.879
<v Speaker 1>repetitive DNA that repeats over and over and you simply

77
00:04:49.920 --> 00:04:53.879
<v Speaker 1>count them to our specific targets which was SNP's where

78
00:04:53.920 --> 00:04:57.279
<v Speaker 1>we did DNA sequencing, looking at all the different letters

79
00:04:57.279 --> 00:04:59.920
<v Speaker 1>within the genome and pulling out the relevant single new

80
00:05:00.040 --> 00:05:05.519
<v Speaker 1>ucleotide polymorphisms that were specifically associated through genealogy, generating through

81
00:05:05.519 --> 00:05:08.720
<v Speaker 1>a pretty extensive process what I'll call an upload file.

82
00:05:09.120 --> 00:05:12.319
<v Speaker 1>These are files that contain all of these SNPs single

83
00:05:12.439 --> 00:05:16.519
<v Speaker 1>nucleotide polymorphisms that contain the information that's needed for upload

84
00:05:16.680 --> 00:05:19.920
<v Speaker 1>into these databases. Okay, I'm going to go back a

85
00:05:19.959 --> 00:05:22.839
<v Speaker 1>little bit with you. So in the traditional str we've

86
00:05:22.879 --> 00:05:26.120
<v Speaker 1>heard some about that today. You said splitting the DNA

87
00:05:26.199 --> 00:05:29.759
<v Speaker 1>open and washing it. Answer. Yes, that's essentially the very

88
00:05:29.800 --> 00:05:33.319
<v Speaker 1>simplified version of a DNA extraction. You can think of

89
00:05:33.319 --> 00:05:37.160
<v Speaker 1>this biological material, these cells as little tiny water filled

90
00:05:37.199 --> 00:05:42.319
<v Speaker 1>balloons with nuclear material with DNA inside of it. DNA extraction,

91
00:05:42.480 --> 00:05:45.879
<v Speaker 1>similar to previous testimony, is just popping those cells open,

92
00:05:46.120 --> 00:05:49.600
<v Speaker 1>pulling the DNA out, and washing all of the cellular garbage.

93
00:05:49.839 --> 00:05:53.360
<v Speaker 1>If you will out to generate an extract, it can

94
00:05:53.399 --> 00:05:56.360
<v Speaker 1>then go down several different pathways depending on what you need.

95
00:05:57.160 --> 00:06:01.680
<v Speaker 1>Traditional forensic testing STRs short tandem repeats is taking that

96
00:06:01.800 --> 00:06:05.360
<v Speaker 1>DNA and looking at repetitive sequences and counting how many

97
00:06:05.439 --> 00:06:08.519
<v Speaker 1>repeats are there. The best way I can explain this

98
00:06:08.959 --> 00:06:11.480
<v Speaker 1>is if you had the word cat and that was

99
00:06:11.480 --> 00:06:14.360
<v Speaker 1>the specific DNA letters that you were looking for. You

100
00:06:14.360 --> 00:06:17.120
<v Speaker 1>could repeat the word cat eight times. That would be

101
00:06:17.160 --> 00:06:20.319
<v Speaker 1>short tandem repeat cat repeat at A times. I would

102
00:06:20.360 --> 00:06:24.480
<v Speaker 1>call that an eight. You have DNA from mom and dad,

103
00:06:24.680 --> 00:06:27.040
<v Speaker 1>so you have two copies of this, So you have

104
00:06:27.120 --> 00:06:29.480
<v Speaker 1>an eight and maybe another eight from mom and dad.

105
00:06:29.800 --> 00:06:32.279
<v Speaker 1>The problem there is that you're actually looking at the

106
00:06:32.319 --> 00:06:35.920
<v Speaker 1>specific letters. So a three letter word like dog would

107
00:06:35.959 --> 00:06:39.839
<v Speaker 1>be completely different. We've got an eight repeat STR and

108
00:06:39.879 --> 00:06:43.480
<v Speaker 1>an eight repeat str that have different sequences but still

109
00:06:43.480 --> 00:06:47.639
<v Speaker 1>have the same short tandem repeat number. SMPS single nucleotide

110
00:06:47.680 --> 00:06:51.800
<v Speaker 1>polymorphism is actually looking at the DNA sequence. So if

111
00:06:51.800 --> 00:06:54.160
<v Speaker 1>we were to look at an SNP, let's say we

112
00:06:54.160 --> 00:06:57.839
<v Speaker 1>were particularly interested in the A and cat, we would

113
00:06:57.839 --> 00:07:01.800
<v Speaker 1>sequence that DNA and then go into the sequence and say,

114
00:07:01.839 --> 00:07:04.360
<v Speaker 1>at this location we have an A, and we have

115
00:07:04.399 --> 00:07:06.680
<v Speaker 1>an A from mom and maybe an A from Dad.

116
00:07:07.399 --> 00:07:10.879
<v Speaker 1>That is the generation of an SNP profile. Same concept,

117
00:07:10.920 --> 00:07:14.000
<v Speaker 1>but we're dialing down to the specific letter. In question

118
00:07:15.079 --> 00:07:18.040
<v Speaker 1>question have you done both kinds of work? Answer? Yes?

119
00:07:18.920 --> 00:07:22.120
<v Speaker 1>In my before IMF I worked at several different laboratories,

120
00:07:22.199 --> 00:07:27.560
<v Speaker 1>all using traditional forensics SDRs and ysdrs. Question if I

121
00:07:27.639 --> 00:07:30.680
<v Speaker 1>understood that part about the water balloon right, It's one

122
00:07:30.759 --> 00:07:34.199
<v Speaker 1>water balloon, same water balloon, same process of popping it open,

123
00:07:34.519 --> 00:07:37.519
<v Speaker 1>stripping out the stuff you don't need, and then it's

124
00:07:37.560 --> 00:07:39.879
<v Speaker 1>where you go from there that makes the difference between

125
00:07:39.879 --> 00:07:44.439
<v Speaker 1>an SDR and an SNP. Ultimately. Answer correct, It's all

126
00:07:44.480 --> 00:07:47.680
<v Speaker 1>associating to the same DNA extract, and that's similar to

127
00:07:47.720 --> 00:07:51.399
<v Speaker 1>what the Idaho State Police Lab does as well. Question so,

128
00:07:51.480 --> 00:07:54.160
<v Speaker 1>the process of getting the DNA water balloon open so

129
00:07:54.199 --> 00:07:57.120
<v Speaker 1>that you can strip things off, does it matter how

130
00:07:57.199 --> 00:08:00.480
<v Speaker 1>you do that? Answer? Well, there's a variety of different

131
00:08:00.519 --> 00:08:03.279
<v Speaker 1>ways to do that, but in forensics there are some

132
00:08:03.319 --> 00:08:06.519
<v Speaker 1>specific and kind of common tools that we use that

133
00:08:06.600 --> 00:08:10.560
<v Speaker 1>the process to get there can vary somewhat. However, it's

134
00:08:10.600 --> 00:08:13.120
<v Speaker 1>a fairly standard practice to use some of the same

135
00:08:13.160 --> 00:08:17.800
<v Speaker 1>tools within the forensic DNA community. Question I think I

136
00:08:17.879 --> 00:08:20.879
<v Speaker 1>understand the repeats as looking for patterns, but the SNP

137
00:08:21.079 --> 00:08:26.600
<v Speaker 1>is looking at the individual characteristic answer individual letter nucleotide.

138
00:08:26.720 --> 00:08:29.199
<v Speaker 1>So if DNA is made up of billions and billions

139
00:08:29.240 --> 00:08:33.480
<v Speaker 1>of letters, STR is looking at repetitive sequences and counting

140
00:08:33.480 --> 00:08:38.000
<v Speaker 1>the repeats, and SNP is looking at specific letters, specific

141
00:08:38.159 --> 00:08:42.480
<v Speaker 1>nucleotides on that genome. Question. All right, I want to

142
00:08:42.519 --> 00:08:45.600
<v Speaker 1>talk about the whole process. How you get from an

143
00:08:45.679 --> 00:08:49.519
<v Speaker 1>SNP to splitting the water balloon open. I'm with you there,

144
00:08:49.720 --> 00:08:52.159
<v Speaker 1>So what do you do after we get that stripped

145
00:08:52.200 --> 00:08:55.960
<v Speaker 1>off answer? If we're going down the path of forensic SNPs.

146
00:08:56.240 --> 00:08:59.039
<v Speaker 1>There's a few different ways you can do that. There

147
00:08:59.120 --> 00:09:02.480
<v Speaker 1>is one way is essentially targeted practice. We're looking at

148
00:09:02.519 --> 00:09:05.600
<v Speaker 1>SNPs in question, and we're going to focus on those

149
00:09:05.759 --> 00:09:10.080
<v Speaker 1>SNPs and do what is called an amplification PCR, basically

150
00:09:10.080 --> 00:09:14.320
<v Speaker 1>a DNA photocopier of that sequence that specific letter and

151
00:09:14.399 --> 00:09:17.320
<v Speaker 1>a little bit around it to specifically target an SNP

152
00:09:17.440 --> 00:09:20.879
<v Speaker 1>in question. The other technique is something more attuned to

153
00:09:20.919 --> 00:09:24.279
<v Speaker 1>the whole genome sequencing. Again, there's a variety of different

154
00:09:24.320 --> 00:09:26.960
<v Speaker 1>ways you can do this, but the idea here is

155
00:09:27.000 --> 00:09:30.039
<v Speaker 1>we're going to sequence the entire genome, the entire length

156
00:09:30.039 --> 00:09:32.960
<v Speaker 1>of DNA on this particular extract, and then we're going

157
00:09:33.000 --> 00:09:36.120
<v Speaker 1>to pull out the SNPs that we need, the relevant

158
00:09:36.120 --> 00:09:39.240
<v Speaker 1>ones that we're looking for. This can be done. I

159
00:09:39.279 --> 00:09:42.879
<v Speaker 1>will refer to that as whole genome sequencing, acknowledging the

160
00:09:42.919 --> 00:09:45.399
<v Speaker 1>fact that there's a variety of different ways you can

161
00:09:45.440 --> 00:09:49.240
<v Speaker 1>accomplish that. Once you sequence the DNA, it comes off

162
00:09:49.399 --> 00:09:53.639
<v Speaker 1>of the instrument in question, and there's several different instruments

163
00:09:53.679 --> 00:09:56.039
<v Speaker 1>you can use, but for the most part, it turns

164
00:09:56.080 --> 00:09:59.080
<v Speaker 1>into raw data, and that raw data is typically found

165
00:09:59.159 --> 00:10:02.799
<v Speaker 1>in a file called the fastq file. This raw data

166
00:10:02.879 --> 00:10:05.879
<v Speaker 1>is massive, it has an insane amount of information, and

167
00:10:05.960 --> 00:10:07.919
<v Speaker 1>it's not refined in a way that we can actually

168
00:10:08.039 --> 00:10:12.200
<v Speaker 1>utilize it. So the steps from taking the raw data

169
00:10:12.240 --> 00:10:15.679
<v Speaker 1>to actually getting that upload file. That end product for

170
00:10:15.840 --> 00:10:21.480
<v Speaker 1>upload into these databases is bioinformatics, essentially a software that

171
00:10:21.519 --> 00:10:24.120
<v Speaker 1>goes into this massive amounts of data and pulls out

172
00:10:24.159 --> 00:10:27.879
<v Speaker 1>the things that we're looking for. Question, is all the

173
00:10:27.919 --> 00:10:31.960
<v Speaker 1>software the same or all bioinformatics programs? Is that all

174
00:10:32.000 --> 00:10:35.159
<v Speaker 1>the same? Answer? No, Each, as far as I can tell,

175
00:10:35.360 --> 00:10:39.799
<v Speaker 1>each laboratory has their own version of bioinformatics that they

176
00:10:39.879 --> 00:10:43.399
<v Speaker 1>use to get the sequence data into usable results. Question

177
00:10:43.799 --> 00:10:46.799
<v Speaker 1>if I understand where we are, We popped open the balloon,

178
00:10:47.120 --> 00:10:50.240
<v Speaker 1>we put the DNA in a sequencer, created a raw

179
00:10:50.320 --> 00:10:53.559
<v Speaker 1>data file called fastq file, and now we need to

180
00:10:53.639 --> 00:10:58.399
<v Speaker 1>do bioinformatics to get an SNP answer. Right question, Okay,

181
00:10:58.440 --> 00:11:02.600
<v Speaker 1>what happens from there? Answer? So in the bioinformatic pathway,

182
00:11:02.840 --> 00:11:05.440
<v Speaker 1>you're doing multiple things. Again, you can think of this

183
00:11:05.519 --> 00:11:08.679
<v Speaker 1>as a DNA sequence that the sequencer has given you

184
00:11:08.679 --> 00:11:13.320
<v Speaker 1>a truckload of information. They're essentially puzzle pieces. The first

185
00:11:13.360 --> 00:11:16.320
<v Speaker 1>thing that we're going to do is reemerging. So we're

186
00:11:16.360 --> 00:11:18.519
<v Speaker 1>going to take all the puzzle pieces that are the

187
00:11:18.559 --> 00:11:20.559
<v Speaker 1>same and we're going to use them to paint the

188
00:11:20.559 --> 00:11:23.519
<v Speaker 1>best picture of that information. You can think of this

189
00:11:23.919 --> 00:11:28.480
<v Speaker 1>as if we're sequencing the DNA this human genome multiple times,

190
00:11:28.960 --> 00:11:33.519
<v Speaker 1>hopefully ten, fifteen, thirty times. Our samples in forensics are

191
00:11:33.559 --> 00:11:35.960
<v Speaker 1>typically difficult, so we don't tend to get that much

192
00:11:35.960 --> 00:11:39.840
<v Speaker 1>information or that level. Essentially, we're trying to sequence this

193
00:11:40.000 --> 00:11:43.440
<v Speaker 1>entire genome multiple times to add more context in what

194
00:11:43.480 --> 00:11:47.840
<v Speaker 1>we're getting e merging. We're talking in DNA sequences that

195
00:11:47.879 --> 00:11:50.720
<v Speaker 1>we have multiple copies of, and we're combining them into

196
00:11:50.759 --> 00:11:54.159
<v Speaker 1>one the best fit for that particular fragment of DNA.

197
00:11:54.879 --> 00:11:57.519
<v Speaker 1>While we still have puzzle pieces to put together. The

198
00:11:57.559 --> 00:12:01.200
<v Speaker 1>next thing is mapping the bioinformatic way will map all

199
00:12:01.240 --> 00:12:04.320
<v Speaker 1>these puzzle pieces and put them into the human genome

200
00:12:04.639 --> 00:12:07.279
<v Speaker 1>in the way that they're supposed to be that takes

201
00:12:07.360 --> 00:12:10.240
<v Speaker 1>multiple fragments and lines them up in the right manner.

202
00:12:10.799 --> 00:12:13.000
<v Speaker 1>Once it's mapped, we are going to start to be

203
00:12:13.039 --> 00:12:15.759
<v Speaker 1>able to find the particular SNPs that we're looking for.

204
00:12:16.480 --> 00:12:18.639
<v Speaker 1>We have a human genome with a whole bunch of

205
00:12:18.639 --> 00:12:21.879
<v Speaker 1>different letters in different positions. We can say, all right,

206
00:12:21.960 --> 00:12:25.320
<v Speaker 1>these are the relevant SNPs that we're looking for. Genealogy

207
00:12:25.360 --> 00:12:29.720
<v Speaker 1>based informative markers. We want to pull these SNPs out,

208
00:12:29.919 --> 00:12:32.039
<v Speaker 1>see what the call is there, and send that up

209
00:12:32.039 --> 00:12:34.679
<v Speaker 1>to an upload file. In some cases we don't have

210
00:12:34.799 --> 00:12:38.120
<v Speaker 1>all of the information. We're missing some sequence data. We've

211
00:12:38.159 --> 00:12:40.759
<v Speaker 1>maybe gotten eighty percent of the genome and most of

212
00:12:40.759 --> 00:12:44.080
<v Speaker 1>the SNPs are missing. There's a process to fill in

213
00:12:44.120 --> 00:12:50.960
<v Speaker 1>those gaps called imputation. Imputation is in most bioinformatic packages. Essentially,

214
00:12:51.000 --> 00:12:53.159
<v Speaker 1>we know that the letters are in front of this

215
00:12:53.200 --> 00:12:56.360
<v Speaker 1>particular spot we're missing, and we know that the letters

216
00:12:56.600 --> 00:13:00.080
<v Speaker 1>are behind what this particular spot that is missing. We

217
00:13:00.120 --> 00:13:02.919
<v Speaker 1>can look at all the human genomes and impute estimate

218
00:13:03.200 --> 00:13:06.039
<v Speaker 1>that most likely the letter in the position that we're

219
00:13:06.039 --> 00:13:10.440
<v Speaker 1>looking for. Once we're through imputation and calls, we're going

220
00:13:10.480 --> 00:13:14.000
<v Speaker 1>to use Bioinformatic package to pull out smps that we're

221
00:13:14.039 --> 00:13:16.639
<v Speaker 1>looking for and put them into the format that will

222
00:13:16.679 --> 00:13:20.840
<v Speaker 1>be usable for upload into these databases. At IMF, we

223
00:13:20.879 --> 00:13:24.759
<v Speaker 1>actually generated two files, one specifically to ged match pro

224
00:13:24.919 --> 00:13:29.399
<v Speaker 1>database and one specific to the family Tree ft DNA database.

225
00:13:30.360 --> 00:13:33.559
<v Speaker 1>In these files, ged match Pro we typically got around

226
00:13:33.559 --> 00:13:36.799
<v Speaker 1>five hundred and fifty thousand plus SNPs that we were

227
00:13:36.799 --> 00:13:40.720
<v Speaker 1>looking for, and the Family Tree DNA database upwards of

228
00:13:40.799 --> 00:13:44.759
<v Speaker 1>six hundred to six hundred and fifty thousand smps. The

229
00:13:44.840 --> 00:13:47.480
<v Speaker 1>end product there though, is a text file or some

230
00:13:47.519 --> 00:13:50.600
<v Speaker 1>sort of upload file which could be in text format

231
00:13:50.679 --> 00:13:54.200
<v Speaker 1>or Excel format, but essentially what it is is a

232
00:13:54.240 --> 00:13:57.080
<v Speaker 1>file with all of these smps the calls at the

233
00:13:57.200 --> 00:13:59.840
<v Speaker 1>SNP locations in a format that allows them to be

234
00:14:00.080 --> 00:14:03.519
<v Speaker 1>uploaded into the database. At this point we hand it

235
00:14:03.559 --> 00:14:06.919
<v Speaker 1>off to the investigative genetic genealogist to do the research.

236
00:14:07.799 --> 00:14:11.919
<v Speaker 1>Question the upload databases you mentioned to ged match pro

237
00:14:12.200 --> 00:14:16.480
<v Speaker 1>and Family Tree DNA answer, that's correct, question. Why did

238
00:14:16.519 --> 00:14:20.399
<v Speaker 1>you mention those answer? Those are the two databases that

239
00:14:20.440 --> 00:14:24.519
<v Speaker 1>allow law enforcement searches. They have specific terms of services

240
00:14:24.559 --> 00:14:27.279
<v Speaker 1>that give a portion of that database, those that have

241
00:14:27.360 --> 00:14:30.559
<v Speaker 1>consented to do so to law enforcement to search for

242
00:14:30.759 --> 00:14:34.879
<v Speaker 1>human remains in some cases and in criminal cases. All Right,

243
00:14:34.879 --> 00:14:36.519
<v Speaker 1>we're going to wrap up right there, and in the

244
00:14:36.559 --> 00:14:38.919
<v Speaker 1>next episode dealing with the topic, we're going to pick

245
00:14:39.000 --> 00:14:42.399
<v Speaker 1>up with question so in this lab process to get

246
00:14:42.480 --> 00:14:45.200
<v Speaker 1>to the SNP. If you'd like to contact me, you

247
00:14:45.240 --> 00:14:48.039
<v Speaker 1>can do that at Bobby Kapuchi at ProtonMail dot com.

248
00:14:48.159 --> 00:14:51.519
<v Speaker 1>That's bo b b Y c ap u Cci at

249
00:14:51.519 --> 00:14:54.600
<v Speaker 1>ProtonMail dot com, or if you prefer, you can find

250
00:14:54.600 --> 00:15:00.440
<v Speaker 1>me on x at Bobby Underscore c ap u Cci.

251
00:15:00.919 --> 00:15:02.720
<v Speaker 1>All of the links that go with this episode can

252
00:15:02.799 --> 00:15:04.519
<v Speaker 1>be found in the description box.
