bacon · 2026-10-08T12:14:26 · BACoN 0.3.8 · last run took 2.1 s, not counting reused steps
4/4
samples assembled
0
assembled with a note
30,000 bp
reference
reference.fasta, 1 sequence(s)
20
SNP sites
5 genomes, 4 distinct
Metadata from metadata.tsv: 4 columns (group, year, origin, note); 4 of 4 samples have a value. The figures are coloured by group (a colour and a shape for each value, the same in every figure and table).
Samples
faileddepth below 20x or length outside 0.8–1.2x the referenceN basesClick a column to sort.
Sample
group
year
origin
note (metadata)
Status
Raw reads
Baited reads
Baited %
Filtered reads
Read N50
Depth
Contigs
Circular
Length
× ref
N bases
Note
alpha
A
2021
site 1
Identical to the reference by design (no planted SNP)
ok
282
235
85.425
218
7,640
47.6
1
NA
30,000
1.000
0
beta
A
2021
site 1
Carries six planted SNPs, in coding sequences, an intron and a spacer
ok
323
253
75.767
231
6,815
47.5
1
NA
30,000
1.000
0
delta
2023
site 2
Ten SNPs of its own, including a nonsense change and one in the pseudogene
ok
308
251
81.665
232
7,031
47.6
1
NA
30,000
1.000
0
gamma
B
2022
site 2
Shares beta's six SNPs, plus four of its own (one in a tRNA)
ok
304
247
80.764
224
6,716
47.7
1
NA
30,000
1.000
0
Figure 1. Estimated depth of the filtered reads over the reference, per sample. The line is the 20x threshold of the table's depth flag; samples below it are labelled. Failed samples are listed without a bar. Hover a bar for the read counts and the note. Each bar has the colour of the sample's group, whose marker is before the name, as in every figure and table (legend under the heatmap; a hollow circle and a grey bar: no value).
Figure 2. N bases in each assembly, largest first (the table flags every assembly with N bases; the 5 largest counts are labelled). Hover a bar for the fraction of the assembly. Colours and markers as in Figure 1.
Tree
Figure 3. SKA2 SNPs; FastTree, midpoint-rooted and ladderized. Numbers on the internal branches are supports. Markers and the muted text after the names give each genome's group: each value has a colour and a shape of its own, as in the legend of the heatmap below (a hollow circle: no value). The scale bar is in substitutions per SNP site, with the equivalent number of SNPs.
SNP distances
Figure 4. Pairwise SNP distances, in tree order. Colour classes are spread on a log scale over the range of the distances (legend); hover a cell for the exact distance of its pair. The inner grey bands on both axes and the blocks on the right mark the groups of identical genomes; the outer bands give each genome's group as a marker, a colour and a shape for each value (legend; a hollow circle: no value).
Identical genomes
No SNP between any two genomes of a group (positions with N or a gap are not compared).
group 1 (2): Reference, alpha
Genomes of each group by group (the reference has no value):
A
B
no value
Total
group 1
1
–
1
2
not in a group
1
1
1
3
5 genomes, 4 distinct at the SNP sites compared.
Genome map
Figure 5. SNP positions along the reference (20 records of the VCF, 0 with a missing call in at least one of the 4 genomes). Hover a tick for the position, the alleles and the number of genomes with the alternate allele, with the gene, its context and the effect of the SNP. Ticks closer than 31 bp are merged; hover shows the SNPs of a tick. Genes from reference.gb (23): the + strand above the centre line, the − strand below; hover a gene for its name and coordinates. Genes with 2 or more SNPs are labelled. The band above the genes shows the LSC/IRb/SSC/IRa regions from the annotated inverted repeats. No N track: none of the 4 templated assemblies has an N base.
SNPs
20 SNPs on the annotated sequences: 13 in the LSC, 7 in the SSC. 12 in coding sequences (5 synonymous, 6 missense, 1 nonsense); 1 in introns; 2 in tRNA genes; 1 in pseudogenes; 4 intergenic. Genes with the most SNPs: orf01 (2), orf04 (2). Click a column to sort.
Position
REF>ALT
Region
Gene
Context
Codon
Amino acid
Effect
ALT genomes
Missing
573
T>C
LSC
orf01
CDS
AAT>AAC
N91N
synonymous
2
0
931
A>C
LSC
orf01
CDS
AAC>CAC
N211H
missense
1
0
2,440
G>T
LSC
orf02
CDS
CGC>AGC
R121S
missense
2
0
3,457
C>T
LSC
orf03
CDS
CGA>TGA
R153*
nonsense
1
0
4,851
C>A
LSC
orf04
intron
–
–
–
2
0
5,400
A>C
LSC
orf04
CDS
AAA>CAA
K234Q
missense
2
0
6,258
A>G
LSC
orf05
CDS
TTT>TTC
F81F
synonymous
1
0
7,301
A>C
LSC
orf06
CDS
ATA>CTA
I201L
missense
1
0
8,550
G>A
LSC
orf03-ps
pseudogene
–
–
–
1
0
9,260
T>G
LSC
orf07
CDS
AAT>CAT
N181H
missense
1
0
10,038
A>G
LSC
trnB-sim
tRNA
–
–
–
1
0
12,000
A>T
LSC
–
intergenic between orf09 and orf10
–
–
–
1
0
15,000
T>G
LSC
–
intergenic between orf11 and rrn16-sim
–
–
–
2
0
19,843
G>A
SSC
orf12
CDS
GGG>GGA
G181G
synonymous
1
0
21,148
A>G
SSC
orf13
CDS
GAT>GAC
D251D
synonymous
2
0
22,138
G>C
SSC
trnC-sim
tRNA
–
–
–
1
0
22,581
A>C
SSC
orf14
CDS
ATA>CTA
I61L
missense
1
0
23,500
C>A
SSC
–
intergenic between orf14 and orf15
–
–
–
1
0
24,448
A>G
SSC
orf15
CDS
TTT>TTC
F151F
synonymous
1
0
26,000
A>T
SSC
–
intergenic between orf15 and trnD-sim
–
–
–
1
0
Methods
Reads aligning to the reference (reference.fasta, 30,000 bp) were extracted with minimap2 2.31-r1302 (-x map-ont). Reads shorter than 500 bp were discarded and the best 95% were kept, up to 100x of the reference length, with Filtlong 0.3.1. Each sample was assembled by reference-guided consensus: reads were aligned to the reference with minimap2 and the consensus called with samtools consensus 1.24 (-X r10.4_sup, minimum depth 3; positions with less support, or where the reads disagree, are N). SNPs were identified with SKA2 0.5.1 from split 31-mers present in all genomes (core SNPs). Pairwise SNP distances count the positions where both genomes have a nucleotide and they differ. The SNPs of each genome relative to the reference were written to a VCF file by mapping the split k-mers to the reference with ska map 0.5.1. A tree was built on the SNP alignment with FastTree 2.2.0 (GTR, SH-like supports from 100 resamples) and rooted at its midpoint. Genes were read from the annotation reference.gb; the effect of each SNP on the coding sequences (codon and amino-acid change) was derived by BACoN with translation table 11. The LSC/IRb/SSC/IRa regions were derived from the annotated inverted repeats. The analysis was run with BACoN 0.3.8 (https://github.com/duceppemo/BACoN).