BACoN report

bacon · 2026-10-08T12:14:26 · BACoN 0.3.8 · last run took 2.1 s, not counting reused steps

4/4
samples assembled
0
assembled with a note
30,000 bp
reference
reference.fasta, 1 sequence(s)
20
SNP sites
5 genomes, 4 distinct

Metadata from metadata.tsv: 4 columns (group, year, origin, note); 4 of 4 samples have a value. The figures are coloured by group (a colour and a shape for each value, the same in every figure and table).

Samples

faileddepth below 20x or length outside 0.8–1.2x the referenceN basesClick a column to sort.

Samplegroupyearoriginnote (metadata)StatusRaw readsBaited readsBaited %Filtered readsRead N50DepthContigsCircularLength× refN basesNote
alphaA2021site 1Identical to the reference by design (no planted SNP)ok28223585.4252187,64047.61NA30,0001.0000
betaA2021site 1Carries six planted SNPs, in coding sequences, an intron and a spacerok32325375.7672316,81547.51NA30,0001.0000
delta2023site 2Ten SNPs of its own, including a nonsense change and one in the pseudogeneok30825181.6652327,03147.61NA30,0001.0000
gammaB2022site 2Shares beta's six SNPs, plus four of its own (one in a tRNA)ok30424780.7642246,71647.71NA30,0001.0000
Depth after filtering, per sample0x20x40x20x flaggammagamma: group Bgamma: 47.7x depth after filtering; 224 reads, N50 6,716alphaalpha: group Aalpha: 47.6x depth after filtering; 218 reads, N50 7,640deltadelta: no groupdelta: 47.6x depth after filtering; 232 reads, N50 7,031betabeta: group Abeta: 47.5x depth after filtering; 231 reads, N50 6,815

Figure 1. Estimated depth of the filtered reads over the reference, per sample. The line is the 20x threshold of the table's depth flag; samples below it are labelled. Failed samples are listed without a bar. Hover a bar for the read counts and the note. Each bar has the colour of the sample's group, whose marker is before the name, as in every figure and table (legend under the heatmap; a hollow circle and a grey bar: no value).

N bases per assembly0alphaalpha: group Aalpha: 0 N bases of 30,000 bp (0.00%)betabeta: group Abeta: 0 N bases of 30,000 bp (0.00%)deltadelta: no groupdelta: 0 N bases of 30,000 bp (0.00%)gammagamma: group Bgamma: 0 N bases of 30,000 bp (0.00%)

Figure 2. N bases in each assembly, largest first (the table flags every assembly with N bases; the 5 largest counts are labelled). Hover a bar for the fraction of the assembly. Colours and markers as in Figure 1.

Tree

Treedelta: no groupdeltaReference: no groupReference reference.fastaalpha: group Aalpha Abeta: group Abeta Agamma: group Bgamma B1.0000.2 substitutions per site (about 4 SNPs)

Figure 3. SKA2 SNPs; FastTree, midpoint-rooted and ladderized. Numbers on the internal branches are supports. Markers and the muted text after the names give each genome's group: each value has a colour and a shape of its own, as in the legend of the heatmap below (a hollow circle: no value). The scale bar is in substitutions per SNP site, with the equivalent number of SNPs.

SNP distances

Pairwise SNP distancesdelta: no groupdelta: no groupdeltadeltaReference: group 1Reference: group 1Reference: no groupReference: no groupReferenceReferencealpha: group 1alpha: group 1alpha: group Aalpha: group Aalphaalphabeta: group Abeta: group Abetabetagamma: group Bgamma: group Bgammagammadelta – delta: 0 SNPs0Reference – Reference: 0 SNPs0Reference – alpha: 0 SNPs0alpha – Reference: 0 SNPs0alpha – alpha: 0 SNPs0beta – beta: 0 SNPs0gamma – gamma: 0 SNPs0beta – gamma: 4 SNPs4gamma – beta: 4 SNPs4delta – Reference: 10 SNPs10delta – alpha: 10 SNPs10delta – beta: 16 SNPs16delta – gamma: 20 SNPs20Reference – delta: 10 SNPs10Reference – beta: 6 SNPs6Reference – gamma: 10 SNPs10alpha – delta: 10 SNPs10alpha – beta: 6 SNPs6alpha – gamma: 10 SNPs10beta – delta: 16 SNPs16beta – Reference: 6 SNPs6beta – alpha: 6 SNPs6gamma – delta: 20 SNPs20gamma – Reference: 10 SNPs10gamma – alpha: 10 SNPs10group 1 (2)group 1 (2)SNPs:01–234–56–20group:A (2)B (1)no value (2)

Figure 4. Pairwise SNP distances, in tree order. Colour classes are spread on a log scale over the range of the distances (legend); hover a cell for the exact distance of its pair. The inner grey bands on both axes and the blocks on the right mark the groups of identical genomes; the outer bands give each genome's group as a marker, a colour and a shape for each value (legend; a hollow circle: no value).

Identical genomes

No SNP between any two genomes of a group (positions with N or a gap are not compared).

Genomes of each group by group (the reference has no value):

ABno valueTotal
group 11–12
not in a group1113

5 genomes, 4 distinct at the SNP sites compared.

Genome map

SNP positions along the referenceorganelle 30,000 bp0 kb5 kb10 kb15 kb20 kb25 kb30 kbLSC: 1–16,000 (16,000 bp)LSC 16,000 bpIRb: 16,001–19,000 (3,000 bp)IRb 3,000 bpSSC: 19,001–27,000 (8,000 bp)SSC 8,000 bpIRa: 27,001–30,000 (3,000 bp)IRa 3,000 bp+−orf01: protein-coding gene, + strand, 301–1,200 (900 bp), simulated protein 1; 2 SNPstrnA-sim: tRNA gene, + strand, 1,401–1,475 (75 bp), simulated tRNA Aorf02: protein-coding gene, − strand, 1,601–2,800 (1,200 bp), simulated protein 2; 1 SNPorf03: protein-coding gene, + strand, 3,001–3,900 (900 bp), simulated protein 3; 1 SNPorf04: protein-coding gene, + strand, 4,201–5,699 (1,499 bp), simulated protein 4; 2 SNPsorf05: protein-coding gene, − strand, 5,901–6,500 (600 bp), simulated protein 5; 1 SNPorf06: protein-coding gene, + strand, 6,701–8,200 (1,500 bp), simulated protein 6; 1 SNPorf03-ps: pseudogene, + strand, 8,401–8,700 (300 bp); 1 SNPorf07: protein-coding gene, − strand, 8,901–9,800 (900 bp), simulated protein 7; 1 SNPtrnB-sim: tRNA gene, − strand, 10,001–10,075 (75 bp), simulated tRNA B; 1 SNPorf08: protein-coding gene, + strand, 10,201–11,400 (1,200 bp), simulated protein 8orf09: protein-coding gene, + strand, 11,601–11,900 (300 bp), simulated protein 9orf10: protein-coding gene, − strand, 12,101–13,300 (1,200 bp), simulated protein 10orf11: protein-coding gene, + strand, 13,501–14,400 (900 bp), simulated protein 11rrn16-sim: rRNA gene, + strand, 16,301–17,800 (1,500 bp), simulated rRNA 16trnD-sim: tRNA gene, + strand, 18,001–18,075 (75 bp), simulated tRNA Dorf12: protein-coding gene, + strand, 19,301–20,200 (900 bp), simulated protein 12; 1 SNPorf13: protein-coding gene, − strand, 20,401–21,900 (1,500 bp), simulated protein 13; 1 SNPtrnC-sim: tRNA gene, + strand, 22,101–22,175 (75 bp), simulated tRNA C; 1 SNPorf14: protein-coding gene, + strand, 22,401–23,000 (600 bp), simulated protein 14; 1 SNPorf15: protein-coding gene, − strand, 24,001–24,900 (900 bp), simulated protein 15; 1 SNPtrnD-sim: tRNA gene, − strand, 27,926–28,000 (75 bp), simulated tRNA Drrn16-sim: rRNA gene, − strand, 28,201–29,700 (1,500 bp), simulated rRNA 16orf01orf01: protein-coding gene, + strand, 301–1,200 (900 bp), simulated protein 1; 2 SNPsorf04orf04: protein-coding gene, + strand, 4,201–5,699 (1,499 bp), simulated protein 4; 2 SNPsSNPsorganelle:573 T>C: alternate allele in 2 of 4 genomes; called in every genome. LSC · orf01 (CDS) · synonymous N91N (AAT>AAC)organelle:931 A>C: alternate allele in 1 of 4 genomes; called in every genome. LSC · orf01 (CDS) · missense N211H (AAC>CAC)organelle:2,440 G>T: alternate allele in 2 of 4 genomes; called in every genome. LSC · orf02 (CDS) · missense R121S (CGC>AGC)organelle:3,457 C>T: alternate allele in 1 of 4 genomes; called in every genome. LSC · orf03 (CDS) · nonsense R153* (CGA>TGA)organelle:4,851 C>A: alternate allele in 2 of 4 genomes; called in every genome. LSC · orf04 (intron)organelle:5,400 A>C: alternate allele in 2 of 4 genomes; called in every genome. LSC · orf04 (CDS) · missense K234Q (AAA>CAA)organelle:6,258 A>G: alternate allele in 1 of 4 genomes; called in every genome. LSC · orf05 (CDS) · synonymous F81F (TTT>TTC)organelle:7,301 A>C: alternate allele in 1 of 4 genomes; called in every genome. LSC · orf06 (CDS) · missense I201L (ATA>CTA)organelle:8,550 G>A: alternate allele in 1 of 4 genomes; called in every genome. LSC · orf03-ps (pseudogene)organelle:9,260 T>G: alternate allele in 1 of 4 genomes; called in every genome. LSC · orf07 (CDS) · missense N181H (AAT>CAT)organelle:10,038 A>G: alternate allele in 1 of 4 genomes; called in every genome. LSC · trnB-sim (tRNA)organelle:12,000 A>T: alternate allele in 1 of 4 genomes; called in every genome. LSC · intergenic between orf09 and orf10organelle:15,000 T>G: alternate allele in 2 of 4 genomes; called in every genome. LSC · intergenic between orf11 and rrn16-simorganelle:19,843 G>A: alternate allele in 1 of 4 genomes; called in every genome. SSC · orf12 (CDS) · synonymous G181G (GGG>GGA)organelle:21,148 A>G: alternate allele in 2 of 4 genomes; called in every genome. SSC · orf13 (CDS) · synonymous D251D (GAT>GAC)organelle:22,138 G>C: alternate allele in 1 of 4 genomes; called in every genome. SSC · trnC-sim (tRNA)organelle:22,581 A>C: alternate allele in 1 of 4 genomes; called in every genome. SSC · orf14 (CDS) · missense I61L (ATA>CTA)organelle:23,500 C>A: alternate allele in 1 of 4 genomes; called in every genome. SSC · intergenic between orf14 and orf15organelle:24,448 A>G: alternate allele in 1 of 4 genomes; called in every genome. SSC · orf15 (CDS) · synonymous F151F (TTT>TTC)organelle:26,000 A>T: alternate allele in 1 of 4 genomes; called in every genome. SSC · intergenic between orf15 and trnD-simSNP called in every genomeSNP with a missing call in some genomesprotein-coding genetRNA, rRNApseudogene

Figure 5. SNP positions along the reference (20 records of the VCF, 0 with a missing call in at least one of the 4 genomes). Hover a tick for the position, the alleles and the number of genomes with the alternate allele, with the gene, its context and the effect of the SNP. Ticks closer than 31 bp are merged; hover shows the SNPs of a tick. Genes from reference.gb (23): the + strand above the centre line, the − strand below; hover a gene for its name and coordinates. Genes with 2 or more SNPs are labelled. The band above the genes shows the LSC/IRb/SSC/IRa regions from the annotated inverted repeats. No N track: none of the 4 templated assemblies has an N base.

SNPs

20 SNPs on the annotated sequences: 13 in the LSC, 7 in the SSC. 12 in coding sequences (5 synonymous, 6 missense, 1 nonsense); 1 in introns; 2 in tRNA genes; 1 in pseudogenes; 4 intergenic. Genes with the most SNPs: orf01 (2), orf04 (2). Click a column to sort.

PositionREF>ALTRegionGeneContextCodonAmino acidEffectALT genomesMissing
573T>CLSCorf01CDSAAT>AACN91Nsynonymous20
931A>CLSCorf01CDSAAC>CACN211Hmissense10
2,440G>TLSCorf02CDSCGC>AGCR121Smissense20
3,457C>TLSCorf03CDSCGA>TGAR153*nonsense10
4,851C>ALSCorf04intron–––20
5,400A>CLSCorf04CDSAAA>CAAK234Qmissense20
6,258A>GLSCorf05CDSTTT>TTCF81Fsynonymous10
7,301A>CLSCorf06CDSATA>CTAI201Lmissense10
8,550G>ALSCorf03-pspseudogene–––10
9,260T>GLSCorf07CDSAAT>CATN181Hmissense10
10,038A>GLSCtrnB-simtRNA–––10
12,000A>TLSC–intergenic between orf09 and orf10–––10
15,000T>GLSC–intergenic between orf11 and rrn16-sim–––20
19,843G>ASSCorf12CDSGGG>GGAG181Gsynonymous10
21,148A>GSSCorf13CDSGAT>GACD251Dsynonymous20
22,138G>CSSCtrnC-simtRNA–––10
22,581A>CSSCorf14CDSATA>CTAI61Lmissense10
23,500C>ASSC–intergenic between orf14 and orf15–––10
24,448A>GSSCorf15CDSTTT>TTCF151Fsynonymous10
26,000A>TSSC–intergenic between orf15 and trnD-sim–––10

Methods

Reads aligning to the reference (reference.fasta, 30,000 bp) were extracted with minimap2 2.31-r1302 (-x map-ont). Reads shorter than 500 bp were discarded and the best 95% were kept, up to 100x of the reference length, with Filtlong 0.3.1. Each sample was assembled by reference-guided consensus: reads were aligned to the reference with minimap2 and the consensus called with samtools consensus 1.24 (-X r10.4_sup, minimum depth 3; positions with less support, or where the reads disagree, are N). SNPs were identified with SKA2 0.5.1 from split 31-mers present in all genomes (core SNPs). Pairwise SNP distances count the positions where both genomes have a nucleotide and they differ. The SNPs of each genome relative to the reference were written to a VCF file by mapping the split k-mers to the reference with ska map 0.5.1. A tree was built on the SNP alignment with FastTree 2.2.0 (GTR, SH-like supports from 100 resamples) and rooted at its midpoint. Genes were read from the annotation reference.gb; the effect of each SNP on the coding sequences (codon and amino-acid change) was derived by BACoN with translation table 11. The LSC/IRb/SSC/IRa regions were derived from the annotated inverted repeats. The analysis was run with BACoN 0.3.8 (https://github.com/duceppemo/BACoN).

Run

Command
$CONDA_PREFIX/bin/bacon -r example_output/data/reference.fasta -i example_output/data/reads -o example_output/bacon -t 8 -p 2 --annotation example_output/data/reference.gb --metadata example_output/data/metadata.tsv
Reference
example_output/data/reference.fasta (1 sequence(s), MD5 a7a21202b0e29a17aff7d34ff0d3fb82)
Annotation
example_output/data/reference.gb (23 genes on 1 sequence(s); copy annotation.gb, MD5 05b2191e4afc5f5ab2ac53ce18e834bd)
Output
example_output/bacon
Python
3.12.15 on Linux
Metadata
example_output/data/metadata.tsv (copy metadata.tsv; colours by group)
minimap2
2.31-r1302 — $CONDA_PREFIX/bin/minimap2
filtlong
Filtlong v0.3.1 — $CONDA_PREFIX/bin/filtlong
samtools
samtools 1.24 — $CONDA_PREFIX/bin/samtools
ska
ska 0.5.1 — $CONDA_PREFIX/bin/ska
FastTree
FastTree 2.2.0 — $CONDA_PREFIX/bin/FastTree