Brett Pike, Anders Gonçalves da Silva, Wilson Terán
We present 4 haplotype-resolved, chromosome scale diploid assemblies of cannabis, assembled from ONT R9.4.1 reads via a novel pipeline based on Hi-C phasing. These assemblies, while low in QV, offer contiguity and genic content comparable to recent HiFi assemblies. Along with a trio-binned assembly previously produced by us and 56 haplotypes selected from the Salk Institute Cannabis Pangenome project, we use these assemblies to create a reference-free pangenome graph. Within a total length of 6.48 Gb, it contains 162.14 M nodes, 228.27 M edges, 14.87 M SNPs, and 6.40 M indels. By optimizing parameters within the Pangenome Graph Builder (PGGB), we avoid many spurious connections among repeat elements, reduce processing time, and arrive at a data structure that visibly recapitulates the linear nature of plant chromosomes. Via k-mer analysis, we corroborate that more genotypes are needed to close the cannabis pangenome, and that, in particular, the region of origin likely remains undersampled.