Complete fungal genome assembly from PacBio HiFi (Sequel II / Revio), Oxford Nanopore PromethION R10.4.1, and hybrid strategies — producing gapless, telomere-to-telomere assemblies for pathogenic fungi, industrial strains, model organisms, and environmental isolates across all fungal phyla
CD Genomics provides fungal whole-genome de novo sequencing on PacBio HiFi, Oxford Nanopore PromethION R10.4.1, and Illumina platforms — producing telomere-to-telomere, gapless assemblies for pathogenic, industrial, and environmental fungi across all major phyla. Our service is tailored to fungal genome complexities including AT-rich repeat resolution, diploid-aware phasing, accessory chromosome capture, and comprehensive annotation with CAZy, antiSMASH, and effector prediction databases.
Fungal whole-genome de novo sequencing is the process of reconstructing a complete genome sequence for a fungal species or strain from scratch, without reliance on a pre-existing reference genome. Unlike resequencing — which identifies differences relative to a known reference — de novo assembly builds the genome directly from sequencing reads, producing the complete set of chromosomes, mitochondrial genome, and any extrachromosomal elements. This approach is essential for characterizing novel fungal isolates, species without a high-quality reference, or strains where the existing reference is incomplete or misrepresentative.
Fungal genomes present unique challenges that make long-read sequencing especially valuable. At 10–100+ Mb with highly repetitive AT-rich regions, complex secondary metabolite gene clusters, variable numbers of accessory chromosomes, and often-difficult telomere-to-telomere resolution, fungal genomes require the read length and accuracy that only PacBio HiFi and Oxford Nanopore technologies can provide. Short-read assemblies of fungal genomes are typically fragmented into hundreds of contigs with collapsed repeats, incomplete telomeres, and misassembled biosynthetic gene clusters — limitations that long-read sequencing overcomes by spanning repetitive elements and producing complete chromosomal assemblies in a single contig per chromosome.
At CD Genomics, we provide fungal WGS de novo sequencing across all major platforms — PacBio HiFi, Oxford Nanopore PromethION, and Illumina NovaSeq — with tailored strategies for genome size, ploidy level, repeat content, and project budget. Our bioinformatics pipelines are optimized for fungal genome complexities including diploid-aware phasing, accessory chromosome resolution, and comprehensive functional annotation with fungal-specific databases.
Fungal whole-genome de novo sequencing assembles a complete genome sequence for a fungal species or strain entirely from sequencing reads, without relying on a reference genome for guidance. The output is a set of chromosomal pseudomolecules — ideally one contiguous sequence per chromosome, with telomeres at both ends, the complete mitochondrial genome, and any plasmid or double-stranded RNA viral elements — representing the first complete genomic description of the target organism.
Fungal genomes span a remarkable range of size and complexity. Unicellular yeasts such as Saccharomyces cerevisiae have compact ~12 Mb genomes with few repeats, while filamentous fungi like Aspergillus species range from 30–40 Mb with moderate repeat content, and plant pathogens like Puccinia or Blumeria exceed 100 Mb with extremely high repeat densities (>80% transposable elements). Many fungal species also carry accessory (conditionally dispensable) chromosomes that are present in some strains but absent in others — often harboring virulence factors, host-range determinants, or secondary metabolite clusters that are invisible in assembly drafts that collapse or discard them. Short-read sequencing systematically under-assembles these complex regions: repeats collapse, accessory chromosomes fragment, rDNA arrays remain as gaps, and telomeres are lost entirely.
Long-read sequencing transforms fungal genome assembly. A single PacBio HiFi read (15–25 kb, Q30+) or ONT ultra-long read (50–100+ kb) can span an entire transposable element, resolve a multi-gene secondary metabolite cluster in its chromosomal context, bridge an rDNA repeat unit, and connect the subtelomeric region to the chromosome end. With sufficient coverage and appropriate assembly algorithms (hifiasm, Flye, CANU, HiCanu), long-read assemblies routinely produce telomere-to-telomere fungal genomes with no gaps, no uncollapsed repeats, and complete chromosomal structure.
At CD Genomics, our fungal WGS de novo service is designed around the specific biological and technical requirements of fungal genomics — from AT-rich genome assembly optimization and diploid phasing to fungal-specific annotation pipelines — delivering complete genome blueprints for research, biotechnology, and agricultural applications.
Fungal genomes typically harbor 10–80% transposable element content, along with long rDNA arrays, subtelomeric repeats, and centromeric regions. Long reads spanning the full length of these repetitive elements allow the assembler to resolve them correctly rather than collapsing them into incorrect consensus sequences — essential for accurate genome size estimation and repeat landscape characterization.
Conditionally dispensable (accessory) chromosomes are common in pathogenic and symbiotic fungi and carry key adaptive genes. Long-read assemblies recover these chromosomes as complete, independent contigs with telomeric ends, whereas short-read assemblies typically fragment them or miss them entirely due to their unusual repeat content and sequence composition.
Many fungi are haploid, but important groups are diploid (Candida albicans, Zygosaccharomyces), dikaryotic (rust fungi, mushrooms), or exhibit variable ploidy. PacBio HiFi reads with hifiasm produce high-quality phased assemblies, separating haplotypes and resolving allelic differences — including heterozygous SNPs, InDels, and gene presence-absence variants — without the collapse or switching errors typical of short-read assemblies.
With modern long-read sequencing and assembly algorithms, telomere-to-telomere fungal genome assembly is achievable at coverage levels that are cost-effective even for large genomes. ONT-only strategies (with HERRO correction) can produce gapless assemblies from simplex reads at 40–65× coverage, while PacBio HiFi provides the highest per-base accuracy for annotation-grade assemblies.
We tailor sequencing strategy to genome size, ploidy, and budget — from cost-effective ONT-only assembly for large genomes to hybrid PacBio HiFi + Illumina strategies for the highest quality assemblies, or PacBio HiFi-only for annotation-grade T2T genomes.
Fungal genome annotation requires specialized databases not covered by standard prokaryotic pipelines: CAZy for carbohydrate-active enzymes (critical for plant cell wall-degrading fungi), antiSMASH for secondary metabolite clusters, PHI (Pathogen-Host Interactions) for effector and virulence proteins, and MITE/RepeatModeler for transposable element characterization — all included in our standard annotation package.
Beyond pathogen and industrial genomics, long-read fungal genome assembly is increasingly applied in natural product drug discovery, where complete resolution of biosynthetic gene clusters (BGCs) in their native chromosomal context enables the identification of novel polyketide synthase (PKS) and non-ribosomal peptide synthetase (NRPS) clusters that are fragmented or collapsed in short-read assemblies — opening access to previously hidden fungal secondary metabolite diversity for pharmaceutical development.
Fungal cell walls require specialized lysis protocols — enzymatic digestion (lyticase, chitinase, glucanase), gentle bead-beating, or CTAB-based methods optimized for each taxonomic group. High-molecular-weight DNA (fragments ≥ 30 kb) is essential for long-read assembly, particularly for spanning repetitive regions and rDNA arrays. DNA purity (removal of polysaccharides, polyphenols, and other fungal-specific contaminants) is verified by UV spectrophotometry, Qubit fluorometry, and pulsed-field gel electrophoresis.
Platform-specific libraries are prepared with barcoded adapters for multiplexed sequencing. For PacBio HiFi, 15–20 kb SMRTbell libraries with HiFi CCS mode and barcoded overhang adapters (up to 384-plex). For Oxford Nanopore, native ligation libraries (SQK-LSK114) with native barcoding (up to 96-plex), with optional ultra-long library protocols for maximum contiguity. For Illumina, 350–550 bp paired-end libraries for hybrid polishing. Optional Hi-C libraries for chromosome-scale scaffolding when required.
Sequencing coverage is scaled to genome size and complexity. For PacBio HiFi: 30–50× CCS for yeasts and small genomes (~10–20 Mb), 50–100× for filamentous fungi (~30–80+ Mb). For Oxford Nanopore: 50–100× raw coverage with R10.4.1 flow cells and Dorado SUP basecalling. For hybrid strategies: 30–60× HiFi or ONT + 50× Illumina polishing.
Figure 1. Complete fungal whole-genome de novo sequencing workflow — from HMW DNA extraction and multi-platform sequencing through assembly (Flye, hifiasm, CANU), polishing, telomere-to-telomere completion, and comprehensive fungal-specific annotation.
Quality-filtered reads are assembled using fungal-optimized pipelines. Primary assemblers: Flye (for ONT and HiFi), hifiasm (for HiFi, with diploid phasing), CANU or HiCanu (for ONT ultra-long). See our genome assembly service for detailed methodology and quality metrics. Assembly polishing: two rounds of Racon or Medaka for ONT, followed by Pilon or Polypolish with Illumina data. Telomere-to-telomere completion: manual curation with IGV, telomere motif identification, and gap closure verification. Assembly QC includes QUAST, BUSCO (fungi_odb10), CheckM, and Merqury k-mer completeness analysis.
Structural annotation: ab initio gene prediction (BRAKER3, Augustus, GeneMark-ES) supported by fungal RNA-seq evidence, with manual curation of problematic loci. Functional annotation: against fungal-specific databases including CAZy, antiSMASH, PHI-base, DFVF, MEROPS, Pfam, and KEGG. Repeat annotation: RepeatModeler + RepeatMasker for TE characterization. Deliverable package includes annotated GenBank files, complete genome maps (CGView/Circos), and a comprehensive project report.
Our fungal genome annotation pipeline is specifically designed for fungal biology, incorporating kingdom-specific tools and databases that go beyond generic genome annotation workflows.
| Analysis Feature | Basic Package | Advanced Package |
| Read QC, assembly & completeness assessment | ✓ FastQC, NanoPlot; Flye/hifiasm/CANU assembly; QUAST statistics; BUSCO (fungi_odb10); CheckM | ✓ + Merqury k-mer completeness; multiple assembler comparison; Hi-C scaffolding; T2T manual curation |
| Repeat annotation & TE characterization | ✓ RepeatModeler + RepeatMasker; TE class/family classification | ✓ + TE insertion time estimation; RIP index calculation; landscape analysis; TE proximity to genes |
| Structural gene prediction | ✓ BRAKER3 / Augustus / GeneMark-ES with RNA-seq evidence integration | ✓ + Manual curation of problematic loci; alternative transcript detection; ncRNA identification |
| Functional annotation (general) | ✓ InterProScan, Pfam, GO, KEGG, COG, Nr alignment | ✓ + KEGG pathway mapping; Enzyme Commission numbers; Transporter classification (TCDB) |
| Carbohydrate-active enzymes (CAZy) | ✓ dbCAN3 / HMMER search against CAZy database; GH, GT, CE, AA, PL family classification | ✓ + Modular CAZyme architecture analysis; signal peptide prediction; secreted CAZyme identification |
| Secondary metabolite gene clusters | ✓ antiSMASH fungal version; core biosynthetic gene identification (PKS, NRPS, terpene, indole) | ✓ + Cluster boundary refinement; comparative cluster analysis; heterologous expression target prioritization |
| Secreted proteins & effector prediction | ✓ SignalP, TargetP, TMHMM for secreted protein identification | ✓ + EffectorP (fungal effector prediction); small cysteine-rich protein identification; RxLR/LFY motif analysis |
| Pathogen-host interaction (PHI) annotation | ✓ BLAST against PHI-base; virulence factor classification | ✓ + DFVF (fungal virulence factor) database; phenotype association; comparative pathogenomics |
| Mating-type locus analysis | — | ✓ MAT locus identification; idiomorph structure resolution; mating-type determination |
| Comparative & population genomics | — | ✓ Orthofinder gene family analysis; phylogenomics; synteny visualization; pan-genome construction |
| Custom reporting & visualization | ✓ Standard project report PDF with summary statistics and annotation tables | ✓ Interactive genome browser (JBrowse2); Circos genome maps; publication-ready figures; NCBI submission files |
The optimal sequencing strategy for fungal whole-genome de novo assembly depends on genome size, ploidy, repeat content, and the intended applications of the assembly. We provide platform-neutral recommendations based on your specific project requirements.
| Feature | PacBio HiFi | Oxford Nanopore | Hybrid (HiFi + ONT + Illumina) |
| Read length | 15–25 kb (CCS) | 20–100+ kb (native) | Multi-platform combination |
| Per-base accuracy | ★★★★★ (Q30+; >99.9%) | ★★★☆☆ (Q14–Q20 raw; Q30+ polished) | ★★★★★ (polished) |
| Assembly contiguity | ★★★★☆ (T2T for most fungi at 50×) | ★★★★★ (ultra-long reads span largest repeats) | ★★★★★ (best of both worlds) |
| Repeat resolution | ★★★★☆ (15–25 kb spans most TEs) | ★★★★★ (50–100 kb spans rDNA, large repeats) | ★★★★★ (comprehensive) |
| Diploid phasing | ★★★★★ (hifiasm phased assembly) | ★★★☆☆ (limited; needs HiFi for phasing) | ★★★★★ (HiFi phasing + ONT contiguity) |
| Coverage needed (yeast ~12 Mb) | 30–50× CCS | 40–80× raw | 30× HiFi + 40× ONT |
| Coverage needed (filamentous ~40 Mb) | 50–100× CCS | 60–100× raw | 50× HiFi + 50× ONT |
| Per-genome cost (multiplexed) | $$$ (higher per-Gb) | $ (lowest per-Gb) | $$–$$$ (varies) |
| Best suited for | Annotation-grade T2T genomes; diploid phasing; studies requiring maximum per-base accuracy for SNP-level analyses | Cost-effective T2T assembly; large repeat-rich genomes; rapid screening; projects with many strains | Ultimate quality assemblies; large complex genomes; multi-omics projects combining assembly with epigenetics |
For most fungal genome projects, we recommend a primary long-read strategy (PacBio HiFi or ONT) supplemented by a modest amount of Illumina data for polishing validation. If the highest possible assembly quality is required — for a reference-quality genome intended for community use — a combined PacBio HiFi + ONT ultra-long approach provides both the per-base accuracy of HiFi and the repeat-spanning capability of ultra-long reads. Contact our team for a free project consultation and platform recommendation.
| Category | Requirement | Notes |
| Sample type | Genomic DNA, mycelium (fresh or frozen), yeast pellet, spore suspension, or glycerol stock | DNA extraction service available for challenging fungal samples (cell wall-rich, polysaccharide-heavy, or low-biomass) |
| Minimum input (gDNA) | 500 ng (Illumina); 5 µg (PacBio HiFi); 5 µg (Nanopore) | Lower input may be accepted for specific protocols; HMW DNA (≥ 30 kb) strongly recommended for long-read assembly |
| DNA quality | OD260/280: 1.8–2.0; OD260/230: ≥ 1.8; no visible degradation; free of polysaccharide/polyphenol contamination | Fungal-specific DNA purification methods used to remove polysaccharides and phenolic compounds that inhibit library preparation |
| Genome information | Estimated genome size, ploidy, and G+C content (if known) | Prior information helps optimize coverage and assembly strategy; genome size estimation service available for novel isolates |
| Sample numbers | Single isolate to large population studies | Multiplexed barcoding enables cost-effective assembly of multiple strains; dedicated project management for large projects |
| Shipping conditions | gDNA: ice pack (4°C) or dry ice; Mycelium/pellet: dry ice | See our Sample Submission Guidelines for detailed instructions |
Dedicated Fungal Genomics Expertise
Fungal genomics is not an afterthought at CD Genomics — it is a core service area with dedicated protocols, pipelines, and expertise that are qualitatively different from our bacterial WGS services. From fungal-specific DNA extraction protocols that remove polysaccharide and polyphenol contaminants, to genome assembly pipelines calibrated for AT-rich genomes and diploid phasing, to annotation pipelines with CAZy, antiSMASH, PHI-base, and effector prediction — every aspect of our service is tailored to the unique biology of fungi.
Multi-Platform Independence
We operate PacBio HiFi (Sequel II / Revio), Oxford Nanopore (PromethION R10.4.1), and Illumina NovaSeq platforms in-house, allowing us to design the optimal sequencing strategy for each fungal genome project based purely on scientific requirements — not platform availability. When a project requires the highest accuracy for annotation-grade genome assembly, we recommend PacBio HiFi. When cost-efficient assembly of multiple large genomes is needed, we design ONT-based strategies.
Telomere-to-Telomere Track Record
Our team has delivered complete, telomere-to-telomere fungal genome assemblies across diverse taxonomic groups, from compact yeast genomes to large filamentous fungal genomes with extreme repeat content. We understand the specific bioinformatics challenges of each fungal lineage and maintain optimized protocols for rDNA array resolution, accessory chromosome identification, and manual assembly curation.
Publication Support
Every fungal genome assembly project includes NCBI GenBank/Genome submission support, detailed methods sections suitable for genome announcement or resource papers, and publication-ready figures. Our team has experience with fungal genome publications and understands the data presentation standards required by journals in the field.
Hoyer LL, Freeman BA, Hogan EK, Hernandez AG. Use of a Candida albicans SC5314 PacBio HiFi reads dataset to close gaps in the reference genome assembly, reveal a subtelomeric gene family, and produce accurate phased allelic sequences. Frontiers in Cellular and Infection Microbiology. 2024;14:1329438. doi:10.3389/fcimb.2024.1329438.
Candida albicans is the most prevalent human fungal pathogen and a leading cause of invasive candidiasis in immunocompromised patients. The SC5314 reference genome (assembly ASM18296v3) was originally assembled from Sanger and short-read sequencing and contained 80 spanned sequence gaps (regions of Ns) — unresolved repetitive regions, rDNA arrays, and complex genomic loci that short-read technologies could not assemble. These gaps included critical genomic features such as subtelomeric regions, the rDNA locus on chromosome R, and genes encoding the ALS (agglutinin-like sequence) adhesin family — key virulence factors whose complete sequences were essential for understanding C. albicans pathogenesis but remained fragmented or missing from the reference.
In this study, Hoyer et al. generated PacBio HiFi sequencing data for the SC5314 strain and used it to produce an updated de novo assembly that systematically resolved all 80 gaps and completed the genome.
PacBio HiFi reads were generated on the Sequel IIe system (two SMRT Cells, 6 kb and 13 kb library sizes) and assembled using hifiasm to produce a collapsed haploid telomere-to-telomere assembly (ASM3268872v1). The new assembly was compared to the existing ASM18296v3 reference using MUMmer alignment, gap analysis, and manual inspection of previously unresolved loci. Telomeric repeats were identified by searching for the C. albicans telomere repeat sequence (GGTGTACGGATGTCTAACTTCTT) at chromosome ends. Individual HiFi read alignments were used to resolve the diploid sequence of the ALS gene family.
Figure 2. Comparison of the previous C. albicans reference assembly (ASM18296v3, 88 contigs, 80 gaps) with the updated PacBio HiFi assembly (ASM3268872v1, 8 contigs — one per chromosome, zero gaps). The new assembly added telomeric repeats to all 16 chromosome ends and discovered a conserved family of DNA helicase-encoding (YRF1) genes at 10 of 16 subtelomeric regions. Adapted from Hoyer et al. (2024), Frontiers in Cellular and Infection Microbiology, CC BY 4.0.
This study demonstrates that PacBio HiFi sequencing alone — without additional platforms — can produce a complete, telomere-to-telomere fungal genome assembly that resolves regions intractable to short-read approaches for over a decade. The 8-contig assembly (one per chromosome), complete telomere resolution, and discovery of a previously hidden gene family illustrate the transformative impact of long-read de novo sequencing on fungal genomics. The approach is directly applicable to any fungal species for which a complete genome sequence is desired.
CD Genomics provides free project consultation to help determine the optimal fungal genomics strategy for your specific research questions. Contact our scientists to discuss your project requirements.
Fungal genomes are typically 5–50× larger than bacterial genomes (10–120 Mb vs. 1–10 Mb) with much higher repeat content (10–80% transposable elements), AT-rich regions, larger rDNA arrays, and complex subtelomeric structures. Fungal DNA extraction is more challenging due to rigid cell walls and polysaccharide/polyphenol contaminants. Assembly pipelines must handle diploid or dikaryotic genomes (requiring phasing), accessory chromosomes, and variable ploidy. Annotation requires fungal-specific databases (CAZy, antiSMASH, PHI-base, DFVF) that are not used in bacterial pipelines. These differences require specialized expertise and tailored experimental and bioinformatics protocols beyond standard bacterial genome assembly workflows.
Coverage requirements depend on genome size, repeat content, and platform. For PacBio HiFi, 30–50× CCS coverage is typically sufficient for yeasts and small genomes (~10–20 Mb), while 50–100× is recommended for filamentous fungi (~30–80+ Mb). For Oxford Nanopore, 40–80× raw coverage with R10.4.1 flow cells produces T2T assemblies for most fungi when using HERRO correction and hifiasm. Higher coverage within these ranges is recommended for genomes with very high repeat content or polyploid genomes. Coverage can be adjusted during project design — we provide specific recommendations after reviewing your target species.
Yes. Many clinically and industrially important fungi are diploid or exhibit variable ploidy. With PacBio HiFi data, hifiasm produces high-quality phased assemblies that separate the two haplotypes and resolve allelic differences — including heterozygous SNPs, InDels, and structural variants — without the collapse or switching errors common in short-read assemblies. For dikaryotic fungi (e.g., rust pathogens, mushrooms), we implement specialized pipelines that separate the two nuclear genomes. The phased assembly output includes both primary and alternate pseudohaplotypes suitable for downstream analysis.
Beyond standard gene prediction and functional annotation (InterProScan, Pfam, GO, KEGG), our fungal annotation includes: CAZy database for carbohydrate-active enzyme classification (critical for plant cell wall-degrading and biomass-converting fungi), antiSMASH for secondary metabolite biosynthetic gene cluster identification (PKS, NRPS, terpene, indole clusters), PHI-base and DFVF for pathogen-host interaction and virulence factor annotation, EffectorP for fungal effector prediction, MEROPS for protease classification, and RepeatModeler/RepeatMasker for comprehensive transposable element annotation. Custom database annotation is available upon request for specialized projects.
Typical turnaround times depend on genome size, sequencing platform, and annotation requirements. For a standard filamentous fungal genome (~40 Mb): 30–45 working days for PacBio HiFi sequencing and assembly, or 25–35 working days for Oxford Nanopore sequencing and assembly. Annotation adds approximately 10–15 working days depending on the complexity of the analysis package selected. Expedited timelines may be available for urgent projects. A detailed project timeline is provided during the project design phase based on your specific requirements.
Deliverable Examples for Fungal WGS De Novo Sequencing Projects
1. Complete genome assembly file (FASTA) with chromosomal pseudomolecules — each chromosome as a single contiguous sequence with telomeric repeats at both ends, plus mitochondrial genome and any accessory chromosomes or extrachromosomal elements.
2. Genome annotation file (GFF3/GenBank format) with structural and functional annotations — protein-coding genes, tRNA, rRNA, ncRNA, repeat annotations, and predicted functional domains with cross-references to CAZy, antiSMASH, PHI-base, and KEGG.
3. Circular genome visualization (CGView/Circos) showing multi-track chromosomal maps with CDS features, repeat density, secondary metabolite cluster locations, GC content, and GC skew — publication-ready format.
4. Comprehensive project report PDF documenting experimental methods, sequencing QC metrics, assembly statistics, annotation summary, and detailed methods text suitable for manuscript preparation.
Figure 3. Representative deliverable formats for fungal WGS de novo sequencing projects. Left: multi-track chromosome map with CDS, repeats, and SM clusters. Center: CAZy and antiSMASH annotation summary. Right: assembly QC metrics and comparative genomics. AI-generated representative data.
References
For research use only. Not for use in diagnostic procedures.