
CD Genomics offers an end-to-end animal and plant whole genome de novo sequencing service combining PacBio Revio HiFi and Oxford Nanopore PromethION long-read platforms with optional Hi-C scaffolding, delivering chromosome-level and telomere-to-telomere (T2T) assemblies without dependence on a reference genome. The resulting assemblies resolve centromeres, telomeres, rDNA arrays, segmental duplications, and repetitive regions that short-read approaches cannot assemble.
For species with no reference genome, complex polyploidy, or heterozygosity profiles that confound existing tools, our dual-platform strategy provides a flexible, independently validated assembly path from raw reads to annotated, publication-ready genome sequences.
De novo genome sequencing reconstructs a genome sequence directly from raw reads, without alignment to a pre-existing reference. It is the only approach capable of capturing genome-specific features such as novel chromosomal rearrangements, species-specific gene family expansions, and structural variants absent from any reference, making it the foundation of new genome projects and a necessity for species where reference genomes are incomplete, outdated, or do not exist.
Historically, the prohibitive cost and assembly complexity of long-read technologies — combined with the requirement for high-molecular-weight DNA — limited de novo assembly to a small number of model organisms. The introduction of PacBio Revio, which roughly quadruples the throughput of Sequel IIe at equivalent or lower per-base cost, and the scaling of Oxford Nanopore PromethION to routine ultra-long read lengths (N50 > 100 kb with optimal library preparation), has transformed de novo assembly into a cost-accessible, high-resolution service for any eukaryotic species.
Modern assemblies routinely achieve contig N50 values exceeding 20 Mb for plant and animal genomes, with chromosome-level scaffolding enabled by Hi-C chromatin conformation data. Where ultra-long ONT reads are combined with HiFi accuracy, gapless telomere-to-telomere assemblies are now achievable for most diploid and some polyploid species, resolving regions that all previous technologies left unassembled.
CD Genomics operates both current-generation long-read platforms and designs each project around the platform combination that best matches genome size, complexity, ploidy, and assembly objectives.
The PacBio Revio system replaced the Sequel IIe as the primary PacBio production platform in 2023. Revio achieves approximately four times the throughput of Sequel IIe per SMRT Cell run, with similar or improved read quality. Each SMRT Cell 25M produces roughly 25–35 Gb of HiFi data; a standard four-SMRT-Cell Revio run yields approximately 90–140 Gb, sufficient for 30–60× coverage of most vertebrate or plant genomes without the additional run costs previously required on Sequel IIe.
| Parameter | PacBio Revio | PacBio Sequel IIe (Legacy) |
| HiFi yield per SMRT Cell | ~25–35 Gb | ~8–10 Gb |
| SMRT Cells per instrument run | 4 (simultaneous) | 1 |
| Single-read accuracy (HiFi) | ≥99.9% (Q30) | ≥99.9% (Q30) |
| Typical HiFi read N50 | 15–20 kb | 12–18 kb |
| Mean read length | 12–18 kb | 10–15 kb |
| Genome coverage per run | 30–60× for most genomes (<2 Gb) | Typically requires multiple runs |
| Primary assembly use case | Accurate contig assembly, SNV/indel calling, haplotype phasing | Same, lower throughput |
The PromethION P24 and P48 instruments are Oxford Nanopore's high-throughput production platforms. Paired with the R10.4.1 pore chemistry and latest Dorado basecalling models, PromethION delivers duplex reads (accuracy >Q30) alongside ultra-long reads (N50 >50 kb, with top libraries >100 kb) that are essential for spanning centromeric satellite arrays, large tandem repeats, and transposable element clusters that HiFi reads alone cannot bridge.
| Parameter | ONT PromethION (R10.4.1) | ONT MinION / GridION (reference) |
| Yield per flow cell | 50–150 Gb (P24/P48) | 5–50 Gb |
| Modal read accuracy (simplex) | ~Q20–Q25 (R10.4.1) | ~Q20–Q25 |
| Duplex read accuracy | >Q30 with Dorado duplex calling | Similar (duplex rate varies) |
| Ultra-long read N50 (optimized prep) | >50 kb; >100 kb with HMW libraries | Variable |
| Max individual read length | 2 Mb+ (ultra-long, rare) | 2 Mb+ (rare) |
| Native modification detection | 5mC, 5hmC, 6mA, others via Dorado | Same |
| Primary assembly use case | Spanning centromeres/telomeres, T2T scaffolding, methylation phasing | Smaller genomes or targeted enrichment |
For projects requiring haplotype-resolved assemblies — including plant polyploids and outbred animal genomes — we also integrate Hi-C chromatin conformation data, which provides long-range contact information for chromosome-scale scaffolding and haplotype separation. See our dedicated Haplotype-Resolved T2T Genome Assembly service page for details.
Not every genome requires the same data strategy. CD Genomics consults with clients on the biological question, genome characteristics (size, ploidy, repeat content, heterozygosity), and downstream use cases to select the most cost-efficient assembly path before sequencing begins.
| Strategy | Data Input | Expected Output | Recommended When |
| HiFi-only (Revio) | PacBio Revio HiFi, 30–60× | Contiguous, highly accurate contigs; contig N50 >10 Mb for most species; suitable for gene annotation | Diploid animals or plants with moderate repeat content; gene-model-focused projects; tight budget |
| HiFi + Hi-C | PacBio Revio HiFi + Hi-C reads | Chromosome-level, phased assembly with full scaffold anchoring; ideal for comparative genomics | Projects requiring chromosome-level scaffolding; polyploids; population-level reference genomes |
| HiFi + ONT Ultra-Long | PacBio Revio HiFi 30× + ONT PromethION 20–50× | Near-T2T contigs; spans most centromeres and large transposon arrays; QV >40 | Species with high repeat content (>50%); highly heterozygous genomes; complex polyploids |
| HiFi + ONT + Hi-C (T2T) | PacBio Revio HiFi + ONT PromethION + Hi-C | Telomere-to-telomere gapless assembly; all chromosomes fully resolved including centromeres; haplotype-phased | Reference genome construction; complete structural variant catalogue; evolutionarily important species |
For established species with reference genomes that need structural variant characterisation or resequencing, see our Animal/Plant Whole Genome Resequencing service. For pan-genome projects requiring multiple high-quality assemblies from the same species, CD Genomics offers scalable multi-accession de novo workflows.
All projects are run on PacBio Revio (not Sequel II) and ONT PromethION (not MinION), delivering the highest currently available per-read accuracy and throughput for de novo assembly.
Long reads spanning repetitive elements produce contig-level assemblies that reveal centromeric structure, segmental duplications, inversions, and gene family copy numbers undetectable by short reads or reference-guided approaches.
HiFi reads carry sufficient heterozygosity signal for phasing algorithms (Hifiasm, Verkko) to separate parental haplotypes without trio sequencing in most cases, supporting allele-specific expression and disease association studies.
ONT PromethION delivers 5mC and 6mA methylation calls from the same reads used for assembly, with no additional bisulfite conversion step. See our Long-Read DNA Methylation Sequencing service for integration options.
From DNA quality assessment and library preparation through sequencing, assembly, polishing, scaffolding, annotation, and final genome report, all steps are performed under one quality management system, eliminating hand-off errors between service providers.
CD Genomics has assembled genomes ranging from compact microbial chromosomes to complex polyploid plant genomes exceeding 10 Gb. Custom data volumes and assembly strategies are designed for each project.
Outputs include FASTA/GFF3 genome files, annotation databases, repeat landscapes, BUSCO completeness scores, and assembly statistics formatted to meet requirements of major sequence repositories (NCBI, ENA) and genome announcement journals.
The workflow applies equally to invertebrates, vertebrates, model and non-model plants, fungi, and protists. No prior genomic knowledge of the target species is required.
Our team reviews available genome size estimates, known repeat content, ploidy, and project objectives to recommend the optimal platform combination, sequencing depth, and bioinformatics strategy before sample submission.
High-molecular-weight (HMW) genomic DNA is either provided by the client or extracted by our laboratory. DNA integrity is assessed by pulse-field gel electrophoresis or Fragment Analyzer (major band >30 kb for ONT ultra-long); concentration and purity are confirmed by Qubit fluorometry and NanoDrop (OD260/280 = 1.8–2.0; OD260/230 ≥ 2.0).
Separate SMRTbell libraries (>15 kb insert, size-selected by BluePippin) are prepared for PacBio Revio HiFi sequencing; ligation-based, unsheared libraries are prepared for ONT PromethION ultra-long sequencing. Hi-C libraries are prepared from fresh or crosslinked tissue for chromosome-scale scaffolding when included in the project design.
HiFi sequencing is performed on the PacBio Revio platform, generating CCS reads with ≥Q30 accuracy. Ultra-long sequencing is performed on Oxford Nanopore PromethION (R10.4.1 chemistry), with Dorado basecalling applied for simplex and duplex read generation. Real-time monitoring confirms target yield and N50 are met before run completion.
Raw reads are assembled using graph-based assemblers optimised for the data combination in use (Hifiasm for HiFi-primary; Verkko for HiFi+ONT hybrid assemblies). Assembly graphs are inspected, tangles resolved, and residual errors corrected. Hi-C contact maps are applied for scaffolding with YaHS or 3D-DNA, followed by manual curation where required.
The final assembly is annotated for repeat elements (RepeatModeler/RepeatMasker), gene models (MAKER/BRAKER/Helixer with RNA-seq evidence), and functional annotation (InterPro, GO, KEGG). BUSCO completeness and QV scores are reported. Optional add-ons include comparative genomics, synteny analysis, and population-level variant analysis.
Workflow of animal and plant whole genome de novo sequencing, from pre-project consultation and HMW DNA quality assessment through PacBio Revio HiFi and ONT PromethION sequencing, Hifiasm/Verkko assembly, Hi-C chromosome scaffolding, and comprehensive genome annotation.
CD Genomics provides a structured bioinformatics pipeline covering assembly, quality evaluation, annotation, and comparative analysis. Analysis scope is agreed before project start and documented in the final delivery report.
| Analysis Module | Content and Tools |
| Assembly and Quality Control | |
| Read QC and filtering | HiFi CCS generation (SMRT Link), ONT basecalling (Dorado), read length/N50/quality distribution reports |
| De novo contig assembly | Hifiasm (HiFi-primary), Verkko (HiFi+ONT hybrid); graph inspection and tangle resolution |
| Assembly polishing | DeepVariant/Racon polishing for error correction; phased polishing using HiFi reads |
| Chromosome scaffolding (Hi-C) | YaHS or 3D-DNA scaffolding; Juicebox manual curation; final chromosome-level AGP generation |
| Assembly quality metrics | Contig/scaffold N50, L50, total length, gap count, BUSCO completeness (lineage-specific odb10 database), QV (Merqury), k-mer completeness |
| Haplotype-resolved assembly | Hifiasm trio or Hi-C phasing; haplotype-specific contig separation and chromosome assignment |
| Genome Annotation | |
| Repeat identification and masking | RepeatModeler (de novo repeat library), RepeatMasker (annotation); transposable element landscape plots |
| Gene model prediction | BRAKER3 / MAKER pipeline integrating ab initio prediction, protein homology, and RNA-seq evidence (if provided) |
| Functional annotation | InterProScan, GO term assignment, KEGG pathway mapping, BLAST-based homology annotation |
| Non-coding RNA annotation | tRNA (tRNAscan-SE), rRNA (RNAmmer/Barrnap), miRNA, and other ncRNA families (Infernal/Rfam) |
| Comparative and Evolutionary Genomics (Optional) | |
| Synteny and collinearity analysis | MCScan / MUMmer whole-genome alignment; synteny block visualisation between target and reference species |
| Phylogenomics | Single-copy orthologue identification (OrthoFinder); maximum-likelihood or Bayesian species tree construction |
| Gene family expansion/contraction | CAFE analysis of orthogroup size evolution across a species tree |
| Whole genome duplication (WGD) analysis | Ks distribution analysis; syntenic block depth for polyploidy inference (MCScan) |
| Structural variant annotation | Sniffles2 / PBSV for SV calling against an existing reference where applicable; inter-species comparison |
| Native DNA methylation (ONT) | Dorado 5mC/6mA modification calling; per-site methylation phased to haplotype where heterozygous |
Different sequencing technologies contribute distinct strengths to de novo assembly. The table below compares the major long-read platforms and their roles in a de novo genome project.
| Feature | PacBio Revio (HiFi) | ONT PromethION (Ultra-Long) | Illumina Short-Read (NGS) |
| Typical read length N50 | 15–20 kb | >50 kb (optimised); >100 kb (ultra-long prep) | 150–300 bp |
| Single-read accuracy | ≥Q30 (>99.9%) | ~Q20–Q25 simplex; >Q30 duplex | ≥Q30 standard |
| Repeat resolution (>10 kb) | Moderate — spans most interspersed repeats | High — spans centromeres, large satellite arrays | Fails — most repeats unresolvable |
| Contig N50 (solo assembly) | >10 Mb typical for diploid genomes | Variable — 1–20 Mb depending on accuracy | <100 kb in most species |
| Haplotype phasing | Excellent — heterozygous sites in long reads | Good — long reads, lower per-site accuracy | Poor — short reads rarely span phasing variants |
| Centromere/telomere spanning | Partial — limited by repeat length vs HiFi N50 | Yes — ultra-long reads bridge most repeat arrays | No |
| Native modification detection | 5mC (kinetics-based; lower sensitivity) | 5mC, 5hmC, 6mA, others — direct, high sensitivity | Bisulfite conversion required (separate experiment) |
| Best use in de novo assembly | Primary contig assembly; accurate polishing; SNV/SV calling | Gap closing; centromere/telomere assembly; scaffolding support | Hi-C scaffolding only; polishing (optional); not recommended for contig assembly |
For most eukaryotic genome projects, CD Genomics recommends a PacBio Revio HiFi + ONT PromethION hybrid strategy, which combines the high accuracy of HiFi reads with the spanning power of ultra-long ONT reads to produce near-T2T assemblies with QV >40. See our PacBio SMRT Sequencing Technology and Oxford Nanopore Sequencing Technology pages for detailed platform descriptions.
High-molecular-weight, high-purity DNA is the single most critical determinant of assembly quality in long-read de novo sequencing. Samples that fail to meet minimum integrity thresholds are typically flagged at QC before library preparation at no additional charge.
| Parameter | PacBio Revio HiFi | ONT PromethION Ultra-Long | Hi-C (Chromatin Conformation) | Notes |
| Sample type | HMW genomic DNA | HMW genomic DNA (unsheared) | Fresh / snap-frozen tissue or cell pellet | DNA for PacBio/ONT; intact cells/tissue for Hi-C |
| Minimum input quantity | ≥ 5 µg | ≥ 5 µg (standard); ≥ 10 µg (ultra-long >100 kb) | ≥ 1 g tissue or ≥ 1 × 10⁷ cells | Higher input improves library size selection and yield |
| DNA fragment length | Major band ≥ 20 kb (≥ 30 kb preferred) | Major band ≥ 50 kb (≥ 100 kb for ultra-long) | Not applicable (chromatin-based) | Assessed by PFGE or Fragment Analyzer |
| Purity (OD260/280) | 1.8–2.0 | 1.8–2.0 | — | Protein or RNA contamination causes library failure |
| Purity (OD260/230) | ≥ 2.0 | ≥ 2.0 | — | Low A260/230 indicates polysaccharide or phenol carryover |
| Concentration | ≥ 50 ng/µL (Qubit dsDNA BR) | ≥ 50 ng/µL (Qubit dsDNA BR) | — | Nanodrop overestimates RNA- or protein-contaminated samples; Qubit required |
| Storage | TE buffer (pH 8.0) at 4°C short-term; −20°C long-term | TE buffer (pH 8.0) at 4°C; avoid freeze–thaw | Snap-frozen; dry ice shipping | Avoid vortexing HMW DNA; pipette with wide-bore tips |
| Shipping | Dry ice for DNA; or cooled if short transit | Dry ice; upright orientation | Dry ice | Label tubes with species, sample ID, and extraction date |
For plant samples with high polysaccharide or secondary metabolite content (e.g., conifers, citrus, legumes), CD Genomics offers in-house HMW DNA extraction optimised for difficult plant matrices. Contact our team before sample preparation to confirm the recommended extraction protocol for your species. For long-amplicon target sequencing rather than whole genome, see our Animal/Plant Long Amplicon Sequencing service.
Current-Generation Platforms — No Legacy Instruments
All animal and plant de novo genome projects are run on PacBio Revio and ONT PromethION. We retired Sequel IIe-only workflows when Revio became available, ensuring clients receive the highest current yield and accuracy per run without carrying legacy platform costs.
Dual-Platform Hybrid Expertise
CD Genomics has extensive experience designing and executing hybrid assemblies that leverage the accuracy of HiFi reads and the spanning length of ONT ultra-long reads. Assembly strategy is adapted to each species rather than applied generically.
T2T and Haplotype-Resolved Assembly Capability
Beyond chromosome-level assemblies, CD Genomics delivers fully phased, gapless T2T genomes using Hifiasm+Hi-C and Verkko+ONT pipelines, with manual curation of centromeric regions. See our Haplotype-Resolved T2T Assembly service for dedicated options.
Integrated Epigenome Profiling
For projects that require both assembly and methylation context, ONT PromethION data is processed through Dorado modification calling, delivering base-resolution 5mC and 6mA maps phased to the assembled haplotypes in a single experiment.
End-to-End Scientific Delivery
Project scientists are available from experimental design through data interpretation. Final deliverables include raw data, assembly files, annotation GFF3/functional tables, QC summaries, and a structured analysis report formatted for journal submission.
Liu F, Wang Y, Geng T, et al. Haplotype-resolved T2T genome assembly of the Populus nigra NL-1976. Scientific Data 2025. DOI: 10.1038/s41597-025-06361-2.
Populus nigra (black poplar) is an ecologically and economically important forest tree whose genome had not been assembled to full haplotype-resolved, telomere-to-telomere quality. The species' high heterozygosity and complex repeat landscape made it an ideal test case for evaluating the combined capacity of PacBio Revio HiFi, ONT PromethION ultra-long, and Hi-C sequencing for a challenging woody plant genome.
High-molecular-weight genomic DNA was extracted from Populus nigra NL-1976 and subjected to three sequencing strategies run in parallel:
Assembly was performed with Hifiasm in HiFi+ONT+Hi-C mode, producing two phased haplotype contig sets that were scaffolded separately to chromosome level using Hi-C contact data and verified for telomere and centromere completeness.
Both haplotype assemblies were delivered at near-T2T quality:
Figure 4 from Liu et al. (2025). Circos plot of the two haplotype-resolved Populus nigra assemblies (nigraHap1 and nigraHap2) showing chromosome-level features including gene density, repeat distribution, and inter-haplotype synteny. Reprinted under CC BY 4.0.
This study demonstrates that PacBio Revio HiFi combined with ONT PromethION ultra-long reads and Hi-C scaffolding can deliver fully haplotype-resolved, chromosome-level T2T assemblies for a complex, heterozygous woody plant genome at QV >40. The approach provides a reproducible template directly applicable to other forest trees, crop species, and animal genomes of comparable complexity.
No. CD Genomics transitioned de novo genome projects to PacBio Revio as the primary PacBio platform. Revio delivers approximately four times the HiFi data yield per SMRT Cell compared to Sequel IIe, which reduces run counts and project costs for most genome sizes. Sequel IIe is no longer recommended for new de novo projects given the throughput and cost advantages of Revio.
HiFi-only assembly with Revio is sufficient for producing high-quality, gene-complete contig assemblies for most diploid animal and plant genomes with moderate repeat content (<40%). A HiFi + ONT hybrid strategy — combining Revio HiFi accuracy with PromethION ultra-long read length — is recommended when the genome has high repeat content (>50%), large centromeric arrays, or when a T2T-quality gapless assembly is required. ONT ultra-long reads are particularly effective at bridging the satellite repeat regions that even 20 kb HiFi reads cannot span.
A chromosome-level assembly anchors contigs onto chromosomes using Hi-C or genetic map data, but may still contain internal gaps — particularly at centromeres, rDNA arrays, and highly repetitive sub-telomeric regions. A T2T (telomere-to-telomere) assembly resolves these gaps so that each chromosome is represented by a single, continuous, gapless sequence from one telomere to the other. T2T requires both the spanning power of ultra-long ONT reads and the accuracy of HiFi reads, plus Hi-C for chromosome anchoring and manual curation of complex regions.
Yes. PacBio Revio HiFi reads carry sufficient heterozygosity information for phasing algorithms to separate sub-genomes in allopolyploids without trio sequencing in many cases. For highly complex polyploids (hexaploid wheat, octoploid strawberry), we design customised strategies combining high-depth HiFi, ONT ultra-long, Hi-C, and in some cases parental Illumina data for trio-based assembly. Please contact our team with your species and ploidy for a tailored project design.
DNA input and integrity are the most critical pre-sequencing factors. For PacBio Revio, we require ≥5 µg of HMW genomic DNA with a major band ≥20 kb by PFGE or Fragment Analyzer. For ONT ultra-long libraries targeting N50 >100 kb, we recommend ≥10 µg with a major band ≥50 kb and avoiding any freeze–thaw cycles or vortexing. Sheared or degraded DNA directly limits maximum read length, which limits the assembly's ability to span repeats and produces lower contig N50 values regardless of sequencing depth.
Standard deliverables include: raw HiFi and/or ONT FASTQ files, assembled FASTA genome sequences (contig and scaffold level), annotation files in GFF3 format, functional annotation tables, repeat annotation outputs, BUSCO and QV reports, and a structured analysis report. For T2T or haplotype-resolved projects, separate files for each haplotype and telomere/centromere verification reports are included. Custom add-ons such as synteny analyses, gene family evolution, and methylation reports can be added to the project scope.
Both options are supported. Client-extracted DNA is accepted provided it passes our incoming QC thresholds (OD260/280 1.8–2.0; OD260/230 ≥2.0; HMW band ≥20 kb). For samples with challenging matrices — high polysaccharide plant tissues, CTAB-extracted samples, or animal tissues with high glycogen content — CD Genomics also offers in-house extraction optimised for HMW DNA recovery. Contact our team before extraction to confirm the optimal protocol for your species and tissue type.
Yes. De novo genome assemblies produced by CD Genomics are regularly published in journals including Scientific Data, GigaScience, DNA Research, G3, and species-specific journals. Deliverables are structured to meet INSDC submission requirements for NCBI/ENA deposition. We can provide assembly and annotation statistics, metadata tables, and methods text templates to support manuscript preparation on request.
1. Assembly Continuity Plot — Contig N50 comparison across assembly strategies (HiFi-only vs HiFi+ONT hybrid) for a representative plant genome
2. BUSCO Completeness Bar Chart — Genome completeness assessment against embryophyta_odb10 database showing C, D, F, M proportions for both haplotypes
3. Contig Length vs QV Scatter Plot — Per-contig quality value (Merqury) plotted against contig length for the final polished assembly

References
For Research Use Only. Not for use in diagnostic procedures.