Abstract
16S rRNA amplicon sequencing is a known cost-effective approach for taxonomic profiling in microbiome studies. While Illumina short-read platforms remain dominant, Oxford Nanopore Technologies (ONT) long-read sequencing offers the potential for higher taxonomic resolution. However, the accuracy of profiling depends critically on the choice of bioinformatics tools, sequencing platforms, and reference databases. This study benchmarks GAIA v3.1.0, an updated alignment-based taxonomic classifier, against established tools for both short-read and long-read 16S rRNA amplicon data. We further evaluate the impact of using GAIA’s 16S NCBI-based database versus a curated SILVA database on classification performance. We utilized two independent benchmarking frameworks: Odom et al. (2023) for short-read data and Zhang et al. (2023) for long-read data. The evaluation included 125 Illumina samples processed by tools such as DADA2, Mothur, and QIIME 2, and 18 ONT samples processed by EMU and LAST+LCA. Performance was assessed using F1 score, precision, recall, and L1 norm error at the genus level for short reads and the species level for long reads. GAIA v3.1.0 demonstrated superior performance across both platforms. For short-read data at the genus level, GAIA with the NCBI database achieved the highest average F1 score (0.787), significantly outperforming Mothur (0.672), the second best method. For ONT long-read data at the species level, GAIA achieved the top average F1 score (0.532), surpassing EMU (0.421) and LAST+LCA (0.353). Regarding database selection, the NCBI-based database generally provided better precision for species-level resolution in long reads, whereas the SILVA database tended to maximize recall. With these findings, GAIA v3.1.0 provides a robust and versatile solution for 16S rRNA taxonomic classification.