TY - JOUR
T1 - Nanopore-m6A-finder, a novel m6A site caller for Nanopore DRS data
AU - Yang, Yuening
AU - Yu, Liqun
AU - Mo, Li
AU - Qi, Changhai
AU - Song, Wei
AU - Jin, Hua
N1 - Publisher Copyright:
Copyright © 2026 Yang, Yu, Mo, Qi, Song and Jin.
PY - 2026/3
Y1 - 2026/3
N2 - Introduction – N 6-methyladenosine (m6A) is a pivotal RNA modification involved in diverse biological and pathological processes. Compared to the m6A detection methods based on second-generation sequencing, Nanopore direct RNA sequencing (DRS) offers the unique advantage of capturing native modifications. Methods – Here, we present Nanopore-m6A-Finder (NP-mFinder), a reference-free m6A prediction computational framework that employs the XGBoost model in the mRNA exonic region and a hard-voting ensemble of XGBoost and random forest models in the poly(A) region. Results and discussion – NP-mFinder can determine m6A sites as well as estimate their methylation levels from Guppy basecalled DRS data. After training with DRS data of in vitro-transcribed RNA, NP-mFinder achieved high performance on held-out test datasets (area under the curve (AUC) ≈0.90; accuracy, precision, recall, and F1-score >0.80). Comparing with canonical m6A detection methods, it recovered 20% of meRIP-seq-defined m6A sites in yeast, and 27% of our HEK293 site prediction overlapped with miCLIP calls. Although single-base overlap with existing DRS-based tools of EpiNano and mAFiA was limited, 73% of our identified m6A-containing genes were validated by at least one of them. Benchmarking our method with GLORI v2.0 revealed concordance of 28% at a site level and 85% at a gene level, as well as a mild correlation on m6A level estimations. Notably, NP-mFinder achieved 93% precision in detecting m6A within the “AAAAA” sequence context in the mRNA exonic region of HEK293T DRS data when compared to high-confidence m6A site annotation in GLORI v2.0, demonstrating the good performance of our method in the region possessing a stretch of continuous A-sequences. Moreover, our method predicted that m6A might exist in the human HEK293 poly(A) region, suggesting a possibly conserved phenomenon of a modified poly(A) tail beyond the previously reported T. brucei variant surface glycoprotein (VSG) transcripts. Together, these results established NP-mFinder as a robust and versatile tool for transcriptome-wide m6A profiling with DRS data at single-read resolution.
AB - Introduction – N 6-methyladenosine (m6A) is a pivotal RNA modification involved in diverse biological and pathological processes. Compared to the m6A detection methods based on second-generation sequencing, Nanopore direct RNA sequencing (DRS) offers the unique advantage of capturing native modifications. Methods – Here, we present Nanopore-m6A-Finder (NP-mFinder), a reference-free m6A prediction computational framework that employs the XGBoost model in the mRNA exonic region and a hard-voting ensemble of XGBoost and random forest models in the poly(A) region. Results and discussion – NP-mFinder can determine m6A sites as well as estimate their methylation levels from Guppy basecalled DRS data. After training with DRS data of in vitro-transcribed RNA, NP-mFinder achieved high performance on held-out test datasets (area under the curve (AUC) ≈0.90; accuracy, precision, recall, and F1-score >0.80). Comparing with canonical m6A detection methods, it recovered 20% of meRIP-seq-defined m6A sites in yeast, and 27% of our HEK293 site prediction overlapped with miCLIP calls. Although single-base overlap with existing DRS-based tools of EpiNano and mAFiA was limited, 73% of our identified m6A-containing genes were validated by at least one of them. Benchmarking our method with GLORI v2.0 revealed concordance of 28% at a site level and 85% at a gene level, as well as a mild correlation on m6A level estimations. Notably, NP-mFinder achieved 93% precision in detecting m6A within the “AAAAA” sequence context in the mRNA exonic region of HEK293T DRS data when compared to high-confidence m6A site annotation in GLORI v2.0, demonstrating the good performance of our method in the region possessing a stretch of continuous A-sequences. Moreover, our method predicted that m6A might exist in the human HEK293 poly(A) region, suggesting a possibly conserved phenomenon of a modified poly(A) tail beyond the previously reported T. brucei variant surface glycoprotein (VSG) transcripts. Together, these results established NP-mFinder as a robust and versatile tool for transcriptome-wide m6A profiling with DRS data at single-read resolution.
KW - Nanopore direct RNA sequencing
KW - epitranscriptomics
KW - mA
KW - machine learning
KW - poly(A)
UR - https://www.scopus.com/pages/publications/105043264590
U2 - 10.3389/fgene.2026.1770769
DO - 10.3389/fgene.2026.1770769
M3 - Article
AN - SCOPUS:105043264590
SN - 1664-8021
VL - 17
JO - Frontiers in Genetics
JF - Frontiers in Genetics
M1 - 1770769
ER -