HelixCore · Precision Genomics. Unlimited Power. The twelve modules
Sign in
FILE OO–2026–08
ENGINE minimap2-lca/3.0.0
CLASS. TECHNICAL · PUBLIC
TESTS 769/769
MEASURED 2026–08–10
03 Ingestion · metagenomics and metabarcoding

OmniOta

On material whose composition is certified by the manufacturer it reaches 99.08% accuracy at genus and 97.24% at species, naming eight taxa and getting all eight right, without a single false positive.

And where the sample arrives in poor condition — degraded DNA, ultra-processed food, an exhausted flow cell — it still returns a genus-level result with 97.70% accuracy, on datasets where the state of the art, which identifies by exact match, does not return a single line.

What holds the figure up
Every number on this page carries its denominator, its chemistry and its year. The "what is not claimed" sections are not small print: they are the reason the rest is credible.
minimap2-lca/3.0.0 + arbitration-bayes/1.0.0
769 tests green · measured 2026-08-10
99,08 %
accuracy at genus on a manufacturer-certified community
97,24 %
accuracy at species, with 15,549 reads resolved out of fifty thousand
8 / 8
taxa named and correct, without a single false positive
22,514
reference genomes, one per prokaryotic species
01 Certified composition

The ground truth is not ours to set. The manufacturer certifies it.

ZymoBIOMICS D6300 — eight bacteria at 12% genomic DNA and two yeasts at 2%, sequenced by Nicholls, Quick, Tang and Loman in GigaScience (2019) on two different platforms. Fifty thousand reads per platform.
Measurement
GridION
PromethION
Genus-level accuracy
99,08 %
99,00 %
Exactitud a especie
97,24 %
97,14 %
Reads assigned at genus
29,824
27,928
Reads resolved to species
15,549
14,731
Largest false positive
0,17 %
0,21 %
Cross-platform invariance
The same community on two sequencers gives an L1 divergence of 0,0181 against the 0.0257 produced by sampling itself, with zero divergent taxa. Not "close to zero": statistically indistinguishable from chance, which is what had to be shown.
Divergence is tested against the divergence sampling itself produces, simulated with the real counts — not against zero. Two subsamples of the same population do not give zero either.
And what is not reproducible
Low-abundance artefacts do not repeat: Macrococcus appears with 23 reads on one platform and 5 on the other. The composition repeats; the traces do not.
It is published because it is a direct warning about how much weight to give a taxon held up by twenty reads — and because it prompted the correction of the naming floor.
02 The competitor's own ground

Their dataset, their metric, their table.

Brown et al. 2017: seven nanopore datasets with declared composition, the metric defined by the authors of the reference tool — correct reads over classified reads — at genus rank. Same files, same accession numbers.
Dataset · genus rank
OmniOta
GAIA
Kraken
One Codex
MG-RAST
Mean across the seven datasets
97,70 %
96,37 %
94,67 %
94,34 %
77,77 %
P. fluorescens — the hardest
94,64 %
85,83 %
84,6 %
84,2 %
84,9 %
M. aeruginosa
96,58 %
96,39 %
85,8 %
95,1 %
53,1 %
Even mixture of four
97,70 %
93,47 %
97,6 %
87,4 %
65,0 %
S. elongatus
100,00 %
100,0 %
98,1 %
97,6 %
87,9 %
The largest advantage appears on the hardest dataset: Pseudomonas fluorescens, where the four published tools land between 84.2% and 85.8% and OmniOta reaches 94.64%.
The aggregate figure — 896 correct reads out of 932 classified, without averaging percentages — is 96.1%, with a 95% confidence interval between 94.7% and 97.2%. Both are published: the per-dataset mean because it is the only one comparable with the published competitor figures, and the aggregate because it is the statistically correct one.
What these figures do NOT license anyone to say
The +1.33 point advantage over the competitor closest to OmniOta's performance does not reach statistical significance with seven pairs: paired t 0.823 against a critical 2.447; sign test p = 0.375. The highest value in the table corresponds to the one obtained by OmniOta; the rigorous course is not to claim the advantage is systematic.
A difference not being significant does not mean the tools are equivalent either: that would require an equivalence test nobody has run. For products that are not open or not accessible — closed services with a proprietary reference — the values used are those of the same benchmarks as officially published by their authors.
FIG. 01–02 · Degraded samples
03 The limit of the degraded sample

When the sample arrives in poor condition, the competition returns nothing.

0 /7
degraded datasets in which sylph returns a single line. None.
Nature Biotechnology 2024, the state of the art, run in-house against exactly the same 22,514 reference genomes. It is not a run failure nor an unfavourable configuration: the tool declares it itself.
The cause is structural and it can be computed. These methods identify by looking for 31-base stretches that match exactly the reference genome. When DNA arrives broken or the read accumulates errors, those exact stretches stop existing: with reads matching on only 74% of their bases, the probability that a 31-base stretch survives intact is 0.74³¹. There is nothing left to count. An alignment tolerates changed bases; an exact match does not.
When identity is mentioned here identity it means that and nothing else: what percentage of a read's bases match the reference. At 99% the sample is clean; at 74%, one base in four does not fit.
FIG. 01 Probability that a 31-base stretch arrives intact
1/1 1/10 1/100 1/1,000 1/10,000 99 % 95 % 90 % 74 % bases matching the reference → 1 in 1.4 1 in 5 1 in 26 1 in 11,319
0.74³¹ — logarithmic scale. Degraded material falls at the amber point.
And this is not argued: it is measured
Between a 2015 dataset and a 2019 one everything changes at once, so attributing the collapse to identity was not justified. It was settled by taking the same certified 2019 file and injecting error at increasing rates: everything else identical by construction. It could have come out against us.
And the variable is not the year: it is identity. A flow cell at the end of its life, DNA degraded by the thermal processing of the food, a difficult extraction from a fatty matrix or an old environmental sample all put a 2026 run in the same regime as a 2015 one.
The collapse reproduces on modern, certified material. More depth delays the limit but does not remove it: at 76.84% there is not a single taxon even with fifty thousand reads, and quadrupling the sequencing buys about four points of identity. The most degraded level of the experiment lands at 74.23% — exactly the identity of the 2015 material, which the model predicts without having been given it.
FIG. 02 Taxa detected by sylph
Measured identity
5,000 reads
20,000
50,000
88,77 % intacto
8 of 8
8
84,91 %
6
8
82,65 %
0
7
8
79,68 %
0
0
6
76,84 %
0
0
0
74,23 %
0
0
0
Depth in reads. Amber: not one taxon at any depth.
What happens to OmniOta in that same regime
−1,4 pts
At genus
97.70% on degraded material against 99.08% under ideal conditions.
−59,8 pts
A especie
37.4% against 97.24%. Declared as such: at that identity no tool holds up a species.
Yes
Does it return a result?
On all seven datasets. The competitor, on none of them and at no rank.
At genus rank performance stays practically intact on data where the competitor does not return a line. At species rank it degrades, and badly — and that is declared: at 74% identity no tool can hold up a species. That is what a laboratory needs to know before buying: with a difficult sample it will still get a genus-level answer — the one that triggers an action — and it will know the species is not reliable, instead of receiving an empty report or, worse, one full of invented names.
04 Why this decides matters in food

The difficult case is not the exception. Here it is the job.

All of the above would be a benchmark curiosity if degraded samples were rare. In food microbiology they are the usual regime, and the literature has documented it for more than a decade.
01
Direct sampling is mandatory, and culture alters the picture
Faced with severe or atypical spoilage the sample has to be read as it is: any enrichment selects whatever grows fast in that medium and erases the community's real proportions. The field itself accepted this years ago — culture-independent methods exist precisely because culture does not describe the ecosystem that spoils the product, and in a thoroughly sampled meat plant there appeared 74 as-yet undescribed bacterial taxa.
Cocolin et al. (2013) Int J Food Microbiol 167:29–43 · 10.1016/j.ijfoodmicro.2013.05.008
Xu et al. (2025) Microbiome 13:25 · 10.1186/s40168-024-02026-1
02
The food matrix comes loaded with inhibitors
Fats, calcium, polysaccharides, proteins, bile salts, phenolic compounds, metal ions: a catalogue of substances that degrade reaction performance and that force matrix-specific protocols. The documented consequence is not slightly less sensitivity: it is false negatives and less usable material, which is exactly what sinks read identity.
Schrader et al. (2012) J Appl Microbiol 113:1014–1026 · 10.1111/j.1365-2672.2012.05384.x
03
Processing fragments the DNA before it reaches the laboratory
Temperature, pH, pressure and time degrade the matrix DNA — reviewed to the point of making the analysis unviable or the quantification unreliable. An ultra-processed product does not deliver long, intact DNA: it delivers fragments. And a short fragment carrying errors is, by definition, a fragment with no exact stretches to offer.
Gryson (2010) Anal Bioanal Chem 396:2003–2022 · 10.1007/s00216-009-3343-2
04
And part of what must be found does not grow on a plate
Under stress — refrigeration, acid, disinfectant, starvation — relevant pathogens enter a viable but non-culturable state: they stay alive, recover their ability to infect on resuscitation, and do not appear on standard medium. If the organism you are looking for does not grow, the only possible reading is molecular. It is not a methodological preference: it is the only route left.
Oliver (2010) FEMS Microbiol Rev 34:415–425 · 10.1111/j.1574-6976.2009.00200.x
What follows from putting the four together
Mandatory direct sampling, a matrix full of inhibitors, already fragmented DNA and organisms that do not culture. All four conditions push in the same direction: reads that match the reference poorly. In this sector, the regime where exact-match state of the art returns nothing is not the rare case — it is the case that reaches the laboratory when something has genuinely gone wrong.
The practical consequence
Where the method breaks stops being a benchmark footnote and becomes a purchasing criterion: it defines which samples can be analysed. Severe, rare or urgent spoilage is exactly what brings in the worst sample. And that is exactly where an empty report costs an entire batch.
With material at that level of degradation OmniOta holds 97.70% at genus and declares the species unreliable. The first triggers the action; the second stops anyone over-defending it.
05 What really sets it apart

Refusing to assert.

Anyone can argue over a point of accuracy either way. What is unusual is a system that stays silent when the data does not support a call — and that says exactly how much more it would need in order to speak.

Telling "it is not there" from "I have not sampled it" is the difference between a measurement and a decorated guess.

Naming floor
From 62 genera to 8, without losing a real member
There was a criterion for assigning a read, but not for naming a taxon: the engine named 62 genera where there were 8, with 53 false positives adding up to just 0.82% of the reads — yet named in the report all the same. Corrected with a floor derived from the run itself. Taxa below it are not deleted: they are flagged.
Not resolvable
"The best match is 62.1% where 90% would be expected"
On the Listeria in the certified mock, the system refuses to name a lineage and explains why: the strain is probably not represented in the panel. Without that criterion it would have asserted a lineage with posterior probability 1.0000 on a 61% match.
How much is missing
"3 of 2,000 positions; about 0.7× more coverage would be needed"
Not "resequence", but exactly how much depth. And in balancing, what computation can solve is distinguished from the 65% of stops only resequencing fixes — so as not to charge for the latter.
Trazas
A pathogen at 0.1% shows up flagged, and the technician decides
A profiler reports eight things and that is it: whatever falls below its threshold does not exist in its output. OmniOta keeps 62 taxa, 54 flagged as trace, with their counts intact. The question "is there any of this in my batch?" is answered without hiding what is scarce and without presenting it as confirmed presence.
06 Engine versus car

Classification accuracy is a capability. It is not the system.

Kraken2, Bracken and sylph are command-line tools that do one thing very well. The comparable commercial offerings are analysis platforms. OmniOta is a diagnostic application inside a laboratory ecosystem.
5
Classification engines
Each with its own rule stamp, and selection assisted by platform, marker, matrix and quality control.
11
Marcadores soportados
Including a customer-defined one: matK, rbcL or COI do not force a change of platform.
9
Quality-control layers
Orthogonal modules, each with its own veto. A silent failure has to cross all nine to reach the report.
24 / 24
Taxonomic complexes
Each with its ISO standard or cited DOI, applied equally to every column of any comparison.
8,213
Infraspecific positions
Panel derived from 149 complete genomes, uncurated: it works with any species that has deposited genomes, without depending on a published scheme.
0,012×
Coverage required
To observe a hundred discriminant positions. A hundredth of a pass, far below what an assembly demands.
One engine does not serve every sample
A degraded 16S amplicon, a fungal ITS, a whole metagenome and a customer's own marker are not well classified by the same algorithm. A single-pipeline system forces the sample to adapt to the tool.
Per-read assignment, not just a profile
OmniOta says what each of the 29,824 reads belongs to with 99.18% accuracy. Without that there is no infraspecific resolution over a taxon's reads and no link to functional analysis.
The scale is derived from the run itself
The noise ceiling varies by 7% between datasets from the same study and the same chemistry. Any absolute threshold chosen for one would be miscalibrated for the rest.
07 Technical sheet
In
Raw reads from any platform — Illumina, Nanopore, Ion Torrent — detected by reading the file headers. Eleven-marker amplicon or whole metagenome.
Out
Taxonomy with declared rank and per-row evidence, diversity, abundance by mean coverage proportion, metabolic potential and Sankey diagrams, in a report of about ten pages with methods, figures and tables ready to publish. Direct hand-off to primer design.
Reference
22,514 genomes, one per prokaryotic species, at the best assembly level available: 1,118,270 sequences, 35.1 GB. Complete download without a single failure.
Traceability
A per-row rule stamp: two runs with different rules are known to be non-comparable. Versioned panels carrying their accessions and hash inside, pinned tool versions and published lockfiles. An automated test compares the published figures against the actual run.
Limit · chemistry
The most recent material measured is from 2019. ONT R10.4 and Illumina remain unmeasured.
Limit · competitors
sylph is run on an equal reference footing. Kraken2/Bracken, MetaPhlAn 4, Ganon and KMCP remain pending; CAMI II was ruled out with written justification. The competitor closest to OmniOta's performance has not been run: it is a closed product and its reference is proprietary; the values used are those of the same benchmark as officially published by its authors.
Limit · scope
The reference is bacterial: fungi and yeasts fall outside it. In the certified mock the two yeasts come out at 0.00% on both platforms — this is not an engine false negative, and the figure being identical on both confirms it.
Limit · calibration
No lineage certainty percentage is issued: it would require a calibration against known-lineage material that does not yet exist. On nanopore, the declared confidence overestimates by 1.9 points, that is published, and it is not extrapolated to Illumina.
Sources: Brown, Watson, Minot, Rivera and Franklin (2017) GigaScience 6(3) · Paytuví-Gallart et al. (2019) bioRxiv 804690 · Nicholls, Quick, Tang and Loman (2019) GigaScience 8(5):giz043 · Shaw and Yu (2024) Nature Biotechnology · ISO 13136:2012. Data from the European Nucleotide Archive, projects PRJEB8672 and PRJEB8716.

There is not a figure here you cannot check. And several we would have loved to include are missing.

Request access See the twelve modules
BIOTECNO.org · Vitoria-Gasteiz EU + Codex