Every number here is reproducible. This is how.
Public, versioned, reproducible. The scoring matches the shipped engine, and every result can be replayed from the raw stream.
RTKTEST operates fixed reference receivers (“probes”) at coordinates surveyed to sub-centimetre accuracy. When a client starts a session, their NTRIP credentials authenticate directly against their provider’s caster. Each correction stream is consumed by the probe nearest the client’s stated location, with the position uplink (GGA) set to the probe coordinates.
The probe’s RTK engine computes an independent position solution every second. Because the true position of the probe is known, accuracy is measured against ground truth, not estimated from solution covariance. Probes on the public network run identical hardware, firmware and antenna models; partner probes are third-party stations and do not, which is set out in the next section.
Most probes sit on a public network. Their observation stream comes from an open reference network (EUREF, IGS), the engine is RTKLIB, and the probe coordinate is published on the map. Everything the measurement needs is available to anyone, so the whole procedure can be run again outside this app: take the same public stream, the same open engine, the same published coordinate, subscribe to the service yourself, and check the numbers. That is what makes this methodology verifiable rather than merely stated.
Some probes sit on a partner network. These are stations of a private operator, reached through an agreement, and their stream is not open. The measurement is made the same way and scored the same way, but a third party cannot replay it without the same agreement. They are a coverage decision: they let the service open in regions where no public stream of usable quality is available today. North America ran on partner probes for that reason until August 2026, when public stations there became usable and took over.
North America measures in NAD83(2011) at epoch 2010.0, the frame and the date a US surveyor works in. A station’s coordinate is only meaningful at a date, because the ground moves — 15 mm a year across the stable interior of the continent, up to 45 mm on the San Andreas fault — and a correction service realises the national frame at an epoch of its own choosing. Naming ours settles the question instead of hiding it: a service that realises the current epoch instead of the official one will show that difference in its trueness, and that is a finding about the service, not an artefact of our reference.
For 46 of the 77 North American probes the coordinate is the antenna reference point published by the National Geodetic Survey, adopted as published. Anyone can look it up by station name and check it against the figures here. Those probes are production. Where the NGS publishes nothing for a station, or where its published value is contradicted by both the live stream and the agency’s own daily solution, the coordinate is derived instead: the broadcast position carried back to 2010.0 with a published station velocity. That costs a few centimetres of trueness, so those 31 probes were labelled beta. Ahead of the public release they were withdrawn from the map instead of being offered with a caveat attached: a number we have to explain away is not a number worth handing over. Every North American probe you can select today is a production one.
What the caveat did and did not affect, while those probes were live. Inside one session the comparison between services was untouched — same station, same seconds, same coordinate for everyone — and so was repeatability. What carried the caveat was the absolute trueness of a beta figure, which is not comparable with a production one. That is the asymmetry that decided the withdrawal. Europe is production for the same reason North America now is: open physical stations, four constellations, and a coordinate published by a geodetic agency in the same frame as the services.
Sessions that were measured on a beta probe remain excluded from the public leaderboard, for the same reason the fixed reference and self-referenced legs are: a board is where two numbers get compared, and these two are not built the same way. They were delivered in full to the person who ran them, and the exclusion rule stays in the engine, which is what any future beta probe would be measured under.
The two kinds are drawn in different colours on the probe map, and every result states which one it was measured on. Two consequences are worth stating plainly. Partner probes are third-party stations, so the identical-hardware property of the public network does not hold across them, which matters when comparing probes rather than services within one session. And their coordinate is supplied by the operator of the network, not by a geodetic agency, so it carries that operator’s realisation of the reference frame.
A session tests up to three services at once. Every stream is consumed simultaneously by the same probe, processing the same GNSS observations from the same antenna during the same epochs. Atmosphere, multipath, and satellite geometry are identical by construction. Any difference in the result is attributable to the correction service.
Every service runs under the same rate limits, the same engine settings, and the same scoring. No network ever serves as the yardstick: the only ground truth in this methodology is the surveyed monument. That includes the reference service described below, which is measured like any other and never used to calibrate, normalise or correct a result.
Every session also measures one fixed reference service, RTK Premium, operated by Premium Positioning, on the same rover stream, the same ephemeris, the same epochs and the same surveyed truth point as the services under test. Absolute scores are not directly comparable between sessions, which are run at different places, on different days and under different ionospheric conditions. A service measured in every session is a common anchor against which sessions can be related.
The reference is not ground truth and is not a yardstick. It is reported in its own block, never ranked against the services you chose, and excluded from the public leaderboard, which aggregates only services users selected. When a user tests RTK Premium on their own initiative, that leg is ranked like any other.
The reference is operated by the company that operates this tool, so the safeguards are structural rather than declarative: same engine, same configuration, same epochs, never ranked, excluded from the published aggregates, and raw data archived so any third party can recompute it.
Every session reports three statistics per component (East-West, North-South, Up-Down): AVE, the mean deviation from surveyed truth (trueness); STD, the dispersion (repeatability); and RMS, the combined error. The distinction matters: a large mean with small scatter is a deterministic bias in the correction model, not measurement noise, and it is invisible to any metric based on solution covariance alone.
The composite score weights four components, computed on fixed-solution epochs, and the four weights sum to 100. The RTK engine is open source and its exact configuration file is published with each methodology version, so any result can be reproduced from the raw stream.
Accuracy and precision are scored per component, horizontal and vertical separately, never as one combined 3D figure. A service that holds 7 mm horizontally and 2.6 cm vertically is a real and common shape, and rolling the two into a single 3D number charges the vertical twice while giving the horizontal no credit at all.
Neither converts in a straight line. Each maps through a knee: flat over the range where a service is doing well, steep through a reference value, and small beyond it. For NRTK the reference is 3 cm for accuracy and 2 cm for precision, on each component. Read off the curve, a component holds above 95 up to 1.44 cm, above 90 up to 1.73 cm, above 80 up to 2.12 cm, reaches 50 at 3 cm and 20 at 4.24 cm.
The reason for the knee is that the two ends of the range are not equally interesting. A few millimetres between two services that both hold well inside tolerance is not a difference a user will meet in the field, and it should not decide a ranking: 1.2 cm scores 97.4 and 1.5 cm scores 94.1. Five centimetres on one component is a different matter, and it scores 10.7, not a passing mark. A straight line, or the exponential this replaced, gets both of those wrong at once: too severe where the services are good, too generous where one is far out.
The two components are combined into one sub-score by a harmonic mean, which is closer to the worse of the two than an average is. That choice is not cosmetic. Under a plain average, one perfect component floors the pair at 50 whatever the other one does, and a service measured here at 0.66 cm horizontally and 18.9 cm vertically came out with a composite of 67.5. An 18 cm vertical error is not a two-thirds service. Both components are reported in the result next to the combined figure, so nothing hides inside the combination.
The accuracy references are the per-component expectation this platform works to, 2 to 3 cm horizontal and 2 to 3 cm vertical for NRTK, taken at the top of that band, so a service sitting exactly at the tolerance limit scores 50 on each component. The precision references are ours, and we would rather say so than imply a standard we are not quoting: no industry figure fixes an acceptable scatter. If you disagree with where any of these sit, the numbers are published above and every raw statistic behind a score is in the session record, so you can put the bar elsewhere and check.
Correction latency is reported as a connection-health flag, not a scored component. The correction age between two synchronous 1 Hz streams is quantised at roughly one second and is not a representative latency, so scoring it would distort the result. A leg is flagged when its correction age exceeds five seconds.
A service is identified from the caster address and mountpoint it broadcasts on — never from a username or password, which are never stored. A network appears on the public leaderboard only after 40 independent sessions in a region within the quarter, anonymised by default. An operator may claim its listing after verifying ownership and consenting to publication. Individual client sessions are never published or attributed.
The number of correction services has grown fast. Comparing them has not become easier. Coverage maps all look complete, every service delivers a FIX, and the differences that matter — how fast, how stable, how accurate under real conditions — stay invisible until you are in the field.
We think buying decisions in this industry should be based on measured performance, not promises.
That is why we built RTKTEST. Live sessions against surveyed reference points, every service tested simultaneously under identical conditions, and all raw data archived so the results can be verified independently.
To make sure the method holds up to scientific scrutiny, we developed the evaluation methodology in cooperation with Geo++, the industry reference for GNSS quality.
A partnership confers no advantage in a measurement. No RTK provider is ranked differently, sees another operator’s data, or influences a score. The only ground truth remains the surveyed monument, and the safeguards and exclusions set out above apply without exception.
The observations a probe delivers are not ours. Each session runs on a reference station operated by a public geodetic network, and every result we publish or hand over is a derivative work of that network’s data. Attribution is a licence condition, not a courtesy, so it travels with the result: the session report names the network the observations came from, and the networks are listed here.
- EUREF EPN — Rover observations from the EUREF Permanent GNSS Network (EPN), provided by its station operators through the EPN Central Bureau.
- IGS — Rover observations from the International GNSS Service (IGS) network, provided by its station operators and data centres.
- EarthScope — Rover observations from the EarthScope Consortium GNSS network (NOTA/GAGE), used under the EarthScope User License Agreement and attributed under CC BY 4.0.
- NOAA NGS CORS — Rover observations from the NOAA National Geodetic Survey CORS network.
- partner network — Rover observations from a partner reference network, used under agreement. The station coordinate is the network's own realisation, not an independent geodetic solution — see the methodology.
None of these networks endorses RTKTEST, reviews its results, or is responsible for them. They provide observations; the method, the scoring and the conclusions are ours.
A probe measures network performance at the probe location, not across the entire service area. Ionospheric conditions vary between sessions, which is why single sessions carry an indicative rank only and the leaderboard requires session minimums. VRS streams are initialised at the probe position, which may differ from a client’s working position by several kilometres.
We publish limitations because a benchmark that hides its weaknesses is marketing. Corrections to this methodology are versioned and archived. v1.1 (2026-08-25): added the fixed reference service. v1.2 (2026-08-25): stated the split between public and partner probes, and narrowed the identical-hardware claim to the public network. v1.3 (2026-08-26): North America labelled beta, with the reason, and beta sessions excluded from the public leaderboard. v1.4 (2026-08-26): named the Geo++ partnership, its purpose, and the boundary that a partnership confers no advantage in a measurement. v1.5 (2026-08-27): named the geodetic networks whose observations the probes deliver, and the attribution their licences require, and reworded the partnership block. v1.6 (2026-08-27): North America moved to public stations; it stays beta for the reference-epoch reason, which replaces the partner-stream one. v1.7 (2026-08-27): North America measures in NAD83(2011) at epoch 2010.0, the surveyor’s frame, and the 46 probes whose coordinate the NGS publishes became production; beta now means a coordinate we had to derive. v1.9 (2026-09-02): precision reweighted from 15 to 25, and accuracy and precision moved from an exponential to a knee, so that a small difference between two good services stops deciding rankings and a result outside RTK tolerance is scored as such. Every stored session was rescored on the new basis, so the board and the archive stay on one scale. v1.10 (2026-09-02): accuracy and precision are scored per component against the per-component tolerance, horizontal and vertical separately, combined by a harmonic mean, instead of on one combined 3D figure that charged a vertical weakness twice and credited a strong horizontal not at all. The archive was rescored again. v1.8 (2026-08-31): the beta probes withdrawn ahead of the public release, so every probe on the map carries a coordinate published by a geodetic agency and no session is measured against a derived one.