Every flight data vendor claims global coverage, real-time positions, and high accuracy. None of those words mean anything until you measure them, and almost nobody publishes a methodology for doing so — so evaluations tend to be an afternoon of eyeballing a map, followed by a purchase decision worth six figures. This is the methodology we'd want you to run. It's vendor-neutral, it's runnable, and yes: run it against us too. We'd rather compete on measurements than adjectives.

Set up a fair test first

Three rules before any numbers, because violating them is how bake-offs produce nonsense:

  1. Sample simultaneously. Query every vendor within the same few seconds. Air traffic changes minute to minute; a comparison across different times measures the sky, not the vendor.
  2. Sample repeatedly over days, not once. Coverage and latency vary by time of day and day of week. A single afternoon tells you very little.
  3. Define your regions before you look. Pick the airspace your product actually cares about and commit to it in advance, so you're not unconsciously selecting regions that flatter a favourite.

Pick at least four region types: a dense terminal area (e.g. a bbox over a major hub), an en-route corridor at cruise altitude, a low-altitude / general-aviation area, and a remote or oceanic stretch. Vendors differ far more in the last two than the first.

Server racks in a machine room

Test 1 — Coverage

The question: how many aircraft does each vendor see in the same box at the same moment? Not the global marketing number — your box.

import time, requests, statistics

BASE = "https://skylink-api.p.rapidapi.com"
H = {"X-RapidAPI-Key": KEY, "X-RapidAPI-Host": "skylink-api.p.rapidapi.com"}

REGIONS = {
    "terminal_lhr":  "51.2,-0.8,51.7,-0.1",
    "enroute_natl":  "48,-30,58,-10",
    "lowalt_socal":  "33.4,-118.6,34.4,-117.4",
    "remote_pacific":"10,-160,25,-140",
}

def sample(bbox):
    t0 = time.perf_counter()
    r = requests.get(f"{BASE}/adsb/aircraft", headers=H, params={"bbox": bbox}, timeout=15)
    elapsed_ms = (time.perf_counter() - t0) * 1000
    r.raise_for_status()
    data = r.json()
    return data["aircraft"], elapsed_ms

for name, bbox in REGIONS.items():
    fleet, ms = sample(bbox)
    print(f"{name:16} aircraft={len(fleet):4}  response={ms:6.0f}ms")

Record counts per region per run. Compare medians across runs, not single samples. Also sanity-check the direction of any big gap: a vendor reporting far more aircraft isn't automatically better — check whether the extras are real aircraft or duplicate/ghost tracks (Test 3).

For a global sanity figure alongside your regional counts, GET /adsb/aircraft/statistics returns fleet-wide totals (total_aircraft, airborne, on_ground).

Test 2 — Latency and staleness

This is the test most evaluations get wrong, because they measure the wrong thing. There are two latencies:

  • Response latency — how fast the HTTP call returns. Easy to measure, and largely a function of your network and their edge.
  • Data staleness — how old the position is. This is the one that determines whether your operational picture is right, and it's the one nobody advertises.

Measure staleness from the record's own timestamp:

from datetime import datetime, timezone

def staleness_seconds(fleet):
    now = datetime.now(timezone.utc)
    ages = []
    for ac in fleet:
        ts = ac.get("last_seen")
        if not ts:
            continue
        seen = datetime.fromisoformat(ts.replace("Z", "+00:00"))
        if seen.tzinfo is None:
            seen = seen.replace(tzinfo=timezone.utc)
        ages.append((now - seen).total_seconds())
    return ages

fleet, _ = sample(REGIONS["terminal_lhr"])
ages = sorted(staleness_seconds(fleet))
p = lambda q: ages[min(int(len(ages) * q), len(ages) - 1)]
print(f"staleness p50={p(0.5):.1f}s  p95={p(0.95):.1f}s  p99={p(0.99):.1f}s")

Report percentiles, never averages. A mean staleness of 6 seconds hides a p99 of two minutes, and the p99 is what your users will complain about. Do the same for response latency (p50/p95/p99), and record them separately — a vendor can be fast to answer and slow to know.

A large antenna array under installation

Test 3 — Position quality

Coverage and latency say nothing about whether the positions are good. Poll one region every few seconds, group fixes by icao24, and look for physically implausible behaviour:

  • Impossible speeds between fixes. Compute great-circle distance between consecutive positions divided by elapsed time. A 900-knot airliner or a 3,000-knot jump is a merge/dedup artifact, not an aircraft.
  • Position jitter while stationary. An aircraft with is_on_ground: true should not wander tens of metres between fixes. Jitter here indicates receiver disagreement that isn't being reconciled.
  • Altitude discontinuities. Jumps of thousands of feet between consecutive fixes with no corresponding vertical_rate.
  • Ghost and duplicate tracks. The same aircraft appearing under two identities, or tracks that persist after an aircraft has landed. This is what inflated coverage counts often turn out to be.
  • Track dropouts. Count gaps: how often does an aircraft vanish for more than N seconds and return? Frequent dropouts break any stateful feature you build, like takeoff and landing detection.

Log each violation type as a rate — per 1,000 fixes — so the number is comparable across vendors and regions.

Test 4 — Enrichment completeness

If you need labelled flights rather than anonymous dots, measure what fraction of aircraft actually arrive with usable metadata:

fleet, _ = sample(REGIONS["terminal_lhr"])
n = len(fleet)
for field in ("callsign", "registration", "aircraft_type", "airline"):
    filled = sum(1 for ac in fleet if ac.get(field))
    print(f"{field:14} {filled/n:6.1%}")

A vendor returning positions only, with a 40% registration fill rate, means you're building and maintaining the enrichment layer yourself — a real cost that belongs in the build vs buy column, not hidden in the integration estimate.

The scorecard

Score each vendor per region, then weight by what your product needs:

CriterionMetricHow to weight it
CoverageMedian aircraft count, per region typeHigh if you serve remote/low-altitude airspace
Data stalenessp50 / p95 / p99 secondsCritical for operational decisions
Response latencyp50 / p95 / p99 msMatters most for interactive UIs
Position qualityViolations per 1,000 fixesHigh for analytics and safety-adjacent use
Enrichment fill% with registration / type / airlineHigh if you display labelled flights
Dropout rateGaps > N seconds per track-hourHigh for stateful/event detection
StabilityVariance of the above across a weekAlways

Then, and only then, put price next to it — and read the SLA to see which of those numbers the vendor will actually stand behind contractually. A great p99 with no commitment attached is a hope, not a guarantee.

Run it against us

We'd genuinely rather be measured than described. The free tier gives you 1,000 requests/month — enough to run this methodology across four regions for several days — and every field the code above uses (last_seen, is_on_ground, registration, aircraft_type, airline) is in the standard ADS-B response. If your evaluation needs more volume than the free tier allows, tell us it's for a bake-off and we'll sort out an evaluation key.

Grab one through the free trial, run the tests against every vendor on your shortlist, and buy on the numbers.