Benchmark № 001 — methodology

Frozen before results.

Every scoring rule below was fixed before any vendor was queried. The full document is published in a public repository, tagged, with its SHA-256 posted here — so anyone can verify nothing changes after results come back. Vendors are invited to dispute the method before publication.

Status: methodology v1.0 frozen 2026-07-25 — tag methodology-v1.0. Verify: shasum -a 256 METHODOLOGY.md.

Verification record
methodology v1.0 · frozen ✓ 2026-07-25
sha256
24fb254ed62b0aef493520e6…f636ee34969d
repo: boundstone/phone-validation-benchmark
dispute window: 14 days, all vendors

Vendors under test

Ordinary public self-serve accounts, paid at list price. No sales contact, no negotiated tiers, no advance notice.

VendorProductTier
1LookupPhone Validationself-serve credits
TrestlePhone Validation API v1$0.015/query PAYG
TwilioLookup v2 + Line Type IntelligencePAYG
NumVerifyStandard APIpaid plan

Boundstone's own API is not in run № 001 — self-exclusion: we author and fund this benchmark, and our product launched only days before the freeze. From the first quarterly re-run after launch, it is tested under this identical protocol and scored in public, win or lose.

The dataset — 1,000 numbers, ground truth stated per category

Synthetic, owned, or written-consent numbers only. No scraped or purchased lists. No calls or texts are ever placed — API queries only.

CategoryCountGround truthScored on
Owned active numbers (two carriers)40strongvalidity · line type · carrier
Self-created VoIP / app numbers25strongvalidity · line type · carrier
Opt-in volunteer numbers (written consent)135strongvalidity · line type · carrier
Owned-then-released (carrier aging)15weakdescriptive only
Ported by us between carriers5moderateported status
Unallocated prefixes (pinned NANPA snapshot)350strongvalidity
Fictional range 555-0100…0199125strongvalidity
Impossible formats (NANP rules)150strongvalidity
Valid format, unknown subscriber155format strong · liveness unknownvalidity; liveness descriptive

Scoring rules that cannot change after freeze

§1Answers score correct (1.0), partial (0.5), wrong (0), or abstain — and abstain is excluded, not penalized. Not knowing is honest; being wrong is not.
§2Partial credit exists in only three pre-named cases: right line-type family wrong subtype; MVNO host network for MVNO brand; one-port-stale carrier.
§3Validity means well-formed and allocated — not "has an active subscriber." Calling a real-but-idle number invalid is wrong.
§4Headline composite = the unweighted mean of validity, line-type, and carrier accuracy, with 95% bootstrap confidence intervals. No other aggregate gets invented after results are known.
§5Abstention above 20% on any dimension is flagged in every table where that number appears.
§6Liveness and disconnected-number outputs are reported descriptively — no accuracy claims in either direction.
§7All four vendors queried inside the same 72-hour window; raw responses archived and available to a neutral auditor under NDA.
§8Volunteer numbers are published redacted (row ID + last four digits); names never. Synthetic and owned rows are published in full.
§9Every vendor gets its per-row verdicts 14 days before publication; corrections land in a public errata log — never as silent edits.
§10Conflict declared: we fund this benchmark and compete in this market. That is exactly why every rule above is frozen first.

Get the results when they drop

One email when the benchmark publishes, plus early API access.

The API is live — start free without waiting. This list gets benchmark № 001 the day it publishes. One email. We validate emails for a living — we're not going to spam yours.