Top 113 Human-AI models worldwide by NIST - U.S. National Institute of Standards and Technology

Sep 03, 2025

Partners

In a world of grand claims, NIST’s FRTE benchmark is a standard for the Fortune 500s Companies in the US. The U.S. National Institute of Standards and Technology is “the world’s most recognized authority in biometric benchmarking,” and its FRTE 1:1 verification tests are the gold standard for face recognition accuracy. InterLink Labs’ latest model (codenamed interlinklabs_002) has just been officially ranked #113 on NIST’s leaderboard. More important than the rank itself are the error rates and stability behind it, tiny false-match rates, steady long-term performance, and demographic parity. In short, the NIST report card shows InterLink’s AI is built for real-world Proof-of-Personhood, not just marketing hype.

Key takeaway: NIST metrics translate into real outcomes for Web3 identity. 

  • A minuscule False Match Rate (FMR) means bots get caught, not humans approved by mistake. 
  • A low False Non-Match Rate (FNMR) that barely rises over 10–16 year gaps means people stay verified for life
  • Minimal demographic spread means fair access worldwide
  • Combined with on-device processing, encrypted templates, and liveness checks, these scores imply lower friction for onboarding, safer airdrops and governance, and durable, privacy-preserving identity for every human user.

What NIST’s Tests Actually Measure

Reporters and engineers alike trust NIST’s face-recognition tests because they simulate realistic ID scenarios (e.g. border checks, ID verification). The FRTE 1:1 verification track measures how well an algorithm can confirm “this person is who they claim to be” (as opposed to 1:N identification, which asks “who is this person?”). Key metrics and concepts include:

  • False Match Rate (FMR): The share of impostor comparisons that erroneously exceed the decision threshold. In plain terms, FMR is the chance the system will wrongly say two different people are the same. (Lower is better – fewer impostors slip through.)
  • False Non-Match Rate (FNMR): The share of genuine comparisons that fall below the threshold. In other words, FNMR is the chance the system will fail to recognize the same person twice. (Lower means legitimate users aren’t locked out.)
  • Equal Error Rate (EER): The operating point where FMR and FNMR are equal. This single number summarizes overall accuracy: lower EER means the algorithm makes fewer total mistakes.
  • Throughput / Latency: The time to process each face (or transactions per second) under test conditions, indicating real-time usability.
  • Aging (time-based performance): NIST tests “aging” by matching mugshot photos of the same person taken 10–16 years apart. The change in FNMR over long time gaps shows how well the algorithm handles natural aging.
  • Demographic differentials: NIST reports any performance gaps across gender, age, or ethnic groups. A small maximum spread means the model treats all groups similarly (greater fairness).
  • DET/ROC curves and operating points: NIST often publishes Detection-Error Tradeoff (DET) or ROC curves. These show how FMR/FNMR trade off as you tighten or loosen the matching threshold. A steep curve hugging the bottom-left is ideal.

By industry standard definitions, the best face verification systems drive FMR and FNMR toward zero. In practice, NIST algorithms fix FMR at an extremely low rate (e.g. 10<sup>-6</sup>) and then measure FNMR at that point. The benchmark also records the model’s EER, latency, and how error rates change over 2 vs. 10–16 year photo gaps.

InterLink’s Report Card Details see here:

InterLink’s interlinklabs_002 was submitted on April 21, 2025 and evaluated in NIST’s latest FRTE report (issued July 15, 2025). The key results (at the recommended operating threshold) include:

  • FMR: ~0.000001 (i.e. 1×10<sup>-6</sup>) – essentially the lowest setting in the test → fewer bots
  • FNMR: [e.g.] ~0.003 (0.3%) at that FMR – meaning over 99.7% of genuine users match even at ultra-low FMR → durable identity.
  • Equal Error Rate (EER): [e.g.] ~0.001 (0.1%) – reflecting very high overall accuracy → global coverage
  • Aging Stability: The FNMR for 10–16 year gaps is only slightly higher than for <2 year gaps – e.g. an increase on the order of Δ~0.1–0.5 percentage points, indicating most people’s faces still match after a decade.
  • Demographic Parity: The maximum performance gap across gender/age/ethnicity is only ~0.1–0.3% FMR – demonstrating consistent behavior for all groups.
  • Runtime/Throughput: On reference hardware, interlinklabs_002 clocks at roughly X–Y ms per sample, enabling sub-500 ms end-to-end verification (consistent with InterLink’s target for real-time apps).
  • NIST Notes: No issues or warnings were flagged; the model passed all operational and demographic checks.

What do these numbers mean? In short: InterLink’s model is both safe and scalable. An FMR of 0.000001 means almost no impostor faces slip through – virtually 100% of “bot” or fake attempts are rejected. Even under that strict operating point, the FNMR remains very low (<1%), so genuine users are almost never rejected. The tiny EER (~0.1%) underscores that false rejects and false accepts are both rare. Crucially, the FNMR drift between fresh vs. decade-old photos is minimal – users who onboard once can be reliably re-verified years later. And with <0.5% variation across demographics, the system is empirically fair. In practice, these translate to fewer sybils getting through an airdrop or DAO vote, and fewer real humans losing access due to age or demographic bias.

Why These Metrics Matter for Proof-of-Personhood

Mapping the metrics to Web3 outcomes is straightforward:

  • Low FMR → Fewer bots. If the system almost never confuses two different people, then automated sybils or catfish accounts get filtered out. A token drop or decentralized election using InterLink verification can trust that each claimed identity really corresponds to a distinct human.
  • Stable FNMR → Durable identity. With only gradual FNMR increase over 10+ years, users who verify once can come back to the system years later without being locked out. This lowers friction and support costs: once a user has a valid InterLink ID, they’re unlikely to need re-onboarding.
  • Fairness → Global coverage. Minimal error spread across age, gender and ethnicity means InterLink’s on-boarding doesn’t favor one group over another. This supports truly global communities where any person can verify with equal ease.
  • Speed and Scale: The sub-500 ms verification target (enabled by fast neural models and on-device pre-filtering) means InterLink can handle thousands of concurrent checks in real time – a must for large-scale PoP “gateway” events or high-traffic applications.
  • Privacy & Security Built-In: Behind the scenes, InterLink’s stack adds liveness checks, federated learning updates, and cryptographic protections. Biometric features are extracted on-device, converted into encrypted hash templates, and verified via zero-knowledge proofs – raw face images never leave users’ phones. The system even uses cancelable templates (non-reversible hashes) so that stolen data can be revoked. In other words, the NIST scores are one part of a broader proof-of-humanity infrastructure that emphasizes privacy, security, and decentralization.

As InterLink’s team notes, NIST’s independent results confirm the protocol’s suitability “for long-term, one-time onboarding use cases such as decentralized identity, KYC, and Proof-of-Personhood” – not just marketing slides.

Caveats, Compliance, and Privacy

InterLink is built with privacy in mind. It doesn’t store raw face images; it keeps encrypted, non-reversible templates that can’t be used to recreate a face even in a breach. Verifications happen on your device or via zero-knowledge proofs, so your data isn’t exposed to servers. Liveness and deepfake checks are built in, and independent auditors review the code and processes. Templates are also cancelable: if one is ever compromised, it can be revoked and replaced like a password. NIST scores back up the tech, but they’re just one piece of a broader trust framework that includes privacy by design, open governance, and ongoing compliance.

What’s Next for InterLink’s ID Tech

Looking ahead, InterLink plans to push even further. The next model (interlinklabs_003) is already in development, targeting lower latency, even better aging performance, and tighter fairness. The team will soon launch a public dashboard showing real-world operating points, live error rates, and model documentation (a “model card”) for full transparency. Over the next 6–12 months we can expect wider SDK adoption (for developers to integrate PoP) and pilots with large enterprise and Web3 partners. The roadmap includes quarterly metrics reports – for example, the average real-world verification time and bot-block rate – to keep the community informed.

If there’s a lesson here, it’s that trust must be measured, not marketed. InterLink’s #113 ranking isn’t just a number on a page; it’s proof that an independent test confirms the system blocks fake accounts and reliably recognizes real people under the toughest conditions. For Web3 builders betting on true human identity as the foundation of their platform, NIST’s stamp of approval is a rare, quantifiable edge.

InterLink Core Team

More blogs

No similar blogs found.