Skip to main content

KERVALT / BENCHMARKS

Independent tests that tell you what works

We run internal and independent comparisons of speed, accuracy, safety, and cost. All tests run in our secure European environment.

Live Results Reproducible EU-Hosted No Data Leaks

01 / LEADERBOARDS

Live results

Selected results from our active test suites. Full datasets and methodology are available to approved research partners.

Test Metric Top Score System
Connected Legal Reasoning Accuracy across linked clauses 0.91 Griot
Long-Document Search Find rate across 128,000 tokens 0.94 Tafari
Structured Financial Questions Exact-match accuracy 0.97 Elimu
Source Accuracy Verified source link rate 0.99 Kervalt combined
Adversarial Safety Harmful output blocked 0.96 Safety controls

Scores reflect Kervalt-hosted tests on European servers. Methodology papers are available in our publications.

02 / EVALUATION

What we measure

We test the things that matter when real people use the system.

Reasoning accuracy

We test how well the system connects facts across documents, long texts, and structured records.

Speed at scale

We measure response speed under real production load, not ideal lab conditions.

Safety under pressure

We test how the system behaves when someone tries to trick, mislead, or push it off policy.

03 / REALITY GAP

Why most AI test scores don't tell the whole story

Public leaderboards reward one-question-at-a-time tasks. Real contracts, health records, and supply chains require tracing connections across many documents.

Problem: Tests ignore real-world connections

Public leaderboards reward one-question-at-a-time tasks. Real contracts, health records, and supply chains require tracing connections across many documents.

Agitate: A high score can hide a real risk

A model that scores 92% on general questions can still miss a connected clause, exposing the firm to costly mistakes or regulatory action.

Solve: Tests that match real data

Kervalt tests search, connections, and records together. Results are reported per system and combined, with citations linked to sources.

NEXT STEP

Run your own comparison

Bring your datasets and compare model behavior on secure infrastructure with full source tracing.