KERVALT / BENCHMARKS
Independent tests that tell you what works
We run internal and independent comparisons of speed, accuracy, safety, and cost. All tests run in our secure European environment.
01 / LEADERBOARDS
Live results
Selected results from our active test suites. Full datasets and methodology are available to approved research partners.
| Test | Metric | Top Score | System |
|---|---|---|---|
| Connected Legal Reasoning | Accuracy across linked clauses | 0.91 | Griot |
| Long-Document Search | Find rate across 128,000 tokens | 0.94 | Tafari |
| Structured Financial Questions | Exact-match accuracy | 0.97 | Elimu |
| Source Accuracy | Verified source link rate | 0.99 | Kervalt combined |
| Adversarial Safety | Harmful output blocked | 0.96 | Safety controls |
Scores reflect Kervalt-hosted tests on European servers. Methodology papers are available in our publications.
02 / EVALUATION
What we measure
We test the things that matter when real people use the system.
Reasoning accuracy
We test how well the system connects facts across documents, long texts, and structured records.
Speed at scale
We measure response speed under real production load, not ideal lab conditions.
Safety under pressure
We test how the system behaves when someone tries to trick, mislead, or push it off policy.
03 / REALITY GAP
Why most AI test scores don't tell the whole story
Public leaderboards reward one-question-at-a-time tasks. Real contracts, health records, and supply chains require tracing connections across many documents.
Problem: Tests ignore real-world connections
Public leaderboards reward one-question-at-a-time tasks. Real contracts, health records, and supply chains require tracing connections across many documents.
Agitate: A high score can hide a real risk
A model that scores 92% on general questions can still miss a connected clause, exposing the firm to costly mistakes or regulatory action.
Solve: Tests that match real data
Kervalt tests search, connections, and records together. Results are reported per system and combined, with citations linked to sources.
NEXT STEP
Run your own comparison
Bring your datasets and compare model behavior on secure infrastructure with full source tracing.