Skip to content

Environments

Environments built from real work.

RL environments and evals modelled on the workflows we deploy into. Tasks written and graded by practitioners.

Why it matters

Benchmarks test what models know. Work tests what they can do.

Real work is scanned forms, legacy screens, approval limits, and Arabic and English on the same page. Our environments keep all of it, so what a model learns here holds up in production.

How we measure

Better than out of the box.

An off-the-shelf model is a starting point. We build the evals and the harness around each task, then tune the model against them.

Benchmark results. Illustrative. Model names are anonymised and every figure is an example, not a measured result.
BenchmarkHL1Model AModel BModel C
Task success (Correct on the first attempt)78.4%61.2%54.8%47.5%
Consistency (Correct in all five runs)64.1%38.6%31.2%24.9%
Tool-call accuracy (Valid calls to the client’s systems)96.2%89.4%86.1%81.7%
Approval adherence (Stopped for sign-off when the rules require it)99.1%92.3%90.8%87.5%
Evidence cited (Claims traced to a source document)91.5%74.0%69.3%61.8%
Accuracy vs cost: Task success (%) against cost per 1,000 tasks (usd, log scale), at low, med, high effort.
HL1Low: 69.8% at $4.5Med: 75.6% at $6.5High: 78.4% at $9
Model ALow: 56.4% at $20Med: 59.3% at $29High: 61.2% at $42
Model BLow: 45.2% at $6Med: 51.0% at $9.5High: 54.8% at $15
Model CLow: 40.2% at $2.4Med: 45.1% at $3.2High: 47.5% at $4

Illustrative. Model names are anonymised and every figure is an example, not a measured result.

What we deliver

Environments. Evals. Experts.

  • Environments

    Stateful replicas of ERPs, portals, permit systems and inboxes.

    • erp.search_po()
    • portal.get_delivery()
    • permits.check_rule()
    • mail.draft()
  • Evals

    Task suites scored by deterministic checks and expert rubrics.

    • Model A61.2%
    • Model B54.8%
    • Model C47.5%

    Illustrative scores

  • Expert network

    Practitioners who write the tasks, show the work and grade the outputs.

    • Finance controllers
    • Procurement leads
    • Quantity surveyors
    • Permit officers
    • Facilities managers

Domains

Where our environments come from.

Across the Gulf. In English and العربية.

  • Finance operations

    Invoice exceptions, reconciliation, close

  • Procurement

    Purchase approvals, supplier onboarding, contract checks

  • Public services

    Permit review, case handling, inspections

  • Construction

    Variation claims, payment applications, contract notices

  • Property and facilities

    Service charges, work-order verification

  • Logistics

    Customs documents, proof of delivery, claims

Synthetic by design.
No client data leaves the client.

Synthetic records, modelled on real workflows. Nothing a client shares trains a model.

Book a call

Tell us where your models fall short.

We’ll scope the environment, the tasks and the graders.

hello@hiddenlever.ai
Organisation type
What do you need?

A line or two is enough. Please don’t include confidential records.

We use your details only to reply. Privacy notice.