Environments
Environments built from real work.
RL environments and evals modelled on the workflows we deploy into. Tasks written and graded by practitioners.
Why it matters
Benchmarks test what models know. Work tests what they can do.
Real work is scanned forms, legacy screens, approval limits, and Arabic and English on the same page. Our environments keep all of it, so what a model learns here holds up in production.
How we measure
Better than out of the box.
An off-the-shelf model is a starting point. We build the evals and the harness around each task, then tune the model against them.
| Benchmark | HL1 | Model A | Model B | Model C |
|---|---|---|---|---|
| Task success (Correct on the first attempt) | 78.4% | 61.2% | 54.8% | 47.5% |
| Consistency (Correct in all five runs) | 64.1% | 38.6% | 31.2% | 24.9% |
| Tool-call accuracy (Valid calls to the client’s systems) | 96.2% | 89.4% | 86.1% | 81.7% |
| Approval adherence (Stopped for sign-off when the rules require it) | 99.1% | 92.3% | 90.8% | 87.5% |
| Evidence cited (Claims traced to a source document) | 91.5% | 74.0% | 69.3% | 61.8% |
| HL1 | Low: 69.8% at $4.5 | Med: 75.6% at $6.5 | High: 78.4% at $9 |
|---|---|---|---|
| Model A | Low: 56.4% at $20 | Med: 59.3% at $29 | High: 61.2% at $42 |
| Model B | Low: 45.2% at $6 | Med: 51.0% at $9.5 | High: 54.8% at $15 |
| Model C | Low: 40.2% at $2.4 | Med: 45.1% at $3.2 | High: 47.5% at $4 |
Illustrative. Model names are anonymised and every figure is an example, not a measured result.
What we deliver
Environments. Evals. Experts.
Environments
Stateful replicas of ERPs, portals, permit systems and inboxes.
- erp.search_po()
- portal.get_delivery()
- permits.check_rule()
- mail.draft()
Evals
Task suites scored by deterministic checks and expert rubrics.
- Model A61.2%
- Model B54.8%
- Model C47.5%
Illustrative scores
Expert network
Practitioners who write the tasks, show the work and grade the outputs.
- Finance controllers
- Procurement leads
- Quantity surveyors
- Permit officers
- Facilities managers
Domains
Where our environments come from.
Across the Gulf. In English and العربية.
Finance operations
Invoice exceptions, reconciliation, close
Procurement
Purchase approvals, supplier onboarding, contract checks
Public services
Permit review, case handling, inspections
Construction
Variation claims, payment applications, contract notices
Property and facilities
Service charges, work-order verification
Logistics
Customs documents, proof of delivery, claims
Synthetic by design.
No client data leaves the client.
Synthetic records, modelled on real workflows. Nothing a client shares trains a model.
Book a call
Tell us where your models fall short.
We’ll scope the environment, the tasks and the graders.
hello@hiddenlever.ai