CASE STUDIES
You don’t have access to {{ permsTarget }}

{{ permsBody }} If you think you should have access, ask your project lead to add you.

C1

Collected Documents

Real legal documents collected from consenting professionals, quality-gated, PII-redacted and reviewed, with expert-written prompts and rubrics on top. Delivered to a frontier AI lab.

SUMMARY
533
gold standard documents
692
verified participants
4,071
total docs collected
Data types

Capture types

Multiple types to demonstrate our capability

REAL WORLD
Commercial

Real-world documents from professionals.

REAL WORLD
Enterprise

A real-world enterprise doc set from a single entity.

REAL WORLD
Personal

Real-world documents from individuals.

TRAJECTORY
Lawyer

Created to a set brief, with full trajectory and prompt history.

OUTPUTS Redacted document Metadata Trajectory Prompt history
The people

Human Selection

Screened & verified with consent, ownership and NDA in place

Lawyers
Selected to write, markup or review documents.
530 eligible 210 verified & consenting
193
selected
Professionals
Selected to provide commercial documents they own.
492 eligible 380 verified & consenting
314
selected
Individuals
Selected to provide real-world legal documents.
1,384 eligible 1,255 verified & consenting
185
selected
Eligible Verified & consenting Selected
Each bar scaled to its group · 2,406 eligible · 1,845 verified & consenting · 692 selected
Document flow

QC Pipeline

The flow of intake, evaulation, rejection and re-evaluation of documents. 

4,071
documents collected
100%
STAGE 1
LLM REVIEW
Intake Gate

Confirms the document meets a strict set of criteria - valid file, substantial content, readable, complete, and the right document type - and verifies consent, NDA and ownership.

The basic criteria.
Valid, uncorrupted file
Readable, extractable text
Complete, with all pages, schedules and annexures
Correct document type for its category
Right to use, on record.

Every contributor is identity verified, confirms ownership of the document, and signs consent and an NDA before anything is submitted. Both are validated here at intake, and the record travels with the document.

Identity verifiedOwnership confirmedConsent signedNDA installed
−1,415removed  by intake gate
2,656
checked
65%
STAGE 2
LLM REVIEW
PII Redaction

Verification by person, consent captured, NDA in place, ownership and rights confirmed. Identifying values replaced with synthetic ones. Every substitution kept in a change log.

PII replaced, not blacked out.

Names, addresses, amounts and identifiers are swapped for synthetic values or labels so the document still reads naturally. Every substitution is kept in a change log, delivered with the data. Documents that lose too much value once redacted are dropped.

Synthetic substitutionChange log deliveredMeaning preserved
−553removed, redaction reduces value too significantly
2,103
scored
62%
STAGE 3
LLM REVIEW
LLM Scoring

Scored in depth against the universal quality rubric and a comprehensive rubric specific to its document type, with every criterion recorded.

Two rubrics per document.
Universal
quality rubric, applied to every document
31
document-type rubrics, matched to what the document actually is
85%
or better on both is the gold standard bar
Shaped by human review.
SOURCE

550 initial documents we hold rights to, both human and LLM authored.

AUTHORSHIP

LLM-authored documents were produced by verified lawyers from our network, then human reviewed before lock-down.

GROUND TRUTHING

Ground truthed by top-tier human experts across 15% of full dataset to improve the rubric accuracy and weighting.

−859scored sub 85% on either rubric
160RETURNED · TO STAGE 2
1,244
checked
Internal Human QC
GATE · HUMAN

Document quality check by Askable Labs team

−235removed at internal QC
54RETURNED · TO STAGE 2
1,009
reviewed
22%
STAGE 4
HUMAN REVIEW
Expert Review

A lawyer reads against the rubric and records reasoning. Every review is written back into the rubric, so the bar sharpens as the run progresses. This step is what makes the set human-verified rather than model-graded.

Judged by practicing lawyers.

Every gold standard call was made by a person. The lawyers who adjudicated were selected and evaluated in four stages.

1
Credential check

Practicing certificate and jurisdiction verified against the bar register.

2
Interview

Onboarded by phone with one of our researchers. Consent and NDA signed.

3
Scored trial

A trial set of documents reviewed against the rubric, checked for reasoning quality.

4
Ongoing calibration

Reviews spot-checked against peers throughout the run. Drift means re-calibration.

−360rejected by a lawyer
47RETURNED · TO STAGE 2
649
crosschecked
18%
STAGE 5
HUMAN REVIEW
Expert Crosscheck

A different (to first review) expert views and validates all scoring.

What the crosscheck covers.
Label accuracyScore consistencyResidual PII sweepMetadata completenessTraining-readiness sign-off
−116downgraded in final expert QA
533
final documents
FINAL DOCUMENT SET
Locked and delivered.

Every surviving document ships with its rubric scores, reviewer reasoning, consent record and redaction log.

DOC & SCORE

Insights

Score distribution.
Unredacted 3,796 Final 2,099
Share of scored documents per bucket
FAILBRONZESILVERGOLD 02550% 0–24 12.7% 0% 0.1% 0.1% 0.3% 0.4%50–54 0.6% 1.4%60–64 2.6% 4.6%70–74 8.3% 12%80–84 16.4% 21.3%90–94 19.5%95–100 12.5% 0–24 51.9% 0% 0% 0% 0.2% 0.4%50–54 0.4% 1.2%60–64 2.5% 4.9%70–74 10.8% 15.9%80–84 19.7% 17.1%90–94 16.2%95–100 10.8% 0–2450–5460–6470–7480–8490–9495–100
Rejected at intake.
Not a legal document 56%
Personal information still visible 25%
Something else 14%
Unreadable or garbled 4%
Wrong document type 2%
Document mix
Type-by-type volume across the collected set.
DOCUMENT TYPE VOLUME CATEGORY
Service agreement (MSA, SOW or SLA) High
Vendor, supplier or subcontractor agreement High
Commercial memos and documents High
Board minutes, corporate resolutions or share purchase High
Commercial NDA or confidentiality agreement High
Employment contract or executive agreement High
Client, consulting or distribution contract High
Vendor or contractor agreement, personal capacity Medium
Commercial property lease or equipment financing Medium
Corporate governance document Medium
Court filing or dispute record Medium
NDA, personal capacity Medium
Vehicle purchase or financing agreement Medium
Legal advice memorandum Medium
No type layer assigned (universal rubric only) Medium
Residential lease or rental agreement Medium
Property deed Medium
Last will and testament or codicil Low
Amendment, variation, rider or termination agreement Low
Intellectual property assignment or licence Low
Settlement agreement or deed of release Low
Data processing agreement Low
Order form or ordering document Low
Side letter, consent or waiver Low
Commercial Personal
C2

World Environments

Fictional legal worlds written from scratch by lawyers. Each world is internally consistent, so prompts can test reasoning across documents with no real-world leakage. Delivered to a frontier AI lab.

Figures
500
total documents
3
independent worlds
78
ready-to-run prompts
The worlds

Three SaaS Worlds, Three Jurisdictions

Each world is a complete fictional company - parties, matters and documents that reference each other the way a real deal file does. Three jurisdictions, so the same task types run like for like across three legal systems.

Verdemont
WORLD ONE · DELAWARE US
Verdemont Systems, Inc.

A Delaware B2B SaaS company across seven years, from scrappy founder-era contracts to polished post-funding ones: MSAs, order forms, DPAs, employment, IP, side letters and a full corporate record.

167 documents
14 doc types
26 prompts
DOCUMENT BREAKDOWN
{{ t.name }}{{ t.count }}
EVALUATION PROMPTS · 26 3 basic 14 intermediate 9 advanced
Arithmetic & reconciliation9
Corpus tracing & consistency9
Contract architecture & scope4
Governance & authority3
Extraction1
Currawong
WORLD TWO · BRISBANE AU
Currawong Software Pty Ltd

A workforce-scheduling SaaS across eight years, with a disputes family and a New Zealand subsidiary. Corporations Act formalities, Fair Work settlement and GST arithmetic throughout.

172 documents
15 doc types
26 prompts
DOCUMENT BREAKDOWN
{{ t.name }}{{ t.count }}
EVALUATION PROMPTS · 26 4 basic 16 intermediate 6 advanced
Arithmetic & reconciliation8
Corpus tracing & consistency7
Governance & authority5
Contract architecture & scope5
Extraction1
Brambling
WORLD THREE · MANCHESTER UK
Brambling Technologies Ltd

A building-safety SaaS built around the Building Safety Act 2022, with a full pre-litigation data dispute, an employment tribunal claim and an Irish subsidiary adding EU GDPR documents. Board minutes, remediation contracts and regulator correspondence sit alongside.

161 documents
15 doc types
26 prompts
DOCUMENT BREAKDOWN
{{ t.name }}{{ t.count }}
EVALUATION PROMPTS · 26 2 basic 16 intermediate 8 advanced
Arithmetic & reconciliation10
Corpus tracing & consistency6
Governance & authority5
Contract architecture & scope4
Extraction1
How it was built

Built by lawyers, checked by lawyers.

Each world was designed by an experienced lawyer, drafted document by document, and checked before it shipped.

1
Environment spec

A lawyer designs the environment (entities, counterparties, timeline) with one specification row per document.

2
Drafting

A lawyer prompts each document from its spec row, with the right cross-references built in.

3
Automated checks

Five checks run on every document, and re-run whenever anything changes. Detail below.

4
Legal review

A lawyer reviewed a 10% sample of documents from each of the three worlds.

The five checks
What the automated gate tests on every document.
1
Every number proven before it was written

Every derived figure computed and verified in code before it entered a document: settlement sums, share subscriptions, vesting schedules, payment deadlines. Figures shared between documents checked to match across them.

2
Every cross-reference resolved

Citations verified against the shipped version of the cited document: exact date, exact title, and that it actually says what the citing document relies on.

3
Required facts confirmed, prohibited facts excluded

Each document’s required facts probed against the rendered text, plus the negative checks where legal realism lives: no privileged advice cited, no case number before filing, no reference to its own future.

4
Every signature validated

Signers checked against a central registry, role start dates and the corporate timeline. Execution formalities match the instrument and jurisdiction for that date.

5
Defects fixed, full gate re-run

Never waived. The process caught real defects in shipped documents, like misdated consents and silently unrendered signature blocks, found and corrected by whole-set rescans.

C3

Multi Modal

Paired working sessions captured across screen and voice: in every session an expert and a naive user work through a real task together, drawn from five everyday themes. Each session ships as one folder: the combined recording, isolated per-participant audio tracks with timing metadata, and a time-coded transcript. Delivered to a frontier AI lab.

Summary
1,000
recorded sessions
5 × 10
themes × topics per theme
2,000
isolated audio tracks
Participants

Recruited for the task, not the camera.

Participants were recruited from the Askable panel for domain fit, then screened for hardware and environment: quiet room, stable connection, screen capture at working resolution.

Experts
Two per topic, each leading ten sessions.
1,860 eligible 645 verified & consenting
100
selected
Naive users
Everyday users paired with each expert.
7,260 eligible 2,834 verified & consenting
1,000
selected
Informed consent + NDA on fileHardware & environment checkOne expert + one naive user per session · 2 experts per topic
Pipeline

From panel to delivered tracks.

Every session moves through the same seven stages, from panel recruitment to secure transfer.

1
STAGE 1
Recruit & screen users from panel

Naive users recruited from the Askable panel and screened for fit, device and environment.

2
STAGE 2
Recruit & screen experts from panel

Domain experts recruited per theme and screened on background and session experience.

3
STAGE 3
Validate consent, NDA and video interview

Informed consent and NDA validated, and every participant passes a live video interview.

4
STAGE 4
Matching, scheduling and introduction

Each user is matched with an expert in their topic, scheduled, and introduced before the session.

5
STAGE 5
Session capture and completion

The pair works through the task while the session records the shared screen view and each participant’s voice as an isolated track.

6
STAGE 6
Human QC and verification

Every session reviewed by a human: track sync, audio quality, screen legibility and task completion verified.

7
STAGE 7
Data processing and transfer

Tracks named consistently, hashed and packaged, then delivered by secure transfer.

Themes

Five themes, fifty topics.

Every session sits under one of five everyday themes. Each theme breaks into ten topics, and every topic ran as twenty sessions: twenty naive users, two experts.

{{ tp.num }} {{ tp.title }} 20 users · 2 experts · 20 sessions
200 sessions per theme · 1,000 total
Combined session video · MP4Isolated audio · OGG / 48 kHzTime-coded transcript · TXT
Deliverables

One folder per session.

Every session lands in the same fixed structure, so a delivery can be validated and ingested by machine. The .complete marker is written only after every file has arrived intact.

Per-session structure 1,000 folders delivered
<syncId>/
└── <studyId>/
└── <videoCode>/ One folder per session, keyed by sync, study and video IDs.
├── <videoCode>.mp4 Combined session recording of the shared view.
├── tracks/ Isolated per-participant tracks plus timing metadata.
│ ├── <track_id>.json Per-track metadata: participant, role and duration.
│ ├── track_timing.json Offsets aligning each track to the session clock.
│ ├── track_timing.debug.json Raw sync diagnostics, kept for audit.
│ └── EG_*.ogg Isolated audio per participant · Opus 48 kHz.
├── transcript.txt Time-coded transcript of the full session.
└── .complete Written last; marks the folder as final.
C4

Conversational Audio

Natural, unscripted speech with human-verified reference transcripts. Conversational sets are full-duplex: every participant on their own simultaneous, isolated single-speaker track, so overlap and turn-taking are preserved. Delivered to a frontier AI lab.

Summary
250
recorded sessions
~60 hrs
unique audio
16,840
verified segments
Participants

Real speakers, six languages.

Speakers were recruited from the Askable panel across languages and accents. Every participant gave informed consent for recording and model-training use, and signed an NDA. Per-speaker demographics ship in each session’s participants.json.

Solo speakers
One participant per session, unscripted.
3,850 eligible 1,310 verified & consenting
120
selected
Conversation groups
Two- to four-party natural conversations.
1,240 eligible 405 verified & consenting
205
selected
368 speaker tracksEnglish · French · Spanish · Japanese · Hindi · Arabic · MandarinConsent covers training use
Pipeline

Every segment touched by a human.

Transcripts start as machine drafts and end human-verified, with the full correction trail kept as the QC record.

1
Recruit & consent

Panel recruitment across languages and accents; informed consent and NDA before any recording.

2
Full-duplex capture

Every speaker on their own simultaneous, isolated track. No cross-talk, no diarization needed.

3
ASR first pass

Draft transcripts generated per track (Qwen3-ASR), segmented for review.

4
Verify & package

Every segment confirmed or corrected by a reviewer against the isolated audio; changelog, manifest and SHA256 integrity ship with each set.

Sample sets

Four capture formats, one standard.

Every segment was confirmed or corrected by a human reviewer against the isolated audio, with the full correction changelog kept as the QC trail.

SOLO
Solo sessions

Unscripted solo sessions, one participant each, across six languages.

160 sessions6 languages~30 hrs5,360 segments
DUAL
Two-party conversations

Arabic-English and Mandarin-English code-switching, plus natural-environment English.

40 conversations~12 hrs4,610 segments
MULTIPARTY
Group conversations

Three- and four-party conversations: debates, advisory sessions and group planning.

20 conversations68 tracks~10 hrs3,480 segments
MULTILINGUAL
Cross-language pairs

Two-person cross-language conversations in English, Japanese and Spanish.

30 conversations~9 hrs3,390 segments
Human-verified transcriptsIsolated speaker tracksOpus / 48 kHzManifest + SHA256 integrity
C5

Terminal-Bench

Real software tasks packaged as reproducible terminal environments for agentic AI training and evaluation, each with a verified human solution trajectory. Delivered to a second-tier AI lab.

Summary
25
task packages
25
verified trajectories
100%
reproducible environments
Participants

Working developers, real codebases.

Task authors and reviewers were recruited from the Askable panel and screened on verified engineering history. Authors build and solve the task; an independent reviewer reproduces it from the package alone.

Developers
Task authors with verified engineering history.
1,240 eligible 386 verified & consenting
32
selected
Informed consent + NDA on fileTechnical screen before selectionIndependent reviewer per task
Pipeline

From real defect to reproducible package.

Each task starts as a genuine piece of engineering work and ends as a self-contained environment anyone can run.

1
Task sourcing

Real defects and features drawn from working codebases, scoped to a clean terminal task.

2
Environment build

Codebase, task brief and dependencies packaged into a self-contained, reproducible environment.

3
Trajectory capture

A developer solves the task in the packaged environment; the full terminal session is recorded.

4
Checks & review

Automated tests define success; an independent reviewer reproduces the run before the task ships.

Tasks

Self-contained task packages.

Twenty-five packages, each holding the codebase, task brief, solution and checks needed to run the task in a clean terminal environment.

01 async-report-loader-cache-fix Bug fix
02 react-contact-book-url-state Feature
03 csv-import-encoding-errors Bug fix
04 rate-limiter-token-bucket Feature
05 flask-session-timeout-fix Bug fix
06 sqlite-migration-runner Feature
07 webpack-build-memory-leak Bug fix
08 cli-todo-app-filtering Feature
09 express-auth-middleware-refactor Refactor
10 pandas-groupby-performance Performance
11 git-history-cleanup-script Scripting
12 docker-compose-healthchecks Config
13 json-schema-validation-errors Bug fix
14 image-thumbnail-worker-queue Feature
15 timezone-handling-in-scheduler Bug fix
16 graphql-n-plus-one-queries Performance
17 log-rotation-and-retention Config
18 markdown-parser-edge-cases Bug fix
19 payment-webhook-retry-logic Feature
20 css-grid-layout-regression Bug fix
21 python-package-dependency-pin Config
22 websocket-reconnect-backoff Feature
23 unit-test-flakiness-fix Testing
24 api-pagination-cursor-support Feature
25 nginx-reverse-proxy-cache Config
task.yaml briefDockerfile environmentScripted checksRecorded solution trajectory
The delivery

{{ deliveryTitle }}

Open the data →
{{ deliveryRootName }} {{ deliveryStats }}
{{ row.name }} {{ row.name }} {{ row.meta }}
The rubrics