{{ permsBody }} If you think you should have access, ask your project lead to add you.
Real legal documents collected from consenting professionals, quality-gated, PII-redacted and reviewed, with expert-written prompts and rubrics on top. Delivered to a frontier AI lab.
Multiple types to demonstrate our capability
Screened & verified with consent, ownership and NDA in place
The flow of intake, evaulation, rejection and re-evaluation of documents.
Every surviving document ships with its rubric scores, reviewer reasoning, consent record and redaction log.
Fictional legal worlds written from scratch by lawyers. Each world is internally consistent, so prompts can test reasoning across documents with no real-world leakage. Delivered to a frontier AI lab.
Each world is a complete fictional company - parties, matters and documents that reference each other the way a real deal file does. Three jurisdictions, so the same task types run like for like across three legal systems.
A Delaware B2B SaaS company across seven years, from scrappy founder-era contracts to polished post-funding ones: MSAs, order forms, DPAs, employment, IP, side letters and a full corporate record.
A workforce-scheduling SaaS across eight years, with a disputes family and a New Zealand subsidiary. Corporations Act formalities, Fair Work settlement and GST arithmetic throughout.
A building-safety SaaS built around the Building Safety Act 2022, with a full pre-litigation data dispute, an employment tribunal claim and an Irish subsidiary adding EU GDPR documents. Board minutes, remediation contracts and regulator correspondence sit alongside.
Each world was designed by an experienced lawyer, drafted document by document, and checked before it shipped.
A lawyer designs the environment (entities, counterparties, timeline) with one specification row per document.
A lawyer prompts each document from its spec row, with the right cross-references built in.
Five checks run on every document, and re-run whenever anything changes. Detail below.
A lawyer reviewed a 10% sample of documents from each of the three worlds.
Paired working sessions captured across screen and voice: in every session an expert and a naive user work through a real task together, drawn from five everyday themes. Each session ships as one folder: the combined recording, isolated per-participant audio tracks with timing metadata, and a time-coded transcript. Delivered to a frontier AI lab.
Participants were recruited from the Askable panel for domain fit, then screened for hardware and environment: quiet room, stable connection, screen capture at working resolution.
Every session moves through the same seven stages, from panel recruitment to secure transfer.
Naive users recruited from the Askable panel and screened for fit, device and environment.
Domain experts recruited per theme and screened on background and session experience.
Informed consent and NDA validated, and every participant passes a live video interview.
Each user is matched with an expert in their topic, scheduled, and introduced before the session.
The pair works through the task while the session records the shared screen view and each participant’s voice as an isolated track.
Every session reviewed by a human: track sync, audio quality, screen legibility and task completion verified.
Tracks named consistently, hashed and packaged, then delivered by secure transfer.
Every session sits under one of five everyday themes. Each theme breaks into ten topics, and every topic ran as twenty sessions: twenty naive users, two experts.
Every session lands in the same fixed structure, so a delivery can be validated and ingested by machine. The .complete marker is written only after every file has arrived intact.
Natural, unscripted speech with human-verified reference transcripts. Conversational sets are full-duplex: every participant on their own simultaneous, isolated single-speaker track, so overlap and turn-taking are preserved. Delivered to a frontier AI lab.
Speakers were recruited from the Askable panel across languages and accents. Every participant gave informed consent for recording and model-training use, and signed an NDA. Per-speaker demographics ship in each session’s participants.json.
Transcripts start as machine drafts and end human-verified, with the full correction trail kept as the QC record.
Panel recruitment across languages and accents; informed consent and NDA before any recording.
Every speaker on their own simultaneous, isolated track. No cross-talk, no diarization needed.
Draft transcripts generated per track (Qwen3-ASR), segmented for review.
Every segment confirmed or corrected by a reviewer against the isolated audio; changelog, manifest and SHA256 integrity ship with each set.
Every segment was confirmed or corrected by a human reviewer against the isolated audio, with the full correction changelog kept as the QC trail.
Unscripted solo sessions, one participant each, across six languages.
Arabic-English and Mandarin-English code-switching, plus natural-environment English.
Three- and four-party conversations: debates, advisory sessions and group planning.
Two-person cross-language conversations in English, Japanese and Spanish.
Real software tasks packaged as reproducible terminal environments for agentic AI training and evaluation, each with a verified human solution trajectory. Delivered to a second-tier AI lab.
Task authors and reviewers were recruited from the Askable panel and screened on verified engineering history. Authors build and solve the task; an independent reviewer reproduces it from the package alone.
Each task starts as a genuine piece of engineering work and ends as a self-contained environment anyone can run.
Real defects and features drawn from working codebases, scoped to a clean terminal task.
Codebase, task brief and dependencies packaged into a self-contained, reproducible environment.
A developer solves the task in the packaged environment; the full terminal session is recorded.
Automated tests define success; an independent reviewer reproduces the run before the task ships.
Twenty-five packages, each holding the codebase, task brief, solution and checks needed to run the task in a clean terminal environment.