DATA LICENSING

The dataset your model can't scrape.

Ten years of real investigative operations — human decisions, machine assists, corrections, and outcomes — captured as it happened, structured for training. Not synthetic, not annotated after the fact. Doctrine-grounded, provenance-tracked, and licensable as datasets or living streams.

CORPUS AT A GLANCE
10 yrs
Continuous operational capture
NAICS + SIC
Industry-mapped taxonomies
All-INT
Fused intelligence disciplines
H+M
Human–machine teaming, end to end
Volume metrics, schemas, and samples are shared under NDA.
WHAT YOU LICENSE

Four asset classes, one provenance chain.

Investigative reasoning traces

Decade-spanning decision chains from live casework: hypotheses formed, evidence weighed, dead ends ruled out, conclusions reached. The scarce commodity in reasoning data — including what didn't work, and why.

Human–machine teaming telemetry

Workflow-level records of operators and models working the same cases: where models accelerate, where humans override, how the division of labor evolved across a decade of model generations.

Cross-model comparison suites

Extensive evaluations of models run against identical live workflows — insight into capability deltas, failure modes, and improvement curves that benchmark suites can't reproduce.

Doctrine & industry taxonomies

Investigative doctrine encoded as machine-usable structure, with NAICS- and SIC-grounded labeling across security, analytical, fintech, healthtech, defense, and maritime domains — the granularity that lets a model know which world it is reasoning about.

WHY THIS DATA

Models are running out of internet. The next capability gains come from grounded operational data.

Long-horizon judgement — pursuing a hypothesis across weeks, integrating contradictory evidence, deciding when to act and when to wait — barely exists in public corpora. It exists here, because we captured our own operations from day one. For labs training agents, harnesses, and models that must hold state over long horizons and act reliably in the world as it actually is, this is the substrate.

01
Frontier & open-weight labs — reasoning and agentic post-training data with provenance you can defend to regulators.
02
Vertical builders — security, analytical, fintech, healthtech, and defense teams training domain models and harnesses.
03
Long-horizon systems — teams whose models must plan, decide, and act over extended timelines with industry-grade granularity.
LICENSING STRUCTURES

Datasets, streams, or both.

Dataset licenses

Versioned, sanitized corpus slices — scoped by asset class, taxonomy, or vertical. Delivered with schemas, lineage documentation, and evaluation baselines.

Point-in-time · Versioned · Non-exclusive or exclusive by field
Living streams
FLYWHEEL

Continuous delivery of newly generated teaming and reasoning data as our operations run — your models improve on the same flywheel ours do.

Subscription · Fresh weekly · Schema-stable
Evaluation & research access

Scoped access for benchmarking and academic work — run evaluations against the corpus on our infrastructure, publish with our review.

Platform-hosted · Data never leaves · See Research

Every licensed asset carries its sanitization lineage and provenance record. GRC review is part of delivery, not an afterthought — see Technology for how compliance is enforced by construction.

Schemas, samples, and volume metrics — under NDA.
Tell us what you're training. We'll tell you what we have.
Start the conversation