How the models are built.
Documentation of data collection, feature engineering, model training, validation standards, and governance protocols applied across the Splitifi intelligence layer.
Real court records. No synthetic data.
All models in the Splitifi production registry are trained exclusively on real legal outcomes harvested from federal and state court systems. No synthetic data. No augmentation. Zero records fabricated. The corpus spans verified court records across federal, state, and international court systems, deduplicated and normalized to one outcome vocabulary before any model sees it.
Federal and state court dockets and public record systems. Four jurisdictions: US, Canada (BC/AB/ON/QC), Australia, UK (England & Wales + Scotland).
Each jurisdiction's outcome schema is normalized to canonical fields: outcome_type, award_amount, duration_days, judge_id, case_characteristics. Jurisdiction-specific fields are retained as auxiliary features.
Inputs are drawn from the structure of litigation itself: judicial behavior, jurisdiction, case posture and timing. The feature specification is documented per vertical and available to qualified institutions on request under NDA.
Calibrated prediction. Not inference.
Each model in the production registry is a versioned artifact with its training window, target, jurisdiction and case-type scope, validation status and change history recorded. The registry is the single source of truth for what is in production. Models do not use language model inference. They do not hallucinate. They do not drift with upstream API changes.
Every production model is cleared individually: a minimum training sample, a temporal holdout, a leakage audit and a calibration review. Model class, feature selection and thresholds are documented per vertical and available on request under NDA.
Registry entries are tracked through a status lifecycle — production, experimental, suspended, and failed. Promotion requires: discrimination threshold, calibration check, leakage audit. Demotion is automatic when quality gates fail.
Actuarial discipline applied to legal prediction.
Every model we ship has been reviewed as if it were actuarial output, not a demo.
Train/test splits respect temporal ordering. No future data leaks into training. Cases from the test period are never seen during training. Holdout periods are set to minimum 90 days.
We require calibration — not just accuracy. A model that says 70% should be right 70% of the time. We validate with calibration curves across decile bins. Models with significant calibration deviation are held at experimental status.
Every feature is audited for temporal leakage before inclusion. Features derived from case outcomes are excluded. Judicial assignment features use pre-case-filing judicial records only. Leakage is treated as a disqualifying defect.
For multi-jurisdiction models, a held-out jurisdiction is used as the test set. This validates that the model generalizes across jurisdictional variation rather than memorizing jurisdiction-specific patterns.
The production registry.
Every model in production has a registry entry documenting: training data hash, training date, discrimination score, calibration result, leakage audit status, promotion decision, and demoted status if applicable. The registry is the authoritative source of truth for all model state.
