Published findings.
Downloadable research papers on outcome prediction, judicial analytics, settlement modeling, and the methodology behind the Splitifi intelligence layer.
Ten papers. One research program.
Calibrated Outcome Prediction in Legal Contexts
Deterministic prediction models trained on real court outcomes produce calibrated probability estimates that hold across jurisdictions. This paper documents the training methodology, discrimination validation standards, time-aware train/test splits, and leakage remediation applied across production models in the Splitifi registry. We argue that calibration — not accuracy alone — is the correct quality standard for legal prediction.
Judicial Behavioral Analytics: Ruling Patterns Across the Federal and State Bench
Analysis of ruling patterns across federal and state judges reveals statistically significant behavioral clustering: judges with similar educational backgrounds, appointment pathways, and case volume exhibit measurably similar outcome distributions. This paper presents the clustering methodology, validates its predictive utility, and documents the ethical framework governing judicial intelligence products.
Settlement Zone Modeling
Settlement zones — the range within which litigants actually resolve contested issues — are predictable from case characteristics, judicial assignment, and comparable outcomes. This paper introduces the settlement zone model, documents its construction from the Splitifi court record corpus, and presents validation results across custody, asset division, and support contexts.
Litigation Finance: Quantifying Legal Risk with Deterministic Models
Litigation funders require precise estimates of win probability, expected award range, and time to resolution to construct IRR models. This paper presents a framework for litigation finance underwriting using calibrated outcome models, documents the Iron Triangle methodology (win probability x award range x recovery probability divided by time), and presents backtested performance across 20 vertical categories.
The Legal Data Lake: Architecture for Court Record Infrastructure
Building a training-ready legal dataset from public court records requires jurisdiction-specific ingestion logic, outcome normalization, feature engineering, and temporal indexing. This paper sets out the data problem, the source coverage, and the quality controls that govern the corpus underlying all production models.
Model Validation Standards for Legal Prediction
Legal prediction models require validation standards that go beyond standard ML benchmarks. This paper argues for actuarial-grade validation: time-aware train/test splits (no temporal leakage), calibration curves (not just discrimination), and out-of-jurisdiction holdouts. We present the validation protocol applied to all models in the Splitifi production registry and document the standards that disqualify models from deployment.
Family Law as Data Science: The Proof-of-Concept Vertical
Family law was chosen as the foundational vertical not despite its complexity but because of it. Jurisdiction variability, judicial discretion, and multi-outcome case structure make family law the hardest prediction domain in law. This paper documents the methodology developed over five years of family law modeling and argues that its success validates the approach for all legal verticals.
Why Deterministic Models Matter for Legal AI
Large language models cannot produce calibrated legal predictions — not because of training data limitations but because inference-based reasoning does not produce probability estimates that hold at scale. This paper distinguishes between LLM-based legal tools and deterministic outcome models, documents the theoretical basis for this distinction, and presents empirical evidence from production deployment.
Multi-Jurisdictional Outcome Patterns in Family Law
The same case, filed in different jurisdictions, produces measurably different outcomes — and those differences are predictable. This paper presents outcome distribution analysis across all 50 US states, four Canadian provinces, and England and Wales, identifying the key jurisdictional variables that drive outcome divergence and validating the jurisdiction-specific model approach.
Custody Prediction Accuracy: A Retrospective Validation Study
We evaluate the accuracy of custody prediction models by comparing predictions to actual court orders across a held-out set of final custody decisions. Production models are validated against held-out final custody orders across jurisdictions. This paper documents the validation methodology, identifies the primary accuracy-limiting factors, and presents responsible deployment guidelines.
Each paper is free to download.
Select any paper above to read the full text and download the PDF. Institutional partnership inquiries, data access requests, and structured research arrangements go through our contact form.
