Independent Holdout Validation

Validation Methodology: N-1 Temporal Holdout

The Stability Engine retention risk scoring system has been independently validated using N-1 temporal holdout methodology across a frozen cohort review. This page explains the method, the calibration evidence, and the honest limits of what the validation does and does not prove.

N-1
holdout method
72.5%
score calibration ±15pts
0.169
mean Cox Brier score
n=51
holdout cohort

Validation methodology: N-1 temporal holdout

The validation uses N-1 temporal holdout — a methodology designed to test whether a scoring system generates useful predictions from only the data available at the pre-hire stage.

How it works:

  1. 01

    Withhold the most recent completed role

    For each candidate in the validation cohort, the most recent completed role — with a known start date, end date, and tenure length — is removed from the career file. This becomes the ground truth.

  2. 02

    Score on prior history only

    Stability Engine runs on the remaining career history — what was visible before that last role started. This mirrors the actual pre-hire information state: the system scores only what would have been available at the moment of the offer.

  3. 03

    Predict 12-month retention

    The system generates a Stability Score and a 12-month retention probability for each candidate, based solely on prior career history.

  4. 04

    Compare prediction to ground truth

    The predicted 12-month retention outcome is compared to the actual tenure of the withheld role. A candidate who scored in the higher risk bands and departed within 12 months is a correct prediction. A candidate who scored in the lower risk bands and remained past 12 months is also a correct prediction.

Validation results

MetricResultWhat it measures
Validation stanceMethodology-firstPublic proof emphasizes the holdout method and logged prediction discipline, not a single headline threshold metric
Score calibration72.5% within ±15Stability Scores within ±15 points of the reference label (calibration, not threshold accuracy)
Mean Cox Brier score0.169Probabilistic calibration quality (0 = perfect, lower is better)
Cohort sizen=51Total candidates in the holdout validation cohort
MethodologyN-1 temporal holdoutPrior career history only; most recent role withheld as ground truth

What the Stability Score measures

The Stability Score analyzes structural career history signals — not interview performance, personality assessments, or self-reported preferences. The relevant signals include:

  • Prior tenure patterns: how long the candidate stayed across completed roles, and the distribution of that tenure
  • Transition density: how quickly the candidate has moved between roles and environments
  • History alignment: whether the prior career pattern matches the stability demands of the role being assessed
  • Environmental fit signals: whether prior operating environments resemble the current one

The score is a directional signal, not a verdict. It does not tell a hiring team to hire or not hire a candidate. It provides a structured basis for calibrating onboarding investment, monitoring cadence, and early intervention — not for replacing the human judgment that belongs in any serious hiring process.

Honesty architecture

  • Predictions logged before outcomes. Prediction snapshots persist at scan time; the system cannot grade itself retroactively.
  • Frozen regression corpus (engineering discipline). An 18/18 frozen corpus locks current scoring semantics in CI so the pipeline cannot silently drift. This is release discipline, not an accuracy claim.
  • Proof chain. Per-candidate: risk flag → intervention → employment outcome → learning eligibility, with honest complete / in-progress / insufficient states.

Forward validation protocol

Holdout calibration is the retrospective test. The forward proof comes from monitored hires whose predictions were logged before outcomes:

  • A read-only day-30 retrospective calibration runner exists and is fixture-verified. It measures whether monitored-state risk shows directional signal against realized tenure.
  • Current live outcome cohort = 0. There is not yet a real departed-outcome cohort to evaluate — we say this plainly. The number grows as monitored hires accrue logged outcomes.
  • Closed-loop learning is built and activates with volume. Frame it as: the system is architected to learn from your logged outcomes — never as already learned from a large outcome corpus.

Honest limits of the validation

The validation establishes predictive signal, not certainty. Several important limitations:

  • The holdout cohort is n=51. This is a meaningful validation data point but not a large-scale epidemiological study. Additional validation is ongoing as outcome data accumulates.
  • The score captures career history pattern risk — not environmental factors, management quality, or post-hire conditions that also affect retention.
  • A high Stability Score does not guarantee retention. A lower score does not mean a hire will fail. Scores are probability distributions, not individual predictions.
  • The N-1 methodology tests the system against prior career history only. It does not test prediction performance in real-time, concurrent hiring conditions.
  • Disclosed methodology ceiling: where visible prior history is genuinely ambiguous, the system can be too pessimistic (pattern-break false-negatives). These cases are tracked, not hidden.

Frequently asked questions

What is the N-1 temporal holdout methodology?

The most recent completed role in a candidate's career history is withheld as the ground truth. The scoring system runs on prior career history only — what was visible before that last role started. The prediction is then compared to what actually happened in the withheld role.

Why does Stability Engine avoid a single headline accuracy number?

Because threshold accuracy can be inflated by base rates and make a retention model look more discriminating than it is. The safer public framing is the validation method itself, the score calibration evidence, and the fact that Ros logs predictions before outcomes so customers can audit the system on their own cohort.

What is a Brier score and what does 0.169 indicate?

The Brier score measures probabilistic prediction accuracy on a 0-to-1 scale. 0 is perfect; higher is worse. A score of 0.169 indicates well-calibrated probabilistic forecasts — the stated probabilities of early departure track closely with observed departure rates.

Where can I read the full validation study?

The public validation PDF is currently under revision because an older version contained retired headline claims. This methodology page is the current public source for the holdout process and its interpretation limits.

Revision Note

Validation study PDF under revision

The previously downloadable 2026 validation PDF was removed from public distribution because it still contained retired headline claims. It will be republished only after a full revision against the current messaging guardrails.