Validation & technical library

Methods are only useful when their limits travel with them.

Current engine methods, score contracts, sample outputs, verification evidence, and known limits. Each statement is restricted to the named source version and deployment architecture.

Theta AI

Pipeline V5.3 · score contract v1

Unidimensional Item Response Theory using a Graded Response Model fitted in R with mirt, then Expected A Posteriori scoring.

Calibration asset

Frozen synthetic item bank synth-bank-v1 · 20 items · synthetic calibration cohort n=500 · seed cal-20260921-b. Input/output SHA-256 values and commit 49380c1e are recorded in the provenance asset.

Score contract

EAP theta, standard error, 95% interval, and ability_percentile = 100 × Φ(theta) against a frozen N(0,1) reference.

Ability level

Low below -1, Moderate from -1 through 1 inclusive, and High above 1. This is a descriptive relative level, not a pass/fail cut score.

Separate constructs

Panels are never weighted, blended, or summed. Turnover risk is not measured in the current canonical version, and no proxy is substituted.

Integrity evidence

Detector features produce reviewer reference signals, not proof of misconduct. The flag guide records possible false-positive causes and follow-up questions.

Synthetic calibration

synth-bank-v1 synthetic calibration (n=500) reported marginal reliability 0.913. Fitted asset SHA begins 19cdb985; calibration run 35553441453.

Synthetic recovery

synth-detectors-v1 untouched synthetic validation reported theta recovery r=0.850 (95% CI 0.794–0.900; 2,000 bootstrap samples), final run 35578002689. This is not job-performance predictive validity.

Synthetic integrity classifier

synth-detectors-v1 untouched synthetic validation: precision 0.900 (95% CI 0.794–0.977), recall 0.900 (95% CI 0.795–0.979), specificity 0.991 (95% CI 0.982–0.998); dev run 35555214723.

Frozen detector configuration

04_model_training_ensemble.R v7.2.2 and 06_model_evaluation.R v8.3.31 use XGBoost 0.35, LightGBM 0.35, Random Forest 0.30, deep learning 0.00. Frozen target-label cutoff 0.495350; classification threshold 0.5.

Known limits

  • Synthetic calibration is not real-world validity evidence.
  • Oracle generating parameters are used only for diagnostic comparison, never for fitted scoring.
  • This asset applies only to synthetic fixed form synth-bank-v1.
  • Promotion to frozen status requires a separate honest calibration cohort and successful artifact/provenance checks.
Open fictional sample report

Titan AI

Live adaptive architecture

A Graded Response Model with Bayesian adaptive EAP scoring and independent per-candidate theta. The engine selects the next scenario adaptively and can produce a synchronous report after the final response without waiting for cohort calibration.

Service architecture

FastAPI service titan-app on Google Cloud Run. Service requests use X-Service-Api-Key and X-Tenant-Id headers.

Competency scale

The report shows competency mastery on a 0–5 scale against a 4.0 proficiency reference.

Evidence design

Per-competency narratives, strengths, risks, and interview questions keep the response evidence available to the reviewer.

Limited evidence

Observation-count logic marks thin evidence as limited-evidence. It is shown rather than hidden.

Integrity evidence

Behavioral indicators are reviewer reference signals, not proof of misconduct.

System of record

Titan’s Neon database is the system of record; the portal queries it live for current results.

Known limits

  • No load testing has been performed.
  • A 4.0 proficiency reference is report context, not an automated employment decision.
  • Thin evidence can limit interpretation even when a score is available.
  • The report is a decision aid for a human reviewer, not an automated hiring decision.
Open fictional sample report

Shared methodology

Guardrails that stay attached to the evidence.

01

Human review

Report faces state that human review is mandatory.

02

Integrity interpretation

Flags are reference signals, not accusations. Review guidance includes false-positive causes and re-verification questions.

03

Two-stage gate

The Theta → Titan flow includes an HR approval gate between engines.

04

Reproducible evidence

Theta CI verifies report rendering, rasterization, and pixel margins for both PDF and HTML outputs.

05

Version discipline

Theta’s pipeline version, score contract, reference, detector, and provenance are recorded rather than silently replaced.

06

Claim boundary

The current technical assets do not establish universal validity, fairness, customer outcomes, or guaranteed job performance.

Source register

Named artifacts

Theta: pipeline/scripts/02_irt_modeling.R, 02a_cheating_feature_engineering.R, pipeline/calibration/synth-bank-v1.json, canonical score contract v1 with ability-level addendum, detector provenance CSV, flag guide, and report-verification workflow. Calibration run 35553441453; detector dev run 35555214723; untouched synthetic final validation run 35578002689.

Titan: lib/engines/titan-client.ts, titan-orchestrator.ts, components/titan-report-panel.tsx, and titan-report-visuals.tsx. Shared review rules: flag guide, report templates, and specification chapter 12.