Research

Evidence-first, published openly as it develops

Our modeling work is grounded in the founder’s peer-reviewed Stacking Generative AI research and evaluated against modern medical-AI standards — not single-point accuracy claims.

How we evaluate

Validation dimensions

VCIE™ assesses every model across multiple complementary dimensions, consistent with FDA SaMD and academic best practices for predictive clinical AI.

DiscriminationROC AUC and precision-recall AUC
CalibrationCalibration slope, intercept, and calibration curves
Probability qualityBrier score
ClassificationSensitivity, specificity, precision, recall, and F1 at use-case-specific thresholds
GeneralizabilityHeld-out, external, temporal, and multi-site evaluation
Subgroup performancePerformance across age, sex, race/ethnicity where available, geography, and disease burden
Clinical utilityDecision-curve and threshold-based operational analysis
ReliabilityMissingness, out-of-range inputs, data drift, and uncertainty handling

Retrospective performance is evidence of technical potential, not proof of clinical effectiveness. We distinguish research results, internal validation, independent validation, prospective performance, and regulatory-cleared claims in every communication.

By condition

Where each module stands today

C

Cardiovascular

Our stacking model has shown up to 98% accuracy and a ROC AUC of 0.993 across nine diverse clinical datasets. On an independent Framingham test set (n=11,627): ROC AUC 0.97–0.98, near-ideal calibration (slope ≈ 0.98).

K

Kidney

CKD Screen reuses the same VCIE™ ingestion, explainability, and validation infrastructure as our cardiovascular module. Independent validation is in progress.

M

Metabolic

Type 2 diabetes risk prediction is in active development on the same reusable platform, designed for consistent evaluation against the dimensions above.