Evidence-first, published openly as it develops
Our modeling work is grounded in the founder’s peer-reviewed Stacking Generative AI research and evaluated against modern medical-AI standards — not single-point accuracy claims.
Validation dimensions
VCIE™ assesses every model across multiple complementary dimensions, consistent with FDA SaMD and academic best practices for predictive clinical AI.
| Discrimination | ROC AUC and precision-recall AUC |
| Calibration | Calibration slope, intercept, and calibration curves |
| Probability quality | Brier score |
| Classification | Sensitivity, specificity, precision, recall, and F1 at use-case-specific thresholds |
| Generalizability | Held-out, external, temporal, and multi-site evaluation |
| Subgroup performance | Performance across age, sex, race/ethnicity where available, geography, and disease burden |
| Clinical utility | Decision-curve and threshold-based operational analysis |
| Reliability | Missingness, out-of-range inputs, data drift, and uncertainty handling |
Retrospective performance is evidence of technical potential, not proof of clinical effectiveness. We distinguish research results, internal validation, independent validation, prospective performance, and regulatory-cleared claims in every communication.
Where each module stands today
Cardiovascular
Our stacking model has shown up to 98% accuracy and a ROC AUC of 0.993 across nine diverse clinical datasets. On an independent Framingham test set (n=11,627): ROC AUC 0.97–0.98, near-ideal calibration (slope ≈ 0.98).
Kidney
CKD Screen reuses the same VCIE™ ingestion, explainability, and validation infrastructure as our cardiovascular module. Independent validation is in progress.
Metabolic
Type 2 diabetes risk prediction is in active development on the same reusable platform, designed for consistent evaluation against the dimensions above.