MindSim BenchmarksContinuously updated
MindSim twins. Trust, earned.
Before you rely on a MindSim twin, you should know how closely it follows the person it represents. See the scores, test counts, protocol controls, and evidence boundary behind each claim.
Wilson 95% interval: 79.0% to 93.6%. Every reviewed response remains visible below.
Ongoing operational cross-check
Real use. Measured outcomes.
We compare MindSim twin predictions with observed human outcomes during ongoing use. Every response is routed through evidence-quality and attribution controls before it can affect operational accuracy. This 88.2% ongoing result stays separate from the 97.6% controlled benchmark below.
Evaluation coverage
All responses are tracked.
Passed the attribution gate.
These responses had enough real-world evidence for a defensible judgment. Sixty-seven matched and nine fell short, so accuracy is 67 divided by 76.
Rounding policy
The measured result is 88.2%.
Elsewhere on the site, this result is rounded to 90% or 9 in 10. Here, the underlying ratio, sample size, evaluation window, and 95% interval remain visible: 67 of 76, with a range of 79.0% to 93.6%.
Controlled benchmark
Controlled conditions. Exact test count.
Each evaluation asks the MindSim twin to predict an outcome it cannot see, then compares its answer with criteria fixed in advance or a verified real-world decision. The controlled result includes only tests that passed every protocol control.
Protocol scope
Protocol-qualified set
Includes only evaluations that passed every blinding and scoring control.
Held-out professional scenarios scored against criteria fixed before evaluation.
Held-out company profiles checked against verified prior investment decisions.
What 97.6% means
19 of 21 MindSim twins passed every qualified test.
The other two each missed one detail-heavy question. Across all 84 tests, that is 82 matches and two misses. Both misses followed the person’s strategic direction but did not satisfy every required tactical detail.
2 MindSim twinsmissed one required detail each
Methodology
Predict first. Score against ground truth.
A MindSim twin can sound convincing and still be wrong. Before evaluation begins, the protocol fixes the scenario, scoring rule, known outcome, and quality gates, then keeps the outcome hidden from the MindSim twin.
Lock the protocol
The scenario, scoring criteria, ground truth, and quality gates are defined before the MindSim twin is evaluated.
Create the MindSim twin
The Prediction Engine derives a person-specific model of values, reasoning, and decision patterns from protocol-qualified evidence.
Run the test blind
The MindSim twin receives a held-out scenario without access to the known outcome or the answer required to pass.
Score against truth
Predefined behavioral criteria or a verified real-world decision determine whether the result is recorded as a match, miss, or protocol holdout.
Study coverage
Qualified tests. Credible scores.
Of 47 investor candidates, 30 completed evaluation. Twenty-four passed every protocol control and six were held outside the controlled result by the blinding gate. Five were prepared for a later test window, and 12 did not meet the study threshold.
Protocol-qualified24
Blinding holdout6
Prepared, not yet tested5
Study threshold not met12
Data dictionary
How each term is used.
- Protocol-qualified
- An evaluation that passed every pre-scoring control, including the final blinding-quality gate. Six evaluations were held outside the controlled set.
- Predictive correctness
- A binary pass against pre-defined behavioral criteria or verified decision ground truth.
- Scored MindSim twin response
- An operational judgment where the observed outcome could be assigned to the MindSim twin rather than the surrounding integration.
- Not scoreable
- An observed response held outside the accuracy calculation because the available evidence does not support a defensible judgment.
Failure analysis
A miss is evidence. We trace every one.
You should know not only when a MindSim twin misses, but why. Every failed criterion is isolated, classified, and turned into a stricter held-out test. The miss stays visible.
A miss cannot disappear into the average.The 97.6% result remains paired with both failed evaluations and the exact criteria they did not satisfy.
- 2
- misses published
- 2
- rubric traces
- 2
- next controls defined
Published misses
Strategy aligned. Execution detail fell short.
One response matched the person’s approach to regulatory engagement, but did not satisfy every pre-registered criterion for execution detail.
Closed-loop control
Each miss sets a stricter next test.
- 01DetectCapture the exact failed criterion
- 02ClassifySeparate reasoning from completeness
- 03ControlDefine the corrective evaluation
- 04RetestRun a new held-out scenario
Evidence boundaries
Clear evidence. Honest limits.
We publish the scope, confidence range, and next validation step beside each claim, so you can judge the evidence on its merits.
01Investment recall is the current measure+
The current investment benchmark checks verified yes decisions. A follow-on protocol will add verified no decisions to measure specificity, false-positive rate, and overall calibration.
02Independent replication is the next layer+
Behavioral rubrics were fixed before scoring, and evaluators were blinded to the intended outcome. Independent replication is planned and will be published as a separate result.
03Confidence ranges reflect the sample size+
The page reports exact denominators and 95% confidence ranges for 60 behavioral tests and 24 protocol-qualified investment tests. Broader samples across people, domains, and time remain part of the published roadmap.
04Every tested MindSim twin must meet the study threshold+
Candidate inclusion follows a predefined evidence-sufficiency standard. Results apply to MindSim twins that meet it, while study coverage is reported separately from accuracy.
05Six evaluations were isolated by the blinding gate+
A pre-scoring quality gate separated six evaluations from the controlled set because they did not meet the final blinding standard. They remain visible under “All evaluated” for auditability.
Versioned evidence program
New tests. Published here.
Add verified declines
Add verified declines so specificity, false-positive rate, and overall decision calibration can be measured.
Independent retest
Replicate behavioral scoring with independent evaluators and publish the protocol and results separately.
Expand coverage
Expand sample sizes and measure how MindSim twin fidelity changes as the qualified evidence set evolves.
Measure fidelity gains
Run controlled ablations to quantify which evidence additions materially improve prediction quality.
Evidence ledger
Current release
- Controlled benchmarkMindSim-Bench controlled study90 total tests, 84 protocol-qualified
- Operational summary248 monitored responsesSnapshot through July 28, 2026
- Attributable detail76 row-level judgments67 matched, 9 fell short
The MindSim Prediction Engine
Test a MindSim twin.
Start with one person, one consequential scenario, and one prediction you can compare with what happens next.
Accuracy measures person-response prediction under the protocols shown here. It does not measure investment returns or broader business outcomes.