MindSim BenchmarksContinuously updated

MindSim twins. Trust, earned.

Before you rely on a MindSim twin, you should know how closely it follows the person it represents. See the scores, test counts, protocol controls, and evidence boundary behind each claim.

Observed outcomes Attribution controls Every response tracked
88.2%operational accuracy
67 of 76scored predictions matched
248 responsestracked across the evidence pipeline

Wilson 95% interval: 79.0% to 93.6%. Every reviewed response remains visible below.

Follow the evidence

Ongoing operational cross-check

Real use. Measured outcomes.

We compare MindSim twin predictions with observed human outcomes during ongoing use. Every response is routed through evidence-quality and attribution controls before it can affect operational accuracy. This 88.2% ongoing result stays separate from the 97.6% controlled benchmark below.

248responses reviewedJuly 7 to July 28, 2026
88.2%scored outcomes matched67 correct / 76 scored
4.6 / 5mean persona-fidelity rating124 scored observations

Evaluation coverage

All responses are tracked.

248 total
Scored responses

Passed the attribution gate.

These responses had enough real-world evidence for a defensible judgment. Sixty-seven matched and nine fell short, so accuracy is 67 divided by 76.

Rounding policy

The measured result is 88.2%.

Elsewhere on the site, this result is rounded to 90% or 9 in 10. Here, the underlying ratio, sample size, evaluation window, and 95% interval remain visible: 67 of 76, with a range of 79.0% to 93.6%.

Controlled benchmark

Controlled conditions. Exact test count.

Each evaluation asks the MindSim twin to predict an outcome it cannot see, then compares its answer with criteria fixed in advance or a verified real-world decision. The controlled result includes only tests that passed every protocol control.

Protocol scope

Protocol-qualified set

97.6%82 passed / 84 tests

Includes only evaluations that passed every blinding and scoring control.

Behavioral prediction58 / 60
96.7%predictive correctness

Held-out professional scenarios scored against criteria fixed before evaluation.

Investment decision retrieval24 / 24
100%predictive correctness

Held-out company profiles checked against verified prior investment decisions.

ScopeThe current investment test measures recall of verified yes decisions. The next protocol adds verified no decisions to measure false positives and overall calibration.

What 97.6% means

19 of 21 MindSim twins passed every qualified test.

The other two each missed one detail-heavy question. Across all 84 tests, that is 82 matches and two misses. Both misses followed the person’s strategic direction but did not satisfy every required tactical detail.

Methodology

Predict first. Score against ground truth.

A MindSim twin can sound convincing and still be wrong. Before evaluation begins, the protocol fixes the scenario, scoring rule, known outcome, and quality gates, then keeps the outcome hidden from the MindSim twin.

01

Lock the protocol

The scenario, scoring criteria, ground truth, and quality gates are defined before the MindSim twin is evaluated.

02

Create the MindSim twin

The Prediction Engine derives a person-specific model of values, reasoning, and decision patterns from protocol-qualified evidence.

03

Run the test blind

The MindSim twin receives a held-out scenario without access to the known outcome or the answer required to pass.

04

Score against truth

Predefined behavioral criteria or a verified real-world decision determine whether the result is recorded as a match, miss, or protocol holdout.

Study coverage

Qualified tests. Credible scores.

Of 47 investor candidates, 30 completed evaluation. Twenty-four passed every protocol control and six were held outside the controlled result by the blinding gate. Five were prepared for a later test window, and 12 did not meet the study threshold.

47candidates

Protocol-qualified24

Blinding holdout6

Prepared, not yet tested5

Study threshold not met12

Data dictionary

How each term is used.

Protocol-qualified
An evaluation that passed every pre-scoring control, including the final blinding-quality gate. Six evaluations were held outside the controlled set.
Predictive correctness
A binary pass against pre-defined behavioral criteria or verified decision ground truth.
Scored MindSim twin response
An operational judgment where the observed outcome could be assigned to the MindSim twin rather than the surrounding integration.
Not scoreable
An observed response held outside the accuracy calculation because the available evidence does not support a defensible judgment.

Failure analysis

A miss is evidence. We trace every one.

You should know not only when a MindSim twin misses, but why. Every failed criterion is isolated, classified, and turned into a stricter held-out test. The miss stays visible.

A miss cannot disappear into the average.The 97.6% result remains paired with both failed evaluations and the exact criteria they did not satisfy.

2
misses published
2
rubric traces
2
next controls defined

Published misses

Case 01Execution-detail gap
Published and classified

Strategy aligned. Execution detail fell short.

One response matched the person’s approach to regulatory engagement, but did not satisfy every pre-registered criterion for execution detail.

What alignedThe person’s strategic approach to regulatory engagement.
What did not passFull execution-detail coverage required by the predefined rubric.
Next controlIncrease domain coverage before high-specificity use.

Closed-loop control

Each miss sets a stricter next test.

Evidence record retained
  1. 01
    DetectCapture the exact failed criterion
  2. 02
    ClassifySeparate reasoning from completeness
  3. 03
    ControlDefine the corrective evaluation
  4. 04
    RetestRun a new held-out scenario

Evidence boundaries

Clear evidence. Honest limits.

We publish the scope, confidence range, and next validation step beside each claim, so you can judge the evidence on its merits.

01Investment recall is the current measure

The current investment benchmark checks verified yes decisions. A follow-on protocol will add verified no decisions to measure specificity, false-positive rate, and overall calibration.

02Independent replication is the next layer

Behavioral rubrics were fixed before scoring, and evaluators were blinded to the intended outcome. Independent replication is planned and will be published as a separate result.

03Confidence ranges reflect the sample size

The page reports exact denominators and 95% confidence ranges for 60 behavioral tests and 24 protocol-qualified investment tests. Broader samples across people, domains, and time remain part of the published roadmap.

04Every tested MindSim twin must meet the study threshold

Candidate inclusion follows a predefined evidence-sufficiency standard. Results apply to MindSim twins that meet it, while study coverage is reported separately from accuracy.

05Six evaluations were isolated by the blinding gate

A pre-scoring quality gate separated six evaluations from the controlled set because they did not meet the final blinding standard. They remain visible under “All evaluated” for auditability.

Versioned evidence program

New tests. Published here.

Operational data throughJuly 28, 2026
Next01

Add verified declines

Add verified declines so specificity, false-positive rate, and overall decision calibration can be measured.

Planned02

Independent retest

Replicate behavioral scoring with independent evaluators and publish the protocol and results separately.

Scaling03

Expand coverage

Expand sample sizes and measure how MindSim twin fidelity changes as the qualified evidence set evolves.

Research04

Measure fidelity gains

Run controlled ablations to quantify which evidence additions materially improve prediction quality.

Evidence ledger

Current release

  • Controlled benchmarkMindSim-Bench controlled study90 total tests, 84 protocol-qualified
  • Operational summary248 monitored responsesSnapshot through July 28, 2026
  • Attributable detail76 row-level judgments67 matched, 9 fell short

The MindSim Prediction Engine

Test a MindSim twin.

Start with one person, one consequential scenario, and one prediction you can compare with what happens next.

Request a pilot Why we built MindSim

Accuracy measures person-response prediction under the protocols shown here. It does not measure investment returns or broader business outcomes.