22 worked examples across six kinds of decision. Each one shows what the team entered, what the engine returned, and the arithmetic in between — because a number nobody can reproduce is a number nobody trusts.
Figures are illustrative sample runs used to show the method, not customer data.
Someone has to choose between people. The panel walks in with four different impressions of the same interview, and the most senior voice tends to win. Here, every interviewer scores the same candidates against the same criteria before the debrief, and the panel sees where it agrees and where it does not.
Five interviewers, three finalists. Each interviewer scores all three candidates as a range before seeing anyone else's numbers.
| Criterion | Priority | Candidate A (domain expert) | Candidate B (hyper-growth leader) | Candidate C (generalist) |
|---|---|---|---|---|
| Execution & track record | 9 | 8.6 | 7.1 | 7.9 |
| Values & team multiplier | 8 | 7.4 | 8.9 | 8.0 |
| Strategic thinking & autonomy | 8 | 6.2 | 8.4 | 7.2 |
| Compensation & start date | 6.5 | 5.8 | 7.0 | 7.6 |
| Weighted total (priority × score, summed) | 223.9 | 247.8 | 242.1 |
Cells are the mean of the panel's estimates. Each interviewer entered a range (for example 7–10, not 8.5), and the ranges — not the means — are what the simulation samples, so one unsure interviewer widens the uncertainty instead of shifting the average.
Four peers plus the engineering manager, two candidates. Scored blind, so the peers cannot anchor on the manager's opinion and vice versa.
| Criterion | Priority | Candidate A (platform) | Candidate B (product) |
|---|---|---|---|
| Scope owned at the next level | 9 | 9.0 | 7.4 |
| Quality & craft | 8 | 8.6 | 9.1 |
| Collaboration & unblocking others | 8 | 7.9 | 8.5 |
| Communication | 7 | 8.2 | 7.6 |
| Weighted total (priority × score, summed) | 270.4 | 260.6 |
Cells are the mean of the panel's estimates. Each interviewer entered a range (for example 7–10, not 8.5), and the ranges — not the means — are what the simulation samples, so one unsure interviewer widens the uncertainty instead of shifting the average.
Six members, three candidates including one self-nomination. Everyone scores all three; nobody sees the running totals while scoring.
| Criterion | Priority | Candidate 1 | Candidate 2 | Candidate 3 |
|---|---|---|---|---|
| Trust | 9 | 8.4 | 7.6 | 8.0 |
| Delivery | 8 | 8.0 | 8.7 | 7.4 |
| Fairness & inclusivity | 8.5 | 8.8 | 7.0 | 8.2 |
| Availability | 7 | 7.2 | 8.6 | 7.9 |
| Weighted total (priority × score, summed) | 264.8 | 257.7 | 256.2 |
Cells are the mean of the panel's estimates. Each interviewer entered a range (for example 7–10, not 8.5), and the ranges — not the means — are what the simulation samples, so one unsure interviewer widens the uncertainty instead of shifting the average.
The quarterly or annual cycle, and the pay decisions that follow it. Everyone on the list is scored on the same criteria, blind, and the person running the review is scored too — their own row is their self-evaluation. The result is a per-person number for reviews, promotion and remuneration that the team can see the reasoning behind.
Everyone scores everyone against the agreed criteria, blind, and each person's own row is their self-evaluation. The lead is on the list like everyone else.
Six people, one pool. The band boundaries are agreed and frozen before scoring opens, so nobody can renegotiate the line after seeing the numbers.
Seven people across two squads. Each person is scored by their own squad plus two reviewers from the other squad.
Not a ranking of people but a two-option decision about one person: promote this cycle, or build the case and revisit.
Quarterly planning: five objectives, four initiatives, three roles, one budget. The work is choosing, and the argument is usually about which criterion matters — not about the facts. Scoring each candidate objective against weighted criteria turns that argument into a number the room can check.
Nine people across product, engineering, sales and support. Criteria: revenue impact (9), strategic fit (8), effort (7), risk (6). Scores below are the team's mean per objective, expressed on a 0–100 weighted scale.
Four initiatives competing for $400k of the quarter's budget. Same scoring as above, with TCO and opportunity cost added as criteria.
Platform engineer, product designer and data engineer are all argued for. Two slots.
Four running programmes, and the decision is which to wind down.
Market entry, positioning, build against buy, go or no-go. These are the decisions where a weighted total is least persuasive, because everyone can argue that their favourite criterion was underweighted. So the run reports the head-to-head result instead: how often option A beats option B when only the two are compared.
Eight people across sales, product and finance. Each scores the three routes per criterion as a range; the run samples those ranges 10,000 times and counts the head-to-head wins in each draw.
Row beats column in that share of 10,000 simulations. Green means the row option wins that pairing, red means it loses.
Six people, three positions, and a genuinely split room.
Row beats column in that share of 10,000 simulations. Green means the row option wins that pairing, red means it loses.
Seven people including finance and legal. Integration cost carries the highest priority and the widest ranges.
Row beats column in that share of 10,000 simulations. Green means the row option wins that pairing, red means it loses.
Roadmap bets, architecture, migrations, whether to stop and pay down debt. Engineering calls are usually not disputed on the facts — they are disputed on one estimate, made by people who see different parts of the system. This layout puts that estimate and the result side by side.
Six engineers and the product lead. Three criteria carry real disagreement: effort, reliability gain, and cost of delay.
Effort
Scored 4 by two engineers who have worked in the pipeline, 8 by two who have not — a 4-point spread, the widest in the run.
One point lower on effort for the rebuild (6.4 → 5.4) and the rebuild wins outright instead of tying.
Five people across platform and product engineering, scoring operational cost at the current team size.
Operational cost at our team size
Platform scored microservices 7–9, product scored it 3–5. The same option, a five-point split, decided entirely by who is carrying the pager.
If the headcount plan doubles, microservices takes the pairwise lead and the recommendation changes.
Four people, two options, and a hard external date.
Risk exposure in the busy window
Scored 3–8 depending on whether the scorer assumed the migration can slip. That assumption, not the technical work, is the whole disagreement.
A migration that can be paused mid-flight flips the answer to "now" on the same scores.
Eight people across engineering and support, with churn data from the previous quarter as shared evidence.
Cost of carrying the debt
Support scored it 8–9, engineering 4–6. Engineering is carrying it; support is hearing about it.
If the freeze runs two sprints instead of four, the debt work no longer outranks the committed features.
Procurement, tool selection, compliance sign-off, deal approval. These decisions are cross-functional by nature: the person who owns the budget, the person who owns the risk and the person who has to run the thing all have a say. A weighted sheet with the weights stated up front is how those three interests get compared instead of traded.
Sales, support, finance and IT score all three platforms before the demo debrief, so the loudest demo of the week cannot set the room.
| Criterion | Weight | Platform Alpha | Platform Beta | In-house build |
|---|---|---|---|---|
| Total cost of ownership | 8.5 | 6.0 | 8.4 | 5.2 |
| Security & compliance | 9.0 | 8.6 | 7.8 | 9.2 |
| Team experience | 8.0 | 7.4 | 8.8 | 5.4 |
| Implementation speed & lock-in | 7.5 | 6.2 | 8.0 | 4.6 |
| Weighted total | 234.1 | 272.0 | 204.7 |
Weights come from the group, not from the person who set up the run: each member's priorities are normalised against their own top criterion, then averaged, so a member who scores everything high does not outweigh a member who scores sparingly.
Six people, including the two who will run the migration and the one who owns the bill.
| Criterion | Weight | Provider X | Provider Y | Stay where we are |
|---|---|---|---|---|
| Total cost of ownership | 8.5 | 7.2 | 8.6 | 9.0 |
| Security & compliance | 9.0 | 9.1 | 8.4 | 8.8 |
| Team experience | 8.0 | 6.4 | 8.9 | 9.4 |
| Migration speed & lock-in | 7.5 | 5.8 | 8.2 | 9.6 |
Weights come from the group, not from the person who set up the run: each member's priorities are normalised against their own top criterion, then averaged, so a member who scores everything high does not outweigh a member who scores sparingly.
Four people: the security lead, the engineer who would build it, and two stakeholders who own the audit.
| Criterion | Weight | Buy | Build |
|---|---|---|---|
| Audit surface we own | 9.0 | 8.8 | 4.6 |
| Fit to our controls | 8.0 | 7.2 | 8.4 |
| Ongoing maintenance | 7.5 | 8.6 | 5.2 |
Weights come from the group, not from the person who set up the run: each member's priorities are normalised against their own top criterion, then averaged, so a member who scores everything high does not outweigh a member who scores sparingly.
Five stakeholders watching three pitch decks, scored before the internal debrief.
| Criterion | Weight | Agency A | Agency B | Agency C |
|---|---|---|---|---|
| Evidence from comparable work | 9.0 | 8.2 | 6.8 | 7.9 |
| Team we actually get | 8.5 | 7.6 | 8.4 | 6.2 |
| Cost & pace | 7.0 | 6.4 | 7.8 | 8.6 |
Weights come from the group, not from the person who set up the run: each member's priorities are normalised against their own top criterion, then averaged, so a member who scores everything high does not outweigh a member who scores sparingly.
The same five steps run under every example on this page, whether there are two options or five, three people or thirty.
A range (say 6–8) records honest uncertainty; a single number records false precision. The priority is 0–10 per criterion, and 0 means "this does not apply to me" — not a low score you are obliged to give.
Your 9 and my 9 are not the same 9, so each member's priorities are divided by their own highest one. Everyone's top criterion becomes 1.0 and the rest become fractions of it. A member who scores everything high no longer outweighs a member who scores sparingly, and a 0 stays "not applicable for me" without dragging the group's other weights down.
The averaged, normalized priorities round to the 0–10 priority shown on the screen, and the option scores are the mean of the members' estimates. If most of the team marks a criterion "does not apply", it is dropped from the run — the whole set is never dropped, and criteria the run drops stay visible in the audit. Members who said they were unsure have their ranges widened before averaging, so their doubt reaches the simulation instead of being flattened out.
Each draw takes one plausible score from every member's range and asks: in this version of the team's judgment, which option does the group prefer? Within each draw, options are compared pairwise. An option that beats every other option in a majority of draws is the Condorcet winner — a stronger claim than winning on a weighted total, because it holds in every one-on-one comparison. If the comparisons form a loop (A beats B, B beats C, C beats A) the run says a cycle was detected and reports the option with the most pairwise wins instead of inventing a winner.
Win share is the percentage of draws the recommendation won, and the margin against the runner-up is stated on the same line — a 51% result is labelled as fragile, not as settled. Disagreement diagnostics then rank the criteria by how far apart the team is, so the discussion goes to the one criterion that actually moves the answer. Sensitivity runs ask what would flip it. Everything is exported as a PDF or JSON, and the roster, reminders and audit trail show who took part.
Pick the closest template, replace the options with yours, and invite the people who should have a say. Free to start, no card.
Start a decision arrow_forward