About this app

The theory behind DCZion

Every team decision tool has a mathematical backbone. Ours rests on three deliberate choices:Condorcet's pairwise majority rule to rank options,Monte Carlo simulation to turn honest uncertainty into probabilities, and a transparent weighted-scoring model that refuses to throw your data away. This page explains why those choices — and why the popular alternatives, like AHP, fell short.

Part 1

Why Condorcet is the right way to rank options

The core question in any multi-option decision is deceptively simple: which option would the group actually prefer over every other option, one at a time? Condorcet's method is the oldest and most rigorous answer to that question — and it has a 240-year history behind it.

A short history

What a Condorcet winner actually means

An option is the Condorcet winner if it would beat every other option in a head-to-head majority contest. It's the only answer that survives being tested individually against each rival — the way a skeptical executive actually probes a recommendation. It is not the option with the highest average score (a popularity contest can crown a mediocre option that nobody hates) and it is not the option with the most first-place votes (a minority's favorite can beat it in every direct matchup).

The Jury Theorem, made practical

Condorcet's theorem assumes voters' judgments are independent and, on average, better than chance. DCZion operationalizes both assumptions: blind scoring creates the independence (no member can be swayed by watching another score first), and domain expertise in each criterion is what makes judgments better than chance. The theorem then does the rest: the group's aggregated judgment is far more likely to be correct than any single member's — even the loudest one in the room.

Part 2

Why not AHP — and why "no solution found" happens

The most famous academic alternative to Condorcet-style ranking is the Analytic Hierarchy Process (AHP), developed by Thomas Saaty in the early 1970s. AHP is taught in business schools and used in procurement and engineering decisions. It is also, in our judgment, a poor fit for real teams. Here is exactly why.

How AHP works (briefly)

AHP asks a decision-maker to compare every pair of criteria (and every pair of options per criterion) on a 1–9 "importance" scale, builds an n×n comparison matrix, then extracts weights via eigenvector math. It validates the result with a consistency ratio — a measure of how logically coherent all those pairwise judgments were.

Problem 1 — "No solution found" is a feature of the method

AHP rejects your answers when the consistency ratio exceeds 0.10 (10%). Real human judgments — made by busy people in a meeting — are rarely that consistent, so teams routinely hit the dead end where the method says: your pairwise comparisons contradict each other; redo them. In a live decision that means more meetings, more re-scoring, and sometimes abandoning the tool. DCZion never rejects the team's input: honest ranges that disagree are reflected in lower agreement and a smaller margin of victory, not an error screen.

Problem 2 — Hard to calculate, impossible to defend in a room

AHP's weights come from an eigenvector of a comparison matrix with a consistency index and random-index lookup table. Very few stakeholders can follow that math, which makes the final weights feel like a black box — and a black box is hard to defend when a losing faction asks why. DCZion uses plain weighted sums: score = Σ (priority × performance). Anyone in the room can reproduce the calculation on a napkin.

Problem 3 — The cognitive burden of n(n−1)/2 comparisons

AHP needs a pairwise judgment for every pair: 6 criteria → 15 comparisons; 10 criteria → 45. Multiply by options and members and the questionnaire becomes the project. Judgments degrade into noise as fatigue sets in — which is precisely what pushes the consistency ratio over 0.10 and triggers Problem 1. DCZion asks each member to score each option once, on its own merits, with a range that captures uncertainty. No matrix, no forced ratios, no second round of homework.

Problem 4 — Rank reversal

A well-documented flaw (first shown by Belton & Gear in 1983, replicated many times since): in AHP, adding or removing an irrelevant alternative can flip the ranking of the remaining options. Two options you were about to compare can swap places simply because a third, unrelated option was added to the list. For a decision tool whose whole purpose is defensibility, that is disqualifying. Condorcet-style pairwise comparison is immune to this class of artifact.

PropertyAHPDCZion
Accepts honest, imperfect human judgment✗ Rejects it (consistency ratio > 0.10 → redo)✓ Always — disagreement is measured, not rejected
Transparent, reproducible math✗ Eigenvector + consistency index✓ Weighted sum you can verify by hand
Uncertainty captured (ranges, not points)✗ Single 1–9 number per comparison✓ Confidence ranges sampled by Monte Carlo
Robust to adding/removing an option✗ Rank reversal✓ Pairwise — immune to irrelevant alternatives
Effort for 10 criteria✗ 45 pairwise judgments × options × members✓ One range per option per factor
Gives win probabilities, not just a ranking✗ No✓ "Beats all rivals in X% of simulations"
Part 3

Why ranges, not points — uncertainty is the honest input

DCZion asks members to score in confidence ranges (e.g. 7–9) instead of points (e.g. 8). This is not a nicety of the interface — it follows from how human judgment actually works, and from what a simulation needs as raw material.

Most of the time, humans don't know the point — they know the band

Decades of calibration research — the psychology of judgment under uncertainty (Kahneman & Tversky), and the classic interval-estimation studies that followed — show that people are systematically overprecise: when they state a 90% confidence interval, the true value lands inside it only about half the time, and even 98% intervals miss far more often than 2%.

An expert who knows "somewhere between 7 and 9" does not possess the number 8. Forced to name a point, they must pick one — and any choice asserts a certainty they don't have, and downstream math treats it as exact. The range is the information; the point is the guess. Formal theories agree: fuzzy logic (Lotfi Zadeh, 1965) was built on the observation that human reasoning is graded and range-like — "roughly," "between," "more or less" — rather than crisp. Even professional forecasts speak in ranges: "70% chance of rain," or the IPCC's "likely = 66–100%."

Intervals are what make the simulation possible

A point contains nothing to simulate. If every score were a fixed number, all 10,000 Monte Carlo draws would be identical — the app would collapse into plain averaging, with all its flaws: false precision, no win odds, no margin of victory, no fragility analysis.

An interval is the raw material of the simulation: each range (7–9) defines a distribution to draw from, and thousands of draws turn those ranges into probabilities. No intervals, no odds — the whole "beats all rivals in X% of simulations" result only exists because the inputs carry width.

The one piece of theory we actually lean on is the law of large numbers — the guarantee that with enough draws, sampling converges to the true answer. It is the same principle behind weather odds, option pricing, and engineering risk analysis. We don't need heavy machinery to use it: intervals wide enough to express honest uncertainty, and enough draws to converge.

A point estimate throws away half of what the expert knows

The width of the range is itself data: it is the member's stated confidence. Discard the width and you discard the second half of the signal — which is exactly why the app's disagreement diagnostics (dispersion, agreement %, spread) exist: they are only possible because inputs carry width. A point-estimate input has nothing left to measure.

Errors also compound. Multi-factor decisions multiply mistakes: when each factor's point estimate is silently treated as exact, small overconfidence cascades through the weighted sum. Ranges keep the uncertainty visible at every step, so the final answer inherits it honestly instead of pretending it was never there.

Used systematically, ranges give the more accurate answer

A range alone decides nothing — the system does. Sampled thousands of times by Monte Carlo, weighted relatively, and aggregated blindly, ranges become calibrated probabilities: a confident team gets a decisive recommendation; an uncertain team gets a hedged one. Both are correct behavior, and no point-estimate method can produce either.

Probability elicitation research reaches the same conclusion: asking for confidence-bearing ranges yields better aggregate forecasts than asking for best guesses, because proper scoring rules punish overconfident points.

"A point estimate is the part of your answer you're least sure about, presented as the part you're most sure about. A range is the part you actually know."

Part 4

Why Monte Carlo gives the best possible answer from your data

Once scores are ranges instead of points, the next question is how to combine them honestly. Monte Carlo simulation is not one option among many for this — for a weighted model with overlapping uncertainty ranges, it is the standard, and for most real configurations it is the only tractable way to get an exact answer.

A short history

There is no closed-form answer to compute

"What is the probability that AWS beats GCP, given six members' overlapping triangular ranges across five weighted factors?" is a high-dimensional integral with no analytic solution. You cannot write it down in a formula — you can only compute it by sampling. Monte Carlo is that computation: draw one value from every range, score every option, record the winner; repeat thousands of times. By the law of large numbers, the empirical win percentages converge to the true probabilities. With 10,000 draws, the answer is not an approximation of the best answer — it is the best answer available from the data, to whatever precision the draws allow.

It preserves uncertainty instead of deleting it

A point-estimate method (average the scores, rank the averages) discards the team's honesty: a member who knows "7–9" is forced to pretend they know "8". Monte Carlo keeps the range, so wide uncertainty becomes wide outcome distributions — and therefore lower win probabilities and smaller margins. The result tells you not just who wins, but how sure you should be. That is the difference between a recommendation and a gamble with a label.

It produces the full distribution, not one number

Because every trial is recorded, the output is a complete picture: beats-all percentage ("AWS beats every rival in 88.4% of runs"), margin of victory vs. the runner-up, pairwise win matrix for every pair, and per-factor average performances. Fragility analysis (what score change flips the winner?) is the same machinery re-run with one factor forced — a what-if question no point-estimate method can answer.

It is reproducible and auditable

With a fixed random seed, the same inputs always produce the same results — which makes the analysis reproducible for audit, dispute resolution, or a board review. Combined with the pairwise transparency of Condorcet, every number on the results screen can be regenerated on demand. That is what "defensible" means in practice.

Part 5

How the pieces fit together

For the exact formulas and the two aggregation modes (weighted score vs. criteria-bloc Condorcet), see themethodology section on the home page.

Part 6

Honest limits of the method

Sources

The theory's primary references

Put the theory to work on your next decision

Blind scoring, Monte Carlo odds, and a Condorcet winner — free to try on your own decision in 60 seconds.

Start a decision arrow_forward