Most assessments that coaches and consultants send to prospects are personality quizzes in a business suit. Twelve questions, a total score, and one of four flattering labels: "You're a Visionary Builder." The prospect enjoys it, shares it, and learns nothing they can act on. The coach learns even less. Every result sounds plausible, which is exactly the problem.
A well-built client assessment is a different instrument, and it deserves a different name: a diagnostic. It works the way a good intake conversation works: it asks everyone a few broad questions, digs deeper only where there is a signal, weighs evidence from more than one angle, and ends with a result that is specific enough to be wrong. That last part matters. A result that could not be wrong cannot be useful either.
This guide explains how to design one, step by step. It is written for business coaches, executive coaches and consultants who want an assessment that qualifies leads, structures the first session and earns trust, rather than one that merely collects email addresses. Throughout, we use a worked example, the Growth Bottleneck Diagnostic, which you can take yourself in about four minutes. Every rule described here is built into it, and every rule was tested before publication. If you like to design on paper first, there is also a free printable worksheet that follows the guide step by step.
Why most coaching assessments fail
In 1949 the psychologist Bertram Forer gave his students a personality test, then handed each of them the same written profile, assembled from statements in an astrology book. The students rated their "personal" profile as highly accurate: on average 4.26 out of 5. Vague, mostly positive statements feel personal because each reader fills them in with their own life. Psychologists call this the Barnum or Forer effect, and it is the engine behind most online quizzes.
Coaching assessments fall into the same trap in three predictable ways:
- They give everyone a label. If every respondent leaves as one of four "types", the assessment has sorted people, not diagnosed them. Nobody hears "nothing here needs fixing" or "this is not a coaching problem".
- They add everything into one total. A total of 62 out of 100 hides the one area that is genuinely broken behind four areas that are fine. The owner who cannot take a week off and the owner whose prices lose money on every job can score the same.
- Their results don't change what happens next. If every result leads to the same "Book a call" button with the same pitch, the assessment was decoration. The prospect senses it.
A diagnostic fixes all three. It can say "nothing to fix", it scores specific patterns rather than a single total, and each result leads to a different next step.
Quiz, survey or diagnostic: know which one you're building
Three kinds of question-based instrument look alike on screen and do completely different jobs:
| Instrument | Its job | What a good one produces |
|---|---|---|
| Quiz | Engage and entertain; sort people into a category | A shareable label |
| Survey | Measure a group; describe opinions or behaviour across many people | Percentages and trends |
| Diagnostic | Decide, for one person, what to look at, how serious it is and what to do next | A specific finding and a matching next step |
The design choices follow from the job. A quiz wants every question answered by everyone, because the fun is in the journey. A survey wants everyone to answer the same questions, because comparability is the point. A diagnostic wants the opposite of both: different people answer different questions, because the questions that matter depend on what the earlier answers revealed.
If you are a coach or consultant, you almost certainly want a diagnostic, even if you will also use the results in aggregate later.
Start with the decisions, not the questions
The most common way to build an assessment is to brainstorm good questions, add them up, and then decide what the results should say. It is backwards. Start from the end: list every distinct thing you would actually do differently for a prospect, and make each one a result.
For the Growth Bottleneck Diagnostic, the list of next actions looked like this:
- Refer them to an exit-planning specialist (they are preparing to sell or close).
- Tell them honestly that now is not the time (they have no hours to give).
- Start coaching on one of five specific patterns, beginning with the most important.
- Flag an early sign before it becomes a pattern, with one preventive step.
- Confirm that the foundations are solid and help them choose a growth goal.
That produced twelve endings: two route endings, five pattern endings, four early-sign endings and one "solid foundation". Only then did we write a single question. Every question in the final assessment exists because it helps choose between those twelve endings. If a question did not, it was cut.
Common mistake: writing more results than next actions. If two results lead to the same conversation, merge them. Extra results feel richer to the designer and read as filler to the respondent.
The four-stage architecture
Clinical screening often uses the same shape. In primary care, for example, a two-question depression screen, the PHQ-2, is asked first; only people who score above its cut-off go on to fuller questioning. Kroenke, Spitzer and Williams (2003) found that the two questions alone identified major depression with 83% sensitivity and 92% specificity at the recommended cut-off. Ask briefly, then ask deeply only where it is needed. The same shape works for business diagnosis.
Each stage has one job:
- Routing (everyone). A few questions that decide whether this is the right conversation at all.
- Screeners (everyone who continues). One short question per area, each able to open that area's deeper questions.
- Depth (only for areas that opened). The questions that actually score the patterns.
- Closing (everyone who continues). The respondent's own view of what matters most, how urgent it feels, and their contact details.
The result is that different people answer very different numbers of questions. In the Growth Bottleneck Diagnostic, someone preparing to sell answers three; someone with problems in every area answers twenty; a typical healthy business answers nine. That variation is the point. Galesic and Bosnjak (2009) found that announcing a longer web survey reduced the number of people who started it, and that answers late in a long questionnaire were faster, shorter and more uniform: classic signs of what Krosnick (1991) called satisficing, giving a good-enough answer instead of a considered one. A diagnostic spends the respondent's attention only where it buys information.
Common mistake: showing a progress bar or question numbers. When paths differ in length, "Question 6 of 20" is false for almost everyone, and a bar that jumps from 40% to 90% looks broken. Turn both off.
Routing questions: some answers outrank every score
Some answers change what the right next step is, whatever else the person says. In our example there are two:
- "I'm preparing to sell, step back or close it." Exit planning is specialist work. A pattern score of 8 out of 8 for Margin Leak does not change that this person needs a different adviser first.
- "Honestly, almost none right now" (in answer to how much time they can give). A growth plan without protected hours fails. The honest result is "make room first".
Two design rules follow. First, route results are checked before anything else, so no score can override them. Second, route results need their own order: if someone is both selling and out of time, which wins? We chose "selling" (exit planning first), wrote that decision down, and tested it. Every tie in a diagnostic needs a rule like this, decided on purpose rather than by accident.
Routing also controls the path. People on a route result skip the screeners and depth sections entirely and go straight to the contact step. There is no point asking about pricing discipline to someone closing the business next quarter.
Screeners that open depth only where it's needed
A screener is one question per area, written so its answers form a ladder of severity. The top rung means "clear", the next means "a mild signal", and the bottom two open the area's depth questions. Here is the Demand screener from our example:
Three details make screeners work.
Write answers as observable behaviour, not adjectives. "It depends on referrals and luck" is something an owner can recognise in their own business. "Poor" is a judgement they may not want to make about themselves. Smith and Kendall (1963) made this idea central to behaviourally anchored rating scales: anchor each point with concrete behaviour so different people read it the same way. Describing what happens, rather than asking for a grade, may also ease some of the pressure to give a flattering answer, a pressure Tourangeau and Yan (2007) document in their review of sensitive questions.
Keep the mild signal. "Mostly, with the odd quiet month" does not justify four more questions, but it is not nothing. A diagnostic that records it can still say something useful to a business with no serious pattern ("early signs: new business is starting to wobble"), instead of a generic "all good".
Let one question open more than one area. Our fifth screener asks, "When a month goes badly, what's usually behind it?" Each answer points at a different area: not enough new work (Demand), work late or over budget (Delivery), key people stretched or leaving (Team), things waiting on the owner (Owner). An owner who answers the Demand screener confidently but blames bad months on missing work has given you a reason to look anyway. Two independent routes into each area catch what either alone would miss.
Score patterns, not totals
Inside each opened area, the depth questions score patterns: specific, named problems with a known shape and a known first move. The Owner area, for example, holds two patterns, Founder Bottleneck and Busy but Stuck, because "too much depends on the owner" and "the owner has no time for the important work" are different problems with different fixes. A single average across them would hide exactly the difference that decides what to do. (We made a shorter case against averaging in Building a Client Diagnostic That Doesn't Feel Like a Form.)
Each pattern is scored from exactly two questions that approach it from different angles. For Founder Bottleneck, one question asks what happens when a team member meets a new problem (do they decide, suggest, wait, or does it sit?). The other asks how the business would cope if the owner took two weeks off with no phone. The first looks at daily decisions, the second at dependence. Answers score 1 to 4 points, so each pattern scores between 2 and 8.
A pattern fires at 6 or more. Here is what that means in practice:
| Answers to the two questions | Points | Fires? | Why |
|---|---|---|---|
| C and C | 3 + 3 = 6 | Yes | Both angles show a real problem |
| D and B | 4 + 2 = 6 | Yes | One severe signal, backed up by the other |
| D and A | 4 + 1 = 5 | Only if the respondent names it | The two angles contradict each other |
| B and C | 2 + 3 = 5 | No | Moderate on both: an early sign, not a pattern |
| A and A | 1 + 1 = 2 | No | Clear |
The interesting row is "D and A". One answer is the worst possible, the other the best. Is the owner a bottleneck or not? The assessment cannot tell from those two answers alone, so it borrows the respondent's own judgement. Near the end, everyone who saw any depth questions is asked: "Of everything you've answered, which feels like the biggest drag on growth right now?" If they name this pattern, the contradictory 5 fires; otherwise it does not. The person living with the business breaks the tie.
Why two questions rather than one, or ten? One question is a single point of failure: a misread word or an unusual week decides the result. Ten is a research instrument, not a sales-stage diagnostic. Two questions from different angles are the smallest number that can disagree with each other, which is exactly what lets the "D and A" case be detected rather than silently averaged away. Be honest about what this is, though: a two-item score is a screening judgement, not a precise measurement. Eisinga, te Grotenhuis and Pelzer (2013) discuss how even estimating the reliability of a two-item scale needs care; the validity section below covers what to do about it.
Common mistake: scoring a pattern from questions in other areas. Screeners open areas; they do not score patterns. Keep every pattern's score inside its own two depth questions, and you can always explain to a client exactly why they got their result.
Ties, primary and secondary results
As soon as more than one pattern can fire, you need a rule for which one leads. Ours, in order:
- The respondent's own pick wins if it fired. If they say the biggest drag is "too much depending on me" and Founder Bottleneck fired, that is the primary result, even if Margin Leak scored higher. People act on the problem they already feel.
- Otherwise, the highest score wins: any fired pattern scoring 8, then 7, then 6.
- Equal scores are broken by a fixed, written order. Ours is Founder Bottleneck, Busy but Stuck, Feast-or-Famine Pipeline, Margin Leak, Hero Culture. Owner patterns come first because they cap every other fix: a new sales process stalls if every quote still waits for the owner.
The secondary result is the highest-scoring of the other fired patterns, by the same rules. It does not get its own ending, but it is not wasted: the primary ending shows a short note ("Also showing up: Margin Leak. Once your first move is under way, it is the pattern to look at next."), and the coach sees it before the first call.
When three or more patterns fire, the respondent sees one extra screen before their result, explaining that several patterns are working together and that the next screen shows where to start. Without it, a person with five problems receives a result about one of them and wonders whether the assessment noticed the other four.
None of these rules is the "right" one. A tie order could reasonably put pipeline first for a sales-led coach. What matters is that the rule exists, is written down, and gives the same answer every time.
The ending ladder: first match wins
All of the rules above combine into a single ordered list, checked from the top. The first rule that matches chooses the ending, and nothing below it is considered.
Writing the endings as a ladder has two benefits. It makes the logic reviewable by someone who is not the designer: a colleague can read six lines and check that the order matches what the practice actually wants. And it guarantees that every respondent lands somewhere sensible, because the bottom rung catches everyone. "Solid foundation" is not a failure mode; for a healthy business it is the correct and valuable answer, and it builds more trust than a manufactured problem.
Data that respects the people answering
An assessment collects candid answers about someone's business, and sometimes about their health, relationships or money. How you handle that data is part of the design, not an afterthought.
- Use stable codes for results, not prose. When a result goes to your CRM, tag it with a short code, such as
GBD-P-MARGIN-LEAK, not the ending's wording. Codes survive copy edits, and automations that depend on them keep working. - Keep answers and personal details out of links. If a result button links to a booking page or a fuller report, pass codes only. Never put an email address, a name or answer text in a URL, where it ends up in browser history, analytics and server logs.
- Decide what is saved before submission. Many form tools save answers as people go. For a business diagnostic that is often fine. For anything sensitive, it means data exists for people who decided not to submit, which you should not collect.
- Plan a safe way out on sensitive topics. If an assessment touches on wellbeing, conflict or anything that could cause distress, give people a visible way to leave that stores nothing, and put support information where they will see it.
- Say what happens to the answers at the point you ask for contact details, in one plain sentence.
Respondents notice care. Tourangeau and Yan's review of sensitive questions in surveys found that worries about disclosure lead people to skip questions or answer less honestly. A diagnostic is only as good as the candour it receives.
Test it like software
A diagnostic with five patterns, four areas and a ladder of endings has thousands of possible paths. "I clicked through it a few times and it looked fine" tests perhaps five of them. The bugs live in the other paths: the tie nobody tried, the area that opens but never routes back, the ending no one can reach.
The fix is to write test paths before you build, straight from the design: a set of answers in, and the expected result out. Here are some of the eighteen we wrote for the Growth Bottleneck Diagnostic (question counts include the contact step):
| Test | Answers (all others healthy) | Expected result | Questions shown |
|---|---|---|---|
| Nothing flagged | None | Solid foundation | 9 |
| Selling beats no time | Preparing to sell; no time | Exit planning first | 3 |
| Area opens, nothing fires | Delivery screener C; pricing B; overruns C | Early signs: Delivery | 12 |
| A 5 with a D fires only when picked | Demand screener D; pipeline answers D and A; picks "next client" | Feast-or-Famine Pipeline | 12 |
| Pick beats a higher score | Pipeline scores 6, Margin Leak 8; picks "next client" | Feast-or-Famine Pipeline, then Margin Leak | 14 |
| Tie at 7 uses the written order | Founder Bottleneck 7, Margin Leak 7 | Founder Bottleneck, then Margin Leak | 16 |
| Three fire | Pipeline 6, Hero Culture 8, Founder Bottleneck 6; picks "a few people" | Overlay screen, then Hero Culture | 18 |
Good test paths share three habits. Each one checks one rule, named in its title. Together they cover every ending and every tie rule. And they check more than the ending: the number of questions shown, the primary and secondary result, the urgency flag and the tags, because a correct ending reached by the wrong route is a bug waiting for a different respondent.
For the example in this guide we went one step further. We wrote the rules a second time, independently, straight from the written design, and compared the two versions on all 1,287 routing and screener combinations and on 2,500 random complete paths. They agreed on every one. You do not need to go that far for every assessment. But if your diagnostic decides who gets a sales call, or what a client is told about their business, the paths deserve the same testing as anything else that makes decisions.
Common mistake: testing only after launch, by watching the results come in. By then real prospects have received wrong results, and you cannot tell which ones.
Validity in plain English, and its limits
"Is it valid?" is the question a thoughtful client will eventually ask. Messick's (1995) influential answer is that validity is not a property of the questions but of the conclusions you draw from the answers, and that it needs several kinds of evidence. The Standards for Educational and Psychological Testing (AERA, APA and NCME, 2014) take the same view. For a coaching diagnostic, four questions cover most of it:
- Does each pattern's pair of questions cover the pattern? Ask two experienced colleagues to read each pair blind and name the problem it describes. If they cannot, rewrite it. This is content validity, and it is cheap.
- Do results match what you find in the first session? After twenty or thirty clients, compare the assessment's primary result with your own judgement after the first conversation. Where they disagree, find out which question misled.
- Does everyone get the same result? Look at how results are distributed. If 70% of respondents land on one pattern, you have rebuilt the Barnum effect. Tighten that pattern's questions or raise its threshold.
- Does the result change what happens? If the next step is the same for every ending, the diagnostic is decoration, however well it scores.
Be candid about the limits. A diagnostic like this is a structured conversation starter and a decision aid. It is not a psychometric instrument, it has not been normed, and two questions per pattern are a screen, not a measurement. Saying so in your materials costs nothing and earns credibility with exactly the clients you want.
What it takes to build this
Everything above is design. The design then has to survive contact with a form builder. Most form tools were built for contact forms, simple quizzes and surveys, and a diagnostic stresses them in specific places. Here is what the job requires, and how MarketKloud Forms, our own form builder, handles each requirement. It is the tool we used to build the example in this guide.
One language from design to build
A diagnostic is designed on paper with names like SC-DEM and FB1, then built, then reviewed. If the builder calls the same question "Question 14", every review becomes a translation exercise. MarketKloud Forms lets every question carry a question code like SC-DEM or FB1, shown in the builder, in the logic, in results and in exports, so the design document and the live form use one language.
Routing that follows the design exactly
Each question has its own ordered list of rules, checked top to bottom, first match wins, with an "otherwise" destination at the bottom: precisely the ending ladder above. Rules can combine any number of conditions with and and or, test answers or calculated values, and jump forward to any question or straight to an ending. A visual map of the whole form shows every path as one left-to-right flow, so you can see at a glance where the depth sections branch off and rejoin.
Points, variables and formulas
Pattern scoring needs more than a quiz's single total. You can give points per answer from a table, into any number of named scores or variables. Rules can add to a counter ("Owner area opened") or set a variable to a formula written in plain text with brackets, comparisons and IF, THEN, ELSE IF chains. Choice answers can be written as the letter the respondent sees ({{FOCUS}} = D), and long formulas display one case per line so a colleague can review them. For ranking problems, a built-in Pick by rank finds the highest or lowest of a set of scores, with explicit tie-breaking.
Endings that say something specific
A diagnostic needs many endings, each with its own message and next step. MarketKloud Forms supports as many endings as the design calls for, each with its own button, and result blocks: extra paragraphs shown on an ending only when a formula is true. That is how the example adds "Also showing up: Margin Leak" or an urgency note without multiplying endings. Ending buttons can carry the respondent's results as codes in the link, and respondents can download their result as a PDF.
Testing built into the builder
This is where most tools stop, and where the difference matters most. In MarketKloud Forms, test paths are part of the form: under Logic, Tests, you save each path's answers with its expected ending, number of questions, values and tags, and Run all tests replays every one through the same engine the live form uses. Check every path plays up to thousands of answer combinations (every one, when the form has fewer) and reports endings nobody can reach and rules that never apply. When a test fails, Why it went there shows, screen by screen, which rule fired and what each value was. A logic check runs before you publish and flags rules that point at deleted questions, compare against options a question does not have, or can never be true.
Privacy and safety controls
For sensitive diagnostics, privacy mode stores answers only when the respondent submits, with nothing saved along the way and nothing left in the browser. A quick exit button can sit on every screen, and safety exit endings store and send nothing at all. Check results on the server replays each submission's answers through the published logic and records any difference from what the browser reported, so the stored result is the one your rules produce.
Results that feed the next step
Every response can carry tags chosen by rules, such as the primary pattern or an urgency flag, and Results can be filtered by ending or tag and exported as CSV. Responses can be sent to Zapier or Make through signed webhooks, with the option to skip chosen endings, so a "make room first" result never triggers a sales sequence.
Changing it safely, and reusing it
A diagnostic gets refined after its first fifty responses. Edits to a live form stay as unpublished changes until you publish them, every saved test runs again automatically when you click Publish, and Version History keeps earlier versions. A finished assessment can be downloaded as a single form file, questions, endings and logic included, and used as the starting point for the next client.
| A diagnostic needs | In MarketKloud Forms |
|---|---|
| Different questions for different people | Per-question rules, first match wins, jumps to any question or ending; visual map |
| Pattern scores, not one total | Points tables into named scores; counters; formulas with IF / ELSE IF; Pick by rank |
| Deliberate tie-breaking | Rule order you control; formulas; ranked tie order |
| Specific results with matching next steps | Unlimited endings, result blocks, per-ending buttons, PDF results |
| Evidence that every path works | Saved test paths, Run all tests, Check every path, Why it went there, publish-time logic check |
| Care with sensitive answers | Privacy mode, quick exit, safety exit endings, server-checked results |
| Results that drive follow-up | Rule-based tags, ending and tag filters, CSV export, signed Zapier/Make webhooks |
| A shared language with the design | Question codes everywhere, including formulas and exports |
| Safe refinement and reuse | Unpublished changes, Version History, form files |
The quickest way to judge all of this is to use it. The Growth Bottleneck Diagnostic is in the MarketKloud Forms template gallery with its logic, endings and all eighteen test paths included. Take it as a respondent, then open it as a template, change the questions to suit your practice, and run the tests again.
The diagnostic design checklist
Before you build, and again before you publish. (This checklist, with a fill-in page for every step above, is in the printable worksheet.)
- Every result leads to a different next action; results with the same action are merged.
- Route answers that outrank any score are checked first, in a written order.
- Everyone answers routing and screener questions; depth questions appear only where opened.
- Progress bars and question numbers are off when paths differ in length.
- Screener answers describe observable behaviour and form a clear ladder from clear to severe.
- Mild signals are recorded, and can produce an early-sign result.
- Each area can be opened by more than one route.
- Each pattern is scored only from its own questions, at least two, from different angles.
- The firing threshold, and what happens to contradictory answers, are written down.
- The respondent's own view of what matters most is asked, and used.
- Primary and secondary results follow one written rule, with a fixed tie order.
- Several patterns at once are acknowledged, not hidden.
- The endings form one ordered list, with a sensible result at the bottom for everyone else.
- Results leave the form as stable codes; no answers or personal details appear in links.
- What is saved before submission is a deliberate decision.
- Sensitive topics have a safe way out and visible support information.
- Test paths cover every ending and every tie rule, and check more than the ending.
- Tests run again before every change goes live.
- Results are compared with first-session judgement after the first twenty or thirty clients.
- The materials say plainly what the assessment is, and what it is not.
A diagnostic built this way does something a quiz never can: it tells some people they are fine, tells others that you are not the right person yet, and tells the rest something specific enough that the first session starts halfway through. That is what earns the call.
Sources
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. AERA.
- Eisinga, R., te Grotenhuis, M., & Pelzer, B. (2013). The reliability of a two-item scale: Pearson, Cronbach, or Spearman-Brown? International Journal of Public Health, 58(4), 637–642.
- Forer, B. R. (1949). The fallacy of personal validation: A classroom demonstration of gullibility. Journal of Abnormal and Social Psychology, 44(1), 118–123.
- Galesic, M., & Bosnjak, M. (2009). Effects of questionnaire length on participation and indicators of response quality in a web survey. Public Opinion Quarterly, 73(2), 349–360.
- Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2003). The Patient Health Questionnaire-2: Validity of a two-item depression screener. Medical Care, 41(11), 1284–1292.
- Krosnick, J. A. (1991). Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3), 213–236.
- Messick, S. (1995). Validity of psychological assessment. American Psychologist, 50(9), 741–749.
- Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149–155.
- Tourangeau, R., & Yan, T. (2007). Sensitive questions in surveys. Psychological Bulletin, 133(5), 859–883.