Last year I wrote that I had let go of the idea that validation had to mean scientific rigour alone, and that measurement is not the point but a portal: the starting point for a deeper conversation. I still believe that. This paper is not a retreat from it.
But a portal has to hold. Coaches build development plans on these numbers and organisations make decisions from them, so the instrument behind the conversation had better be sound. We now have 11,015 completed assessments from 10,417 people, which is finally enough to test it properly. What follows is what we found. Some of it is good. Some of it calls for caution about how the numbers are used rather than about the instrument itself. A little of it we need to fix, and we say which, because a validation that reports only good news is marketing.
The assessment measures something real and repeatable at the level of the whole profile, it detects development in proportion to how much development work was done, and its individual skill scores are noisier than the framework’s structure implies, which is why the profile, not any single number, is what we ask coaches to work from.
What follows on this page is a summary. It carries the headline findings and the reasoning behind them, but not the full analysis: the complete white paper runs to thirteen pages and includes the method, the per-scale figures, the statistical tests behind each claim, every limitation we are aware of, and the mistakes made while producing it. If you intend to rely on any of this, or to challenge it, read the paper rather than this page.
The Full Analysis
How reliable and valid is the IDG assessment?
Thirteen pages. The complete study: method, every figure, all limitations, and the corrections we had to make. This page covers roughly a third of it.
Two different questions
“Is it validated?” bundles together two things that are tested differently and can come apart. Reliability asks whether the instrument measures consistently: if several questions capture the same quality, people should answer them consistently. Validity asks whether it measures the right thing. Ten questions about shoe size would have excellent reliability and tell you nothing about presence.
This is the distinction most assessment marketing blurs. A high reliability score is often quoted as though it settled validity. It does not, and we have kept them apart throughout.
Reliability: strong as a whole, weak in parts
| Scale | Questions | Reliability |
|---|---|---|
| The whole instrument | 83 | 0.91 |
| Being | 54 | 0.89 |
| Thinking | 37 | 0.85 |
| Relating | 30 | 0.81 |
| Collaborating | 30 | 0.81 |
| Acting | 41 | 0.87 |
| Individual skills reaching 0.70 | — | 5 of 25 |
That gap looks alarming until you see why it happens. Reliability depends on how much the questions in a scale agree with each other and on how many questions there are. At this instrument’s level of agreement, the same questions of the same quality produce very different scores depending only on count: five questions can reach 0.47, six 0.51, thirteen 0.70, and all eighty-three 0.94.
A six-question scale cannot reach 0.70 here however well written it is. Critical thinking rests on six questions. So a low score on a single skill is substantially arithmetic rather than evidence of a badly built scale. In plain terms: a short scale is a short ruler. It still measures, just with wider gradations.
Short is not the same as faulty, and the two can be told apart. How well a scale’s questions agree with each other does not depend on how many there are. Measured that way, 19 of the 20 scales below 0.70 agree about as well as the rest of the instrument: they are short, not incoherent. Exactly one is genuinely weak, where the questions agree at 0.073 against a median of 0.138 across the 25. That one is a defect, and it is on the list of what we are changing.
The overall score and the five dimension scores can carry weight. A single skill score for a single person cannot, on its own. The shape of a profile is more trustworthy than any one number in it, and that is how our reports are written and how coaches should read them.
Do the five dimensions hold?
They do, but weakly, and the first test we ran said otherwise. Compared directly, two skills in the same dimension are no more related than two skills in different ones: 0.547 against 0.544, a difference of 0.003. Read alone, that says the dimensions are decorative.
But the 83 questions are shared between skills. Each question feeds 2.7 skills on average, and two skills built partly from the same answers are bound to look alike. Across all 300 pairs of skills, correlation rises directly with how many questions the pair has in common: 0.461 with none, 0.568 with one, 0.664 with two, 0.748 with three or more.
So the honest test is to compare only skills built from entirely separate questions.
| Pairs sharing no questions | Pairs | Average correlation |
|---|---|---|
| Same dimension | 34 | 0.494 |
| Different dimensions | 122 | 0.452 |
| Difference | +0.042, beyond chance |
The dimensions are real, and the question overlap was hiding them. Shared questions turn out to be spread across dimensions more often than within one, so the overlap inflates the cross-dimension side more, cancelling the genuine within-dimension excess. We had assumed the confound ran the other way and first read the null as ‘cannot tell’. It was masking a real effect.
Stated at its strongest and no further: the separation is small. Three of the five dimensions carry it clearly; two do not. All 25 skills remain strongly related to one another whichever way you cut it, and one broad factor still explains 57% of all the variation. The framework’s structure has support, not proof.
What the instrument gets right
It behaves as the theory predicts. If inner development genuinely accumulates with life experience, scores should rise with age. They do, in the right order, with no exceptions:
| Generation | Mean overall score | People |
|---|---|---|
| Baby Boomers | 81.3 | 150 |
| Generation X | 78.8 | 612 |
| Millennials | 76.2 | 799 |
| Generation Z | 73.4 | 1,470 |
It also carries a warning. Because the age effect is this large, comparing any group against a general average mostly measures its age composition. Every cohort comparison we publish is matched on generation for that reason, and any comparison that is not should be treated with suspicion, including comparisons of our own published averages.
Everyone’s profile is uneven, and consistently so. The distance between a person’s highest and lowest skill averages 26.7 points, and 97% of people exceed 15 points. An uneven profile is the normal condition, not a sign of imbalance, which is worth saying to anyone reading their own report for the first time.
The same organisation gives the same reading. One client was assessed three times over more than two years, 2,113 people in total. Taking each wave’s profile relative to everyone else on the platform, the shapes correlate at 0.93, 0.91 and 0.98. The same skills stand out every time.
Does it detect development?
Stability is only half an argument. An instrument that never moves is not stable, it is blind. At that same organisation, 293 individuals could be matched across two assessments by employee number: the same people, the same instrument, two points in time.
| Measure | Change | Beyond chance? |
|---|---|---|
| Overall score | +1.30 | Yes |
| All five dimensions | +1.16 to +1.64 | Yes, all five |
| Individual skills moving beyond chance | 18 of 25 | About 1 expected by chance |
An honest note on method: compared at organisation level rather than person level, the same change reads +0.51 and is indistinguishable from noise, because the people taking part differed between waves. Pairing individuals removed that and revealed the effect.
Most telling is that the size of the change tracks what was actually done:
| What was done | Change |
|---|---|
| Two years of individual development work | +6.6 |
| Comparable period, no targeted development | +3.0 |
| Structural redesign aligning people to their strengths, no inner-development work | +1.30 |
A structural intervention that never targeted inner development produced about a fifth of the movement seen in dedicated individual development. The instrument does not simply move; it moves in proportion to how much of the relevant work was actually done. That pattern is difficult to produce by accident.
What this does not establish
- It is a self-assessment. Every score reflects how a person sees themselves. We have no data linking scores to behaviour observed from outside the instrument, and until we do, no claim of that kind should be made on our behalf.
- The 25 individual skills are still not shown to be distinct from each other. The five dimensions separate; the 25 skills within them are a finer cut than the current question set can resolve.
- The change evidence rests on one organisation. 293 paired individuals is a reasonable sample, but it is one client, one sector, one country. No second organisation has yet been assessed twice at a size that could confirm it, so the finding is currently unreplicated. That is the biggest gap.
- Repeat measurement is not neutral. Growing self-awareness can lead someone to rate themselves lower because they see more clearly what a skill involves. A flat result after genuine development work is not necessarily a failure.
- A small repeat assessment cannot settle anything. Individual change varies widely, so detecting a shift of the size we observed needs roughly 75 people measured twice. Under about 20, the numbers cannot carry a verdict whatever they appear to say.
What we are changing
- 1.Reduce question sharing. 83 questions currently produce 221 skill assignments, and only 17 questions belong to a single skill. That overlap did not just obscure the dimension structure, it inverted the test for it.
- 2.Rebuild the one scale that does not hold together. Inclusive mindset is the single scale whose weakness is not explained by its length, and three individual questions do not move with the scale they belong to. Those are the genuine defects the analysis found.
- 3.Report confidence honestly in the product. Overall and dimension scores carry weight; individual skill scores are indicative, and we will make that distinction explicit.
- 4.Build the evidence we do not have. Replicate the change finding elsewhere, and design a study linking scores to something observed from outside the instrument.
- 5.Republish these figures as the population grows. Every number here is generated from the platform data by script, so this can be regenerated rather than rewritten.
Who did this analysis, and what that is worth
The analysis was carried out by an AI system (Claude, made by Anthropic), working directly against the platform database under my direction. I am saying so for the same reason we published the unflattering findings: you should be able to judge the work by how it was done.
What that buys is reproducibility, since every figure traces to a script that can be re-run and nothing was transcribed by hand, and an analyst with no career, grant or citation record riding on the answer. What it does not buy is independence. The work was commissioned by the company that sells the instrument, run on that company’s data, and reviewed by me before publication. An AI engaged by a vendor is not a disinterested third party and I am not going to claim it is.
Nor does it buy freedom from error. Three conclusions reached while producing this paper were wrong:
| First concluded | Why it was wrong | How it was caught |
|---|---|---|
| The five dimensions do not separate; the data cannot settle it | The test was confounded by shared questions, in the opposite direction to the one assumed. | I asked whether the framing was more negative than the evidence warranted. |
| Individuals cannot be matched across repeat assessments | Stated three times from incomplete checks. An employee number was populated for every participant, and it is the basis of this paper’s strongest finding. | I asked for it to be looked at again. |
| A client’s three assessments formed one series | Only two of them did. Treating the third as part of the series produced an apparent decline that was partly an artefact. | I knew the client’s history. |
In every case the error was caught by a person who knew the business, not by the system that made it. So the honest description of what an AI analyst contributed here is speed, consistency and reproducibility inside a process that still needed someone able to say ‘that does not sound right’. Anyone claiming an AI is more objective than a human institution should look at that table first.
We would rather this were checked than believed. If you work on measurement in this field and want to re-run any of it, we will help. The findings we would most like someone else to attack are the change result and the dimension structure.
Read the full study
Everything above is the short version. The white paper sets out how each figure was produced, reports the numbers this page leaves out, states the limitations in full, and shows the workings well enough for someone else to disagree with them properly. Anyone evaluating the instrument, rather than reading about it, should start there.
The Full Analysis
Download the complete white paper
Thirteen pages, including the method, the per-scale figures, the statistical tests behind every claim, and what we are changing as a result.
