Thought LeadershipHow Reliable and Valid Is the IDG Assessment?
22 September 2026Mark Vandeneijnde

How Reliable and Valid Is the IDG Assessment?

How Reliable and Valid Is the IDG Assessment?

Last year I wrote that I had let go of the idea that validation had to mean scientific rigour alone, and that measurement is not the point but a portal: the starting point for a deeper conversation. I still believe that. This paper is not a retreat from it.

But a portal has to hold. Coaches build development plans on these numbers and organisations make decisions from them, so the instrument behind the conversation had better be sound. We now have 11,015 completed assessments from 10,417 people, which is finally enough to test it properly. What follows is what we found. Some of it is good. Some of it calls for caution about how the numbers are used rather than about the instrument itself. A little of it we need to fix, and we say which, because a validation that reports only good news is marketing.

Executive Summary

The assessment measures something real and repeatable at the level of the whole profile, it detects development in proportion to how much development work was done, and its individual skill scores are noisier than the framework’s structure implies, which is why the profile, not any single number, is what we ask coaches to work from.

What follows on this page is a summary. It carries the headline findings and the reasoning behind them, but not the full analysis: the complete white paper runs to thirteen pages and includes the method, the per-scale figures, the statistical tests behind each claim, every limitation we are aware of, and the mistakes made while producing it. If you intend to rely on any of this, or to challenge it, read the paper rather than this page.

The Full Analysis

How reliable and valid is the IDG assessment?

Thirteen pages. The complete study: method, every figure, all limitations, and the corrections we had to make. This page covers roughly a third of it.

Download PDF

Two different questions

“Is it validated?” bundles together two things that are tested differently and can come apart. Reliability asks whether the instrument measures consistently: if several questions capture the same quality, people should answer them consistently. Validity asks whether it measures the right thing. Ten questions about shoe size would have excellent reliability and tell you nothing about presence.

This is the distinction most assessment marketing blurs. A high reliability score is often quoted as though it settled validity. It does not, and we have kept them apart throughout.

Reliability: strong as a whole, weak in parts

ScaleQuestionsReliability
The whole instrument830.91
Being540.89
Thinking370.85
Relating300.81
Collaborating300.81
Acting410.87
Individual skills reaching 0.705 of 25
By convention 0.70 is acceptable and 0.80 good. The instrument as a whole and every dimension clear that comfortably. Most individual skills do not.

That gap looks alarming until you see why it happens. Reliability depends on how much the questions in a scale agree with each other and on how many questions there are. At this instrument’s level of agreement, the same questions of the same quality produce very different scores depending only on count: five questions can reach 0.47, six 0.51, thirteen 0.70, and all eighty-three 0.94.

A six-question scale cannot reach 0.70 here however well written it is. Critical thinking rests on six questions. So a low score on a single skill is substantially arithmetic rather than evidence of a badly built scale. In plain terms: a short scale is a short ruler. It still measures, just with wider gradations.

Short is not the same as faulty, and the two can be told apart. How well a scale’s questions agree with each other does not depend on how many there are. Measured that way, 19 of the 20 scales below 0.70 agree about as well as the rest of the instrument: they are short, not incoherent. Exactly one is genuinely weak, where the questions agree at 0.073 against a median of 0.138 across the 25. That one is a defect, and it is on the list of what we are changing.

The Practical Rule

The overall score and the five dimension scores can carry weight. A single skill score for a single person cannot, on its own. The shape of a profile is more trustworthy than any one number in it, and that is how our reports are written and how coaches should read them.

Do the five dimensions hold?

They do, but weakly, and the first test we ran said otherwise. Compared directly, two skills in the same dimension are no more related than two skills in different ones: 0.547 against 0.544, a difference of 0.003. Read alone, that says the dimensions are decorative.

But the 83 questions are shared between skills. Each question feeds 2.7 skills on average, and two skills built partly from the same answers are bound to look alike. Across all 300 pairs of skills, correlation rises directly with how many questions the pair has in common: 0.461 with none, 0.568 with one, 0.664 with two, 0.748 with three or more.

So the honest test is to compare only skills built from entirely separate questions.

Pairs sharing no questionsPairsAverage correlation
Same dimension340.494
Different dimensions1220.452
Difference+0.042, beyond chance
Not a scale-length artefact: the two groups are balanced on questions per scale (8.2 against 8.3) and on reliability (0.561 against 0.563).
A Correction We Had To Make

The dimensions are real, and the question overlap was hiding them. Shared questions turn out to be spread across dimensions more often than within one, so the overlap inflates the cross-dimension side more, cancelling the genuine within-dimension excess. We had assumed the confound ran the other way and first read the null as ‘cannot tell’. It was masking a real effect.

Stated at its strongest and no further: the separation is small. Three of the five dimensions carry it clearly; two do not. All 25 skills remain strongly related to one another whichever way you cut it, and one broad factor still explains 57% of all the variation. The framework’s structure has support, not proof.

What the instrument gets right

It behaves as the theory predicts. If inner development genuinely accumulates with life experience, scores should rise with age. They do, in the right order, with no exceptions:

GenerationMean overall scorePeople
Baby Boomers81.3150
Generation X78.8612
Millennials76.2799
Generation Z73.41,470
A near eight-point climb from the youngest group to the oldest. This is known-groups validity: the instrument distinguishes groups that theory says should differ, in the direction theory predicts.

It also carries a warning. Because the age effect is this large, comparing any group against a general average mostly measures its age composition. Every cohort comparison we publish is matched on generation for that reason, and any comparison that is not should be treated with suspicion, including comparisons of our own published averages.

Everyone’s profile is uneven, and consistently so. The distance between a person’s highest and lowest skill averages 26.7 points, and 97% of people exceed 15 points. An uneven profile is the normal condition, not a sign of imbalance, which is worth saying to anyone reading their own report for the first time.

The same organisation gives the same reading. One client was assessed three times over more than two years, 2,113 people in total. Taking each wave’s profile relative to everyone else on the platform, the shapes correlate at 0.93, 0.91 and 0.98. The same skills stand out every time.

Does it detect development?

Stability is only half an argument. An instrument that never moves is not stable, it is blind. At that same organisation, 293 individuals could be matched across two assessments by employee number: the same people, the same instrument, two points in time.

MeasureChangeBeyond chance?
Overall score+1.30Yes
All five dimensions+1.16 to +1.64Yes, all five
Individual skills moving beyond chance18 of 25About 1 expected by chance
Replicated on a separate subset of 102 people measured at a third point: +1.29.

An honest note on method: compared at organisation level rather than person level, the same change reads +0.51 and is indistinguishable from noise, because the people taking part differed between waves. Pairing individuals removed that and revealed the effect.

Most telling is that the size of the change tracks what was actually done:

What was doneChange
Two years of individual development work+6.6
Comparable period, no targeted development+3.0
Structural redesign aligning people to their strengths, no inner-development work+1.30
The Strongest Evidence Here

A structural intervention that never targeted inner development produced about a fifth of the movement seen in dedicated individual development. The instrument does not simply move; it moves in proportion to how much of the relevant work was actually done. That pattern is difficult to produce by accident.

What this does not establish

  • It is a self-assessment. Every score reflects how a person sees themselves. We have no data linking scores to behaviour observed from outside the instrument, and until we do, no claim of that kind should be made on our behalf.
  • The 25 individual skills are still not shown to be distinct from each other. The five dimensions separate; the 25 skills within them are a finer cut than the current question set can resolve.
  • The change evidence rests on one organisation. 293 paired individuals is a reasonable sample, but it is one client, one sector, one country. No second organisation has yet been assessed twice at a size that could confirm it, so the finding is currently unreplicated. That is the biggest gap.
  • Repeat measurement is not neutral. Growing self-awareness can lead someone to rate themselves lower because they see more clearly what a skill involves. A flat result after genuine development work is not necessarily a failure.
  • A small repeat assessment cannot settle anything. Individual change varies widely, so detecting a shift of the size we observed needs roughly 75 people measured twice. Under about 20, the numbers cannot carry a verdict whatever they appear to say.

What we are changing

  1. 1.Reduce question sharing. 83 questions currently produce 221 skill assignments, and only 17 questions belong to a single skill. That overlap did not just obscure the dimension structure, it inverted the test for it.
  2. 2.Rebuild the one scale that does not hold together. Inclusive mindset is the single scale whose weakness is not explained by its length, and three individual questions do not move with the scale they belong to. Those are the genuine defects the analysis found.
  3. 3.Report confidence honestly in the product. Overall and dimension scores carry weight; individual skill scores are indicative, and we will make that distinction explicit.
  4. 4.Build the evidence we do not have. Replicate the change finding elsewhere, and design a study linking scores to something observed from outside the instrument.
  5. 5.Republish these figures as the population grows. Every number here is generated from the platform data by script, so this can be regenerated rather than rewritten.

Who did this analysis, and what that is worth

The analysis was carried out by an AI system (Claude, made by Anthropic), working directly against the platform database under my direction. I am saying so for the same reason we published the unflattering findings: you should be able to judge the work by how it was done.

What that buys is reproducibility, since every figure traces to a script that can be re-run and nothing was transcribed by hand, and an analyst with no career, grant or citation record riding on the answer. What it does not buy is independence. The work was commissioned by the company that sells the instrument, run on that company’s data, and reviewed by me before publication. An AI engaged by a vendor is not a disinterested third party and I am not going to claim it is.

Nor does it buy freedom from error. Three conclusions reached while producing this paper were wrong:

First concludedWhy it was wrongHow it was caught
The five dimensions do not separate; the data cannot settle itThe test was confounded by shared questions, in the opposite direction to the one assumed.I asked whether the framing was more negative than the evidence warranted.
Individuals cannot be matched across repeat assessmentsStated three times from incomplete checks. An employee number was populated for every participant, and it is the basis of this paper’s strongest finding.I asked for it to be looked at again.
A client’s three assessments formed one seriesOnly two of them did. Treating the third as part of the series produced an apparent decline that was partly an artefact.I knew the client’s history.
The Pattern Is The Point

In every case the error was caught by a person who knew the business, not by the system that made it. So the honest description of what an AI analyst contributed here is speed, consistency and reproducibility inside a process that still needed someone able to say ‘that does not sound right’. Anyone claiming an AI is more objective than a human institution should look at that table first.

We would rather this were checked than believed. If you work on measurement in this field and want to re-run any of it, we will help. The findings we would most like someone else to attack are the change result and the dimension structure.

Read the full study

Everything above is the short version. The white paper sets out how each figure was produced, reports the numbers this page leaves out, states the limitations in full, and shows the workings well enough for someone else to disagree with them properly. Anyone evaluating the instrument, rather than reading about it, should start there.

The Full Analysis

Download the complete white paper

Thirteen pages, including the method, the per-scale figures, the statistical tests behind every claim, and what we are changing as a result.

Download PDF

Continue Exploring

More Thought Leadership

Shop